Driving behavior identification method and device
By generating attention zone sequences and dynamically adjusting budget curves, the driver's intentions and hand states are identified, solving the problem of inaccurate driver distraction behavior recognition in existing technologies, and achieving more precise intervention timing and improved safety.
Patent Information
- Application Number
- CN202511742571.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies struggle to effectively identify driver distraction, especially in smart cockpit scenarios. Current methods are difficult to correlate with driver gaze or gestures, resulting in imprecise use of thresholds and weights, which limits the applicability and widespread adoption of these technologies.
By generating attention zone sequences, identifying the driver's intentions, and dynamically adjusting the attention zone budget curve, combined with the driver's hand occupancy status, the timing of intervention and risk level are determined, providing personalized driving behavior prompts.
It increases tolerance for normal driving behavior and decreases tolerance for non-driving behavior, improving the accuracy of intervention timing and driving safety, and supporting implementation and promotion in different scenarios.
Smart Images

Figure CN121582902A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent cockpit, and particularly relates to a driving behavior recognition method and device. BACKGROUND
[0002] Under the background of rapid development of intelligent cockpit and in-vehicle infotainment system, the driver inevitably has a visual line shift and hand operation in the driving process, such as checking the instrument, confirming the rearview mirror, short pressing the central control or interacting with the mobile phone and the like. A large number of researches and accident reviews show that the duration of gaze and the frequency of repeated entry of "non-front road surface" are highly related to the distraction risk, but not all deviations are equal to danger, for example, the legal mirror confirmation and short-time instrument reading are necessary for completing the driving task. Therefore, in the intelligent cockpit scenario, the detection of the distraction risk is particularly important.
[0003] In the prior art, the visual gaze tracking and head pose estimation are mainly used as the core to alarm the duration and number of times of the driver "leaving the front road surface"; some schemes also add the hand detection mode to determine whether the hand operation of the driver has the distraction or risk probability. In the above prior art, the threshold and weight used are mostly used under the global or weak conditionalization, and it is also difficult to associate with the driver's eye direction or hand gesture, which limits the influence scene and is also difficult to land and promote. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a driving behavior recognition method and device, which fully considers the driving task and the hand occupation state of the driver, reasonably recognizes the distraction behavior of the driver and determines the intervention time, and thus relieves the above technical problems.
[0005] In a first aspect, an embodiment of the present application provides a driving behavior recognition method, the method comprising: in response to a change in a focus area corresponding to a line of sight of a driver, generating a focus area sequence; wherein the focus area is a cabin partition presented based on a driver's perspective; the focus area sequence is a sequence of identifiers of a plurality of target focus areas recorded in time sequence, the target focus area is a focus area corresponding to a change in the line of sight within a preset continuous time period; identifying an intention of the driver based on the focus area sequence to obtain an intention label corresponding to the intention and a reasonable duration expectation of the intention; dynamically adjusting a focus area budget curve of each target focus area preconfigured according to the intention label and the reasonable duration expectation to obtain a dynamic budget curve corresponding to each target focus area; determining the behavior of the target focus area based on the dynamic budget curve to determine whether the behavior of the target focus area reaches a preset intervention condition; if so, generating an intervention strategy including a risk level and an intervention timing to prompt the driving behavior of the driver based on the intervention strategy.
[0006] In combination with the first aspect, a first possible implementation manner of the first aspect is provided, wherein the step of generating the focus area sequence in response to a change in the focus area corresponding to the line of sight of the driver comprises: determining the number of the focus areas intersected by the line of sight direction based on the line of sight; if the number of the focus areas is one, the focus area is a target focus area, and the identifier of the target focus area is output; if the number of the focus areas is more than one, the target focus area is determined from the plurality of focus areas according to a pre-set focus area priority and a hand-reachable object contained in the focus area, and the identifier of the target focus area; the focus area sequence is generated based on the identifier of the target focus area within the preset continuous time period.
[0007] In combination with the first aspect, a second possible implementation manner of the first aspect is provided, wherein the step of identifying the intention of the driver based on the focus area sequence to obtain the intention label corresponding to the intention and the reasonable duration expectation of the intention comprises: obtaining a pre-constructed focus area semantic mapping, the focus area semantic mapping records the identifier of each focus area, the target object identifier of the hand-reachable object contained in each focus area, and the mapping relationship of the spatial boundary of each focus area; based on the target object identifier and the driving task semantics corresponding to each focus area recorded in the focus area semantic mapping, a candidate intention corresponding to each target focus area is obtained; the dwell time sequence information of the line of sight in the target focus area is extracted from the focus area sequence, and the intention label representing the intention and the corresponding reasonable duration expectation are generated based on the dwell time sequence information and the candidate intention.
[0008] With reference to the first aspect, in a third possible implementation of the first aspect, the step of dynamically adjusting the attention budget curve of each of the target attention areas according to the intention label and the reasonable duration expectation comprises: determining an adjustment direction of the attention budget curve of each of the target attention areas based on the intention label and the reasonable duration expectation; and dynamically adjusting the attention budget curve of each of the target attention areas based on the adjustment direction.
[0009] With reference to the third possible implementation of the first aspect, in a fourth possible implementation of the first aspect, the step of determining the adjustment direction of the attention budget curve of each of the target attention areas based on the intention label and the reasonable duration expectation comprises: if the intention label is an intention label corresponding to a driving task, and the reasonable duration expectation is a short duration expectation requirement, the adjustment direction of the attention budget curve is a widening direction; and if the intention label is an intention label corresponding to a non-driving task, and the reasonable duration expectation is a non-driving task occupation time, the adjustment direction of the attention budget curve is a tightening direction.
[0010] With reference to the fourth possible implementation of the first aspect, in a fifth possible implementation of the first aspect, the step of dynamically adjusting the attention budget curve of each of the target attention areas based on the adjustment direction comprises: obtaining a hand occupation state of the driver; determining an adjustment intensity of the attention budget curve in the adjustment direction based on the hand occupation state; and dynamically adjusting the attention budget curve of each of the target attention areas based on the adjustment direction and the adjustment intensity.
[0011] With reference to the fifth possible implementation of the first aspect, in a sixth possible implementation of the first aspect, the step of determining the adjustment intensity of the attention budget curve in the adjustment direction based on the hand occupation state comprises: when the hand occupation state is a single-hand occupation driving state, and the intention label is an intention label corresponding to a non-driving task, increasing a tightening amplitude in the tightening direction according to a pre-set tightening strategy; and when the hand occupation state is a double-hand holding driving state, and the intention label is an intention label corresponding to a driving task, decreasing a widening amplitude in the widening direction according to a pre-set widening strategy.
[0012] With reference to the first aspect, the embodiments of the present application provide a seventh possible implementation manner of the first aspect, wherein the dynamic budget curve is used to represent a continuous staying time length criterion of the gaze landing point in the target attention area, and a repeated number criterion of the gaze landing point repeatedly entering the target attention area within a preset time range; the step of determining the behavior of the target attention area based on the dynamic budget curve to determine whether the behavior of the target attention area reaches a preset intervention condition comprises: extracting a behavior sequence corresponding to the attention area sequence in the preset continuous time period; determining whether a cumulative continuous staying time length in the behavior sequence reaches the continuous staying time length criterion represented by the dynamic budget curve, or whether a cumulative repeated number in the behavior sequence reaches the repeated number criterion represented by the dynamic budget curve; and if either of the determination results is yes, determining that the behavior of the target attention area reaches the preset intervention condition.
[0013] With reference to the seventh possible implementation manner of the first aspect, the embodiments of the present application provide an eighth possible implementation manner of the first aspect, wherein the step of generating the intervention strategy comprising a risk level and an intervention time comprises: taking a time point at which the behavior sequence first reaches the continuous staying time length criterion or the repeated number criterion as the intervention time; determining an intervention rule of the target attention area based on a task attribute of the target attention area, determining the risk level according to the intervention rule and the continuous staying time length criterion or the repeated number criterion reached by the behavior sequence; and determining the intervention strategy based on the intervention time and the risk level.
[0014] In a second aspect, the embodiments of the present application further provide a driving behavior recognition device, the device comprising: a response module configured to generate a sequence of attention zones in response to a change in a line-of-sight landing point of a driver corresponding to the attention zones; wherein the attention zones are cabin partitions presented based on a driver's perspective; the sequence of attention zones is a sequence of identifiers of a plurality of target attention zones recorded in time sequence, the target attention zones being the attention zones corresponding to the change in the line-of-sight landing point within a preset continuous time period; a recognition module configured to recognize an intention of the driver based on the sequence of attention zones, and obtain an intention label corresponding to the intention and a reasonable duration expectation of the intention; an adjustment module configured to dynamically adjust an attention zone budget curve of each of the target attention zones pre-configured according to the intention label and the reasonable duration expectation, and obtain a dynamic budget curve corresponding to each of the target attention zones; a judgment module configured to judge the behavior of the target attention zones based on the dynamic budget curve, to determine whether the behavior of the target attention zones reaches a preset intervention condition; and an intervention module configured to generate an intervention strategy including a risk level and an intervention timing when the judgment result of the judgment module is yes, to prompt the driving behavior of the driver based on the intervention strategy.
[0015] The embodiments of the present application bring the following beneficial effects: The driving behavior recognition method and device provided by the embodiments of the present application can generate a sequence of attention zones in response to a change in a line-of-sight landing point of a driver corresponding to the attention zones; recognize an intention of the driver based on the sequence of attention zones, and obtain an intention label and a reasonable duration expectation; dynamically adjust an attention zone budget curve of each of the target attention zones pre-configured according to the intention label and the reasonable duration expectation, and obtain a dynamic budget curve corresponding to each of the target attention zones; judge the behavior of the target attention zones based on the dynamic budget curve, to determine whether the behavior of the target attention zones reaches a preset intervention condition; and if yes, generate an intervention strategy including a risk level and an intervention timing, to prompt the driving behavior of the driver based on the intervention strategy. The above-mentioned method of dynamically adjusting the attention zone budget curve can realize an increase in tolerance for normal driving behavior, and a decrease in tolerance when the driver exhibits non-driving behavior, so that the determined intervention timing or risk level is more consistent with the driving behavior of the user, which can effectively avoid interrupting legal driving behavior, and help to promote early intervention of non-driving behavior, thus helping to land and promote, and helping to improve driving safety.
[0016] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and achieved by means of the structures particularly pointed out in the description and the claims.
[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a method for recognizing driving behavior provided in an embodiment of the present invention; Figure 2 A cockpit schematic diagram provided for an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the output of intervention timing and risk level according to an embodiment of the present invention; Figure 4 A schematic diagram of the structure of a driving behavior recognition device provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Currently, existing technologies primarily focus on visual gaze tracking and head pose estimation, supplemented by AOI (region of interest) mapping, to coarsely estimate the driver's gaze point. They then use fixed or speed-based function thresholds to issue warnings for the duration and frequency of "leaving the road ahead." Some solutions also incorporate hand-holding detection, mobile phone target detection, or central control touch event detection, employing heuristic rules for multimodal fusion, such as "warning if leaving the road ahead for more than 2 seconds or repeatedly entering a non-frontal area N times within T seconds." Other studies use temporal models (such as LSTM) to output distraction probabilities or risk scores, then accumulate decisions within a window using uniform weights. A few works attempt to slightly relax the mirror confirmation threshold based on turn signal status or navigation prompts, but overall, they still use global thresholds or generalized weights for generalization adjustments. Furthermore, existing technologies typically use a cross-regional accumulation approach (without distinguishing task attributes) to provide "distracted / not distracted" or coarse-grained behavior labels. Their explanatory power is mostly limited to the level of "exceeding the threshold," making it difficult to directly link specific driving tasks with when to intervene in distracted behavior. This makes it difficult to implement and promote, and also makes it difficult to effectively identify distracted driving behavior.
[0022] Based on this, the driving behavior recognition method and device provided in this embodiment of the invention can fully consider the driving task and the driver's hand occupation status, reasonably identify the driver's distracted behavior and determine the timing of intervention, thereby alleviating the above-mentioned technical problems.
[0023] To facilitate understanding of this embodiment, a method for recognizing driving behavior disclosed in this embodiment of the invention will first be described in detail.
[0024] In one possible implementation, embodiments of the present invention provide a method for recognizing driving behavior, such as... Figure 1 The flowchart shown illustrates a method for recognizing driving behavior, which includes the following steps: Step S102: In response to a change in the attention zone corresponding to the driver's line of sight, an attention zone sequence is generated; In this embodiment of the invention, the attention area is a cockpit partition based on the driver's perspective; the attention area sequence is a sequence of identifiers of multiple target attention areas recorded in chronological order, and the target attention area is the attention area corresponding to the change of the line of sight within a preset continuous time period.
[0025] Specifically, in this embodiment of the invention, the cabin partitioning is based on driving task semantics. Driving task semantics defines the categories of tasks directly related to driving safety during the driving process, including monitoring the road ahead, checking the rearview mirror, reading the instrument panel, and short-term operation of the central control system. It is used to standardize the association between the setting of attention areas and subsequent judgment logic. Therefore, in this embodiment of the invention, according to driving task semantics, when partitioning the cabin, the road ahead, the instrument panel, the left rearview mirror, the right rearview mirror, the central control system, and the passenger / mobile phone area are set as attention areas.
[0026] Furthermore, regarding the aforementioned attention zones, in this embodiment of the invention, the driver's gaze point and behavior can be monitored in real driving scenarios so as to respond when the attention zone of the gaze point changes.
[0027] Step S104: Identify the driver's intention based on the attention area sequence, and obtain the intention label corresponding to the intention and the expected reasonable duration of the intention; In real-world driving scenarios, a driver's gaze is typically focused on the area of attention corresponding to the road ahead, meaning real-time monitoring of the roadside. However, if the driver becomes distracted, they may occasionally shift their gaze elsewhere, resulting in their gaze falling on a different area of attention. In such cases, the cockpit-mounted image monitoring device can detect the change in the gaze's focus and respond accordingly. This can be done by recording which area of attention the driver's gaze falls on and the duration of that gaze's focus within that area. Then, based on the identifier of the area of attention, the duration of the gaze's focus within that area is recorded, thereby generating the aforementioned attention area sequence. Therefore, the attention area sequence in this embodiment of the invention is actually a sequence of attention area identifiers recorded in chronological order, used to characterize the changes in the area of attention corresponding to the driver's gaze within a recent continuous period.
[0028] Furthermore, in this embodiment of the invention, interpreting the attention zone sequence can reveal the driver's intention, thereby facilitating the identification of dangerous actions such as driver visual deviation and determining the timing of subsequent intervention.
[0029] Step S106: Based on the intent label and the expected reasonable duration, dynamically adjust the attention budget curve of each pre-configured target attention area to obtain the dynamic budget curve corresponding to each target attention area. Step S108: Determine the behavior of the target attention area based on the dynamic budget curve to determine whether the behavior of the target attention area meets the preset intervention conditions; Step S110: If yes, generate an intervention strategy that includes risk level and intervention timing to prompt the driver's driving behavior based on the intervention strategy.
[0030] In practical use, the attention zone budget curve is actually a risk budget curve used to determine behaviors related to gaze deviation, including the permissible range of gaze shifting behavior, the duration of the shift, and the number of permissible gaze shifts, etc., and to determine whether the intervention conditions are met. In this embodiment of the invention, the attention zone budget curve is dynamically adjusted according to the driver's intentions. For example, the threshold for the permissible shifting duration and the number of permissible gaze shifts can be dynamically adjusted. This can increase the tolerance for normal driving behaviors, such as checking the left and right rearview mirrors, and decrease the tolerance when the driver shows signs of non-driving behavior, thereby obtaining a more reasonable and accurate risk level and intervention timing.
[0031] Therefore, the driving behavior recognition method provided in this embodiment of the invention can increase tolerance for normal driving behavior and decrease tolerance when the driver engages in non-driving behavior by dynamically adjusting the attention zone budget curve. This makes the determined intervention timing or risk level more consistent with the user's driving behavior, effectively avoids erroneous interruption of legitimate driving behavior, and helps to improve the early intervention of non-driving behavior. This not only helps in implementation and promotion but also helps to improve driving safety.
[0032] In this embodiment of the invention, in step S102 above, when generating the attention zone sequence and recognizing the driver's intention, it is often based on attention zone semantic mapping. Specifically, attention zone semantic mapping is to associate pre-set attention zones such as the road ahead, instrument panel, left rearview mirror, right rearview mirror, central control, passenger / mobile phone area, etc., with the unique identifier, spatial boundary and target object identifier of each attention zone, forming a mapping relationship from the identifier of the attention zone to its spatial boundary and the target object identifier.
[0033] Typically, the above mapping relationship is established in the cockpit coordinate system. Specifically, the cockpit coordinate system is established with reference to the center of the steering wheel and the center of the instrument panel. Based on this, attention zones are further defined, whose members include the road surface area in front, the instrument panel area, the left rearview mirror area, the right rearview mirror area, the central control area, and the passenger / mobile phone area. Then, a unique identifier is assigned to each attention zone, and the corresponding spatial boundary and target object identifier are determined in the cockpit coordinate system.
[0034] The spatial boundaries of each attention zone are described using a visual field that aligns with the driver's line of sight. For example, the road surface area corresponds to a fan-shaped visual field along the windshield opening, covering the main visual field of the road in front of the vehicle; the instrument panel area and the center console area each correspond to a rectangular visual field centered on their respective physical centers, defining the effective range of the instrument display area and the center console operation area; the left and right rearview mirror areas correspond to a narrow visual field pointing towards the center of their respective mirror surfaces, to accommodate the elongated shape of mirror observation; and the passenger / mobile phone area corresponds to a partial visual field located between the front passenger seat and the seat, to cover the common mobile phone placement and interaction areas in the vehicle.
[0035] Furthermore, to support subsequent processes such as intent recognition, target object identifiers for reachable objects are set within each attention zone. These include object identifiers directly associated with the attention zone, such as the instrument display area, rearview mirror surface, central control buttons, passenger-side placement area, and mobile phone placement area. These identifiers provide semantic anchors at the object level when gaze and hand position are combined. This forms an attention zone semantic mapping, which represents the semantic association from the attention zone's identifiers to its spatial boundaries and the set of target object identifiers, used for unified management of the binding information between the eye, hand, and object.
[0036] When generating the attention area sequence as described above, the actual process involves determining the attention area where the gaze point will ultimately be located, i.e., the target attention area. Specifically, when generating the attention area sequence, based on the gaze point, the number of attention areas where the spatial boundary intersects with the gaze direction is determined. If there is only one attention area, it is the target attention area, and its identifier is output. If there are multiple attention areas, the target attention area is determined from among the multiple attention areas according to the pre-set attention area priority and the reachable objects contained in the attention area, along with its identifier. The attention area sequence is then generated based on the identifiers of the target attention areas within a preset continuous time period.
[0037] For example, using the driver's gaze point as input, when the gaze direction corresponding to the gaze point intersects with the spatial boundary of any attention area, that attention area is designated as the target attention area, and its identifier is output. When the gaze direction intersects with two or more adjacent attention areas simultaneously, a single attention area is selected as the target attention area according to attention area priority, i.e., the order of the road ahead, instrument panel, rearview mirror, central control panel, and passenger / mobile phone area, and its corresponding identifier is obtained to avoid interference with short-term confirmations related to driving tasks. Then, using reachable objects as input, which include the steering wheel, steering lever, central control buttons, and the passenger-side mobile phone storage area, the identifier is output based on the reachable objects. The binding relationship between the target object identifier and the attention area of the reachable object is established, and the attention area corresponding to the reachable object is output, along with its corresponding identifier. Finally, the consistency between the attention area identifier obtained from the gaze point and the attention area identifier obtained from the reachable object is determined. If they are consistent, the current attention area is directly identified as the target attention area. If they are inconsistent, and the reachable object corresponds to the passenger / mobile phone area, the attention area obtained from the reachable object is taken as the target attention area. This is used to increase the recognition priority of driver distraction detection when non-driving tasks occur. In other inconsistent cases, disambiguation is performed according to the above attention area priority order to ensure that the determination of the attention area is consistent with the semantics of the driving task.
[0038] Furthermore, in the time dimension, the attention areas where the gaze falls are observed using time as an index, and the identifier of the attention area where the gaze falls at each moment is recorded, thereby generating the aforementioned attention area sequence in chronological order. In addition, to ensure that the recorded gaze points accurately reflect the driver's attention shifts, this embodiment of the invention can also impose a constraint on the attention area sequence regarding the shortest dwell time. That is, a valid attention shift is only confirmed when the dwell time of an identifier of a certain attention area within consecutive frames reaches a threshold. Specifically, if the gaze point's dwell time in a certain attention area exceeds a certain duration, or the length of consecutive frames reaches a length threshold, the gaze point is determined to have stayed in that attention area, thereby suppressing the interference of the driver's occasional squinting on the decision logic.
[0039] Furthermore, when identifying the driver's intent based on the aforementioned attention zone sequence, it is necessary to further utilize attention zone semantic mapping. This attention zone semantic mapping is used to restrict the interpretation of intent categories to only within the target attention zone and to establish a correspondence between the expected reasonable duration of the intent and the task attributes of the attention zone. Finally, the attention zone budget curve is adjusted based on the intent, thereby binding the intervention timing with the specific attention zone and its task attributes, directly serving the risk assessment of dangerous actions such as line-of-sight deviation.
[0040] Furthermore, in step S104 above, when identifying the driver's intention, a pre-constructed semantic mapping of attention areas can be obtained. This semantic mapping is established in the aforementioned manner, that is, it records the identifier of each attention area, the target object identifier of the hand-accessible object contained in each attention area, and the mapping relationship of the spatial boundaries of each attention area. Then, based on the target object identifier and driving task semantics corresponding to each attention area recorded in the semantic mapping of attention areas, candidate intentions corresponding to each target attention area are obtained. The dwell time sequence information of the gaze point in the target attention area is extracted from the attention area sequence. Based on the dwell time sequence information and candidate intentions, an intention label representing the intention and the corresponding reasonable duration expectation are generated. Among them, the dwell time sequence information is used to represent that the dwell time of the gaze point in a certain attention area exceeds a certain duration, or the length of consecutive frames reaches a length threshold.
[0041] Specifically, the aforementioned intention recognition process is actually a process of interpreting the relationship between the driver's gaze point, hand occupancy status, and the attention area to which the hand can reach the object—that is, the eye-hand-object ternary relationship or the eye-hand-object combination relationship—and then outputting the intention and a reasonable expected duration.
[0042] In practical use, an intent-conditional gaze-hand joint recognition network can be pre-constructed. The input of this network includes the driver's gaze direction vector, gaze point, and set of objects reachable by the hand, as well as the attention area semantic mapping and attention area identifier. The input is the intent label representing the intent and the expected reasonable duration.
[0043] In practical applications, the aforementioned intention-conditional gaze-hand joint recognition network may include an attention zone conditional gating unit, a gaze interpretation unit, a hand occupancy determination unit, a relationship interpretation unit, an intention generation unit, and a duration prediction generation unit. Specifically, the attention zone condition gating unit ensures that the subsequent intent interpretation process is confined to the relevant attention zone, i.e., limited to the aforementioned target attention zone; the gaze interpretation unit interprets the gaze vector to obtain the gaze point and the corresponding target attention zone; the hand occupancy determination unit determines the hand occupancy state and the identifier of the attention zone to which the reachable object belongs; the relationship interpretation unit performs consistency determination and association mapping on the gaze point, hand occupancy state, and the identifier of the attention zone to which the reachable object belongs, forming a gaze-hand-object combination relationship for determining intent; and the intent generation unit outputs the intent and corresponding intent label based on the gaze-hand-object combination relationship and the semantics of the driving task. The intent includes mirror confirmation, instrument reading, short-term central control operation, and handheld phone / non-driving task, i.e., determining the purpose of the driver's gaze shift, such as whether it is necessary to perform driving tasks such as mirror confirmation, instrument reading, or short-term central control operation, or to perform non-driving tasks such as handheld phone, and maintaining a one-to-one correspondence between the intent and the attention zone. Furthermore, in the duration expectation generation unit, a reasonable duration expectation corresponding to the intent is generated based on the intent tag and the task attribute of the current attention area, combined with the dwell time sequence information in the current attention area sequence. In this embodiment of the invention, the reasonable duration expectation is represented by category-level results as short-term confirmation, short-term reading, short-term operation, or non-driving task occupancy.
[0044] Typically, the process of jointly interpreting the eye-hand-object ternary relationship described above, in the aforementioned intention-conditional gaze-hand joint recognition network, takes the output of the attention zone conditional gating unit as a prerequisite. The eye direction interpretation unit and the hand occupancy determination unit provide identifiers in sequence, and the relationship interpretation unit completes consistency determination and association mapping. Finally, the intention generation unit and the duration expectation generation unit provide intention labels and reasonable duration expectations, which are used to determine the intervention timing for dangerous actions such as gaze deviation.
[0045] Specifically, in this embodiment of the invention, before determining the timing of intervention, it is necessary to further determine the driver's gaze deviation-related behaviors, that is, to determine whether intervention is required. In this process, this embodiment of the invention is based on the attention zone budget curve. That is, in the above step S106, the attention zone budget curve is first dynamically adjusted, and then it is further determined whether intervention is required.
[0046] Specifically, in step S106, dynamically adjusting the attention area budget curve includes: determining the adjustment direction of the attention area budget curve for each target attention area based on the intent label and the expected reasonable duration; and dynamically adjusting the attention area budget curve for each target attention area based on the adjustment direction.
[0047] In practical use, the aforementioned attention zone budget curve is also called the attention zone risk budget curve. For the function of dynamically adjusting the aforementioned attention zone budget curve, the in-vehicle semantic attention zone risk budget gating architecture can be preset. Specifically, it can include a budget setting unit, a budget adjustment unit, and a budget determination unit. Each unit is used to generate the basic attention zone budget curve, perform dynamic adjustment of the attention zone budget curve, and determine the timing of intervention, etc.
[0048] The setting of the attention zone budget curve can be implemented in the budget setting unit. For example, the budget setting unit generates a basic attention zone budget curve for each attention zone based on the semantic mapping between driving task and attention zone. Typically, the inherent requirements of driving tasks for each attention zone are specifications for the necessity and tolerability of forward road monitoring, rearview mirror confirmation, instrument reading, short-term central control operation, and non-driving tasks in the passenger / mobile phone area in safe driving, which are used to constrain the composition of the basic attention zone budget curve.
[0049] Attention zone budget curves are budget sets defined for a single attention zone, including two types of parameters: continuous dwell budget and repetition tolerance. These are used to determine the permissible range of gaze shifting behavior under unintentional modulation. The permissible duration of gaze shifting and repetition tolerance under unintentional modulation are parameter pairs in the basic attention zone budget curve used to limit the tolerable time range for a single continuous dwelling and the tolerable range of repetitions within a preset observation period.
[0050] Specifically, the attention zone budget curve in this embodiment of the invention is set as follows: In the budget setting unit, the mapping between driving task semantics and attention zone semantics is used as input. First, based on the attention zone semantic mapping, the attention zone set is determined to be the road surface ahead, instrument panel, left rearview mirror, right rearview mirror, central control unit, and passenger / mobile phone area. Then, a task attribute is assigned to each attention zone as confirmation, reading, operation, or non-driving occupancy. Next, the tolerance strategy under different task attributes is constrained according to the driving task semantics. The confirmation category supports short-term confirmation and necessary verification; the reading category supports short-term reading and allows a small number of repetitions; and the operation category supports a single short-term operation and restricts continuous operation. Multiple operations, including non-driving occupancy categories, are used to limit continuous dwelling and repetitive behavior. Under the above constraints, combined with the dwelling time sequence information in the current attention zone sequence, a corresponding attention zone budget curve is generated for each attention zone. The attention zone budget curve consists of a continuous dwelling budget and a repetition tolerance, that is, the continuous dwelling time of the gaze point corresponding to each attention zone, and the number of times the gaze point is allowed to repeatedly enter the attention zone, and it is kept in one-to-one correspondence with the task attribute of the attention zone. Finally, the attention zone budget curves of each attention zone are summarized into a basic budget that can be called by the budget determination unit, which is used to determine the intervention timing for gaze deviation-related behaviors when there is no intentional modulation.
[0051] Specifically, the attention zone budget curves for each attention zone are as follows: For the road ahead, the attention zone budget curve is set to a baseline where the driver's gaze is unrestricted, so that no budget is consumed when the driver's gaze is fixed on the road ahead; For the left and right rearview mirrors, the attention zone budget curve is set to a confirmation type curve that supports one short-term confirmation and allows one optional review, so as to maintain tolerance under legal confirmation before lane changes or turns; For the instrument panel, the attention zone budget curve is set to a short-term reading type curve that allows a small number of repeated readings, so as to provide limited tolerance when checking vehicle speed or warnings while driving; For the central control unit, the attention zone budget curve is set to an operation type curve that allows one short-term operation and restricts multiple consecutive operations, so as to limit continuous occupation of the central control unit; For the passenger / mobile phone area, the attention zone budget curve is set to a non-driving occupation type curve that restricts continuous stay and does not allow repeated non-driving occupation, so as to reduce tolerance when signs of non-driving tasks appear.
[0052] Specifically, the above setting process can be represented as follows:
[0053] in, Budget for the duration of a single continuous dwell time allowed during unintentional modulation; The number of repeated entries allowed within the preset observation period T; Lmin is the minimum stay constraint, providing a lower limit for the time dimension to eliminate invalid scans; This is an indicator function used to select the item corresponding to the task attribute; where the task attribute... ∈{"Confirmation Class","Read Class","Operation Class","Non-Driving Occupation Class"}; Budget increments for continuous dwell time for confirmation, reading, operation, and non-driving occupancy categories; These represent the number of times each of the four categories was tolerated during the observation period.
[0054] The above representation method can unify the two dimensions of continuous dwell time and repetition count, and ensure consistency with the semantic mapping of the attention area.
[0055] Furthermore, the attention zone budget curve is implemented by binding the continuous stay budget with the repetition tolerance. The continuous stay budget is used to limit the tolerable range of a single continuous stay, and the repetition tolerance is used to limit the number of repetitions within a preset observation period. The two together constitute the judgment criteria in the case of unintentional modulation and are consistent with the task attributes of each attention zone.
[0056] Furthermore, in this embodiment of the invention, the process of dynamically adjusting the attention zone budget curve described above ultimately achieves online widening or tightening of the attention zone budget curve of the target attention zone in the budget adjustment unit. For example, when the intention is mirror confirmation, instrument reading, or short press of the central control, the budget is appropriately widened in the attention zone to tolerate one or a small number of re-checks; when the intention is a handheld mobile phone or a non-driving task, the budget is tightened to shorten the tolerance time and reduce the repetition tolerance. The dynamic adjustment is driven only by the intention output, which is different from the existing fixed threshold or generalized weight adjustment, so that the intervention timing is directly bound to the semantics of the driving task, thereby improving the accuracy of the intervention timing.
[0057] Furthermore, when dynamically adjusting the attention zone budget curve, if the intent label is the intent label corresponding to the driving task, and the reasonable duration expectation is a short-term expectation requirement, then the adjustment direction of the attention zone budget curve is the widening direction; if the intent label is the intent label corresponding to a non-driving task, and the reasonable duration expectation is occupied by a non-driving task, then the adjustment direction of the attention zone budget curve is the tightening direction.
[0058] Furthermore, when making dynamic adjustments based on the adjustment direction, it is necessary to further obtain the driver's hand occupancy status; determine the adjustment intensity of the attention area budget curve in the adjustment direction based on the hand occupancy status; and then dynamically adjust the attention area budget curve of each target attention area based on the adjustment direction and adjustment intensity. Specifically, when the hand occupancy status is a single-handed driving state, and the intention label is an intention label corresponding to a non-driving task, the tightening amplitude is increased in the tightening direction according to a pre-set tightening strategy; when the hand occupancy status is a two-handed driving state, and the intention label is an intention label corresponding to a driving task, the widening amplitude is decreased in the widening direction according to a pre-set widening strategy.
[0059] For example, after locating the corresponding attention zone budget curve based on the current target attention zone identifier, a broadening or tightening strategy is selected based on the intent label and the expected reasonable duration. For instance, when the intent label is a driving task such as mirror confirmation or instrument reading, and the expected reasonable duration is a short-term expectation such as short-term confirmation or short-term reading, the continuous dwell budget is broadened online within the corresponding target attention zone to cover the expectation, and the repetition tolerance is increased to allow one review or a small number of repetitions; when the intent label is a driving task involving short-term operation of the central control system, and the expected reasonable duration is short... During operation, within the corresponding target attention area, the continuous dwell budget is expanded online to cover a short-term operation. Simultaneously, the repetition tolerance is set to disallow multiple consecutive operations. When the current attention area sequence is displayed within a short-term window, and the same short-term central control operation occurs again, the repetition tolerance can be tightened immediately to limit continuous occupation. When the intent label is a non-driving task such as holding a mobile phone, and the reasonable expected duration of non-driving task occupation, the continuous dwell budget in the passenger / mobile phone area is tightened online to be lower than this expectation, and the repetition tolerance is reduced to the lowest level to facilitate rapid judgment. Subsequently, the intensity is adjusted based on the hand occupation status. For example, when the hand occupation status is a one-handed driving state with one hand off the handlebars, and the intent label is a non-driving task, the tightening strategy is increased. When the hand occupation status is a two-handed driving state, and the intent label is a driving task corresponding to mirror confirmation, instrument reading, or a short-term central control operation, the expansion strategy is appropriately reduced to maintain necessary constraints.
[0060] Based on the above adjustment process, the final output is a dynamic budget curve corresponding to the identifier of the current target attention area, which can be called by the budget judgment unit to determine the timing of intervention.
[0061] The aforementioned dynamic budget adjustment process is driven solely by intent tags, reasonable duration expectations, and hand occupancy status, and is implemented at the attention zone level by widening or tightening. It does not involve uniform modification of general weighting factors, ensuring that the timing of intervention remains directly bound to the semantics of the driving task.
[0062] For ease of understanding, Figure 2 A schematic diagram of a cockpit is shown, in which, Figure 2 In the center, the cockpit's attention areas are presented from the driver's shoulder-back perspective: including the road ahead, instrument panel, left and right rearview mirrors, center console, and passenger / phone area. The gaze arrow falls on the instrument panel, corresponding to the intention of "instrument reading," with the annotation "pause / minor repetition" indicating expanded content; the passenger / phone attention area is marked with "intervention" at the edge and displays the status of the phone and hand. Figure 2 The instrument panel also displays the intent and dynamic budget of the subject area. The dynamic adjustment of the attention budget curve for the corresponding attention area is a widening strategy to cover pauses and minor repetitions. The intent shown in the passenger / mobile phone attention area is the intent label corresponding to non-driving tasks, such as holding / focusing on the mobile phone, which are non-driving occupancy states. At this time, the attention budget curve tightens to reduce repetitions to zero. At the same time, the widening is weakened when the hands are "both hands on the handlebars", and the tightening is increased when the hands are "one hand off the handlebars" during non-driving occupancy. Figure 2 As can be intuitively seen, in this embodiment of the invention, the intent is the sole driving force. Within a specific attention area, the attention area budget curve is widened or tightened online, directly binding the intervention timing with the task attribute. This avoids accidentally interrupting legitimate mirror confirmation, instrument reading, and short-term central control operations. At the same time, the tolerance for non-driving activities such as using a handheld mobile phone is rapidly reduced in terms of duration and repetition, thus triggering the intervention conditions.
[0063] Furthermore, the dynamic budget curve obtained after the above dynamic adjustment is used to characterize the benchmark for the continuous dwell time of the gaze point in the target attention area, and the benchmark for the number of times the gaze point repeatedly enters the target attention area within a preset time range. When determining whether the intervention condition is met, it is necessary to further extract the behavior sequence within a preset continuous time period corresponding to the attention area sequence; determine whether the cumulative continuous dwell time in the behavior sequence reaches the benchmark for the continuous dwell time represented by the dynamic budget curve, or whether the cumulative number of repetitions in the behavior sequence reaches the benchmark for the number of repetitions represented by the dynamic budget curve; if any judgment result is yes, then it is determined that the behavior in the target attention area has met the preset intervention condition.
[0064] The aforementioned judgment process involves assessing the behavior within the current target attention zone to determine whether intervention conditions have been met and to output the risk level of driver distraction and the appropriate intervention timing. The aforementioned behavior sequence refers to a set of time-ordered pause segments formed by the identifier of the target attention zone within a preset time range, including the start and end times of each pause segment and the number of cross-segment entries, used to support cumulative judgment. Typically, based on the task attributes of the target attention zone and the exceeded budget item, a corresponding risk level is output, with the first time the limit is exceeded serving as the intervention timing. The output intervention strategy can be used for subsequent intervention execution.
[0065] Specifically, during the judgment process, based on the dynamically adjusted dynamic budget curve, the behavior sequence of the target attention area within a preset continuous time period is first extracted from the current attention area sequence. The continuous dwell time of the gaze point in the target attention area is taken as the cumulative amount of the dwelling channel, and the number of re-entries shortly after the gaze point leaves is taken as the cumulative amount of the repetition channel. Then, the adjusted continuous dwell time benchmark and repetition number benchmark in the dynamic budget curve are used as the comparison benchmarks for the dwelling channel and the repetition channel, respectively, and the cumulative amounts of the two channels are compared synchronously over time. When neither channel reaches the corresponding benchmark, the judgment remains that it has not exceeded the limit, that is, the intervention condition has not been met. When the cumulative amount of any channel reaches or exceeds the corresponding benchmark for the first time, it is determined that the intervention condition has been met, and the budget limit corresponding to that channel is marked as exceeded. The type of exceedance is determined to be either continuous dwell time exceedance or repetition number exceedance. The above judgment process is only carried out within the dynamic budget curve range corresponding to the target attention area and does not cross other attention areas, ensuring that the judgment result is consistent with the specific driving task attributes.
[0066] Furthermore, after determining that the intervention condition has been met, i.e., exceeding the limit, the intervention timing is taken as the moment when the continuous dwell time benchmark or the repetition count benchmark is first reached in the behavior sequence; and, the intervention rules for the target attention area are determined based on the task attributes of the target attention area, and the risk level is determined according to the intervention rules and the continuous dwell time benchmark reached in the behavior sequence, and / or the repetition count benchmark; then the intervention strategy is determined based on the intervention timing and the risk level.
[0067] For example, when the continuous dwell time reaches or exceeds the continuous dwell time benchmark, the intervention rule is first determined based on the task attributes of the current target attention area. For confirmation and reading tasks, the continuous dwell time is the primary threshold, and the number of repetitions is the secondary threshold. The continuous dwell time is used as the primary basis for determining the intervention rule, while the number of repetitions is used as a secondary basis. For operation tasks, a single short operation is the primary tolerance, and multiple consecutive operations are restricted. For non-driving occupancy tasks, any budget item is used as the primary threshold; that is, the continuous dwell time benchmark and / or the number of repetitions benchmark can both be used as the primary basis. Then, the moment the primary threshold is first reached or exceeded is used as the intervention timing. If the secondary threshold is triggered before the primary threshold, it is recorded as a warning, and escalated to intervention when the primary threshold is triggered. When the current target attention area is marked as the road ahead, the dynamic budget curve is stationary and unrestricted, and no intervention will occur. The process for determining intervention conditions involves several steps. When the current attention area is marked as the left or right rearview mirror or instrument panel, if the continuous dwell time exceeds the limit (i.e., reaches or exceeds the continuous dwell time benchmark), a confirmation or reading intervention rule is triggered. If the number of repetitions exceeds the limit (i.e., reaches the repetition count benchmark), a frequent review-related intervention rule is triggered. When the current target attention area is marked as the central control area, if the continuous dwell time exceeds the limit, or if the number of repetitions exceeds the limit within a short period, an operation-related intervention rule is output. When the current target attention area is marked as the passenger / mobile phone area, if any budget item exceeds the limit, a non-driving-related intervention rule is output. Finally, the determined intervention timing and corresponding risk level are output together as the intervention strategy. The risk level can be graded based on the exceeded budget item and its task attribute to reflect the intervention intensity.
[0068] Specifically, the above determination process can be expressed as the following formula:
[0069] In the formula, To be within the preset observation period length The effective dwell segment that satisfies the shortest dwell time constraint Lmin is z. t Markers indicating the area to be noted; For at any time The cumulative duration of continuous stay; To be within the preset observation period length The number of repetitions within; The set of valid dwell segments that are consistent with the current target attention zone marker; For indicator functions; For the shortest dwell time constraint; and These are the tolerances for dynamic continuous dwell time and dynamic repetition count, namely the aforementioned continuous dwell time benchmark and repetition count benchmark; and For both channels, an over-limit indication is provided; The earliest moment when any dynamic budget threshold is first reached; To match task attributes A consistent priority matrix is used to select channels based on task attribute preferences when triggering simultaneously. Confirmation and reading classes take priority over "continuous dwell budget", operation classes are mainly based on "continuous dwell budget" and constrained by "repetition tolerance", and non-driving occupancy classes are treated equally between the two channels. The selected overlimit type.
[0070] In obtaining and Afterwards, the risk level and intervention timing can be determined: When the current target attention zone is marked as the road ahead, since it is always stationary and unrestricted, no intervention strategy is output at this time; when the current target attention zone is marked as the attention zone corresponding to the left or right rearview mirrors or the instrument panel, if Then it is confirmed that the intervention rule is related to the reading class or the confirmation class. This confirms the intervention rule as related to frequent review; when the current target attention area is identified as the central control attention area, Or within a short period of time All were confirmed as intervention rules related to operation occupancy; when the current target attention area is identified as a passenger / mobile phone area, triggering any channel is confirmed as an intervention rule related to non-driving occupancy, and the intervention timing is directly determined. To maintain consistency with the over-limit of the dynamic budget curve.
[0071] In practical use, while outputting the intervention timing and risk level, the triggering reason can also be recorded. The triggering reason includes at least the identifier of the current target attention area, the intent label, the type of budget item that was exceeded (e.g., exceeding the limit for continuous stay duration or the limit for the number of repetitions), the time of exceeding the limit and the corresponding actual duration value, and may also include the start and end times of the stay segment related to the limit in the current attention area sequence. The triggering reason, together with the intervention timing and risk level, constitutes the intervention decision, which is used to achieve interpretable linkage prompts and intervention execution.
[0072] Specifically, the risk levels, intervention timings, and corresponding intervention strategies identified above can be organized into structured, interpretable outputs for use in cockpit human-machine interaction and audit records.
[0073] For example, if κ*="stay" and the target attention area is the attention area corresponding to the instrument panel, the output could be "Intervene when the continuous stay budget read by the instrument panel is reached or exceeded"; if κ*="rep" and the target attention area is the attention area corresponding to the left or right rearview mirrors, it could be "Frequent checks exceed the repetition tolerance"; if the target attention area is the passenger / mobile phone area, it could be "Non-driving occupancy has reached the dynamic budget threshold in terms of duration or repetition." The above outputs can be simultaneously sent to the cockpit's human-machine interface to generate front-end prompt text and be written into the audit log for subsequent compliance review and accountability determination.
[0074] For ease of understanding, Figure 3 This diagram illustrates the output of intervention timing and risk level, such as... Figure 3 As shown, the interpretable output is displayed using three information cards (a), (b), and (c). In (a), the attention area is the instrument attention area, the intent is "instrument reading", and the over-limit item is "continuous stay budget", that is, the continuous stay duration exceeds the limit. At the same time, the triggering reason lists the actual stay duration, dynamic budget (the above continuous stay duration benchmark), number of repetitions, and hand occupation status, and concludes with a low-risk warning. In (b), the attention area is the central control attention area, the intent is "central control short-term operation", and the number of repetitions exceeds the limit within the short window, triggering medium-risk intervention. (c) shows the passenger / mobile phone attention area, the intent is "non-driving occupation". At this time, the attention area budget curve tightens in both the continuous stay duration and the number of repetitions. Exceeding the limit in either dimension is a high-risk level. At this point, (c) presents "the attention zone, intent, over-limit items, intervention time, continuous dwell time and repetition number are all quantified, and the hand occupation status and risk level are given. In addition, different risk levels can be coded with colors, such as dark blue indicating a low or medium risk level, such as a legal short-term tolerance extension, and dark red indicating a high risk level, such as tightening and intervention when not driving, intuitively forming a traceable chain of triggering causes."
[0075] In summary, the driving behavior recognition method provided by this invention is an improved method for joint recognition of intent-conditional gaze and hand gestures, along with dynamic gating of attention zone risk budget. It uses semantic attention zone conditional joint decoding of the object-hand-object ternary relationship to generate an intent label semantically consistent with the current target attention zone and a reasonable expected duration. This intent is then used as the sole adjustment factor for the continuous dwell budget and repetition tolerance of each attention zone, enabling online widening or tightening. The improved method directly binds the intervention timing to the specific attention zone and its task attributes. In scenarios involving mirror confirmation, instrument reading, and short-term central control operation, the budget is appropriately widened with low-interference prompts to avoid mistakenly interrupting legitimate short-term behaviors. In non-driving occupancy scenarios involving passenger / mobile phone areas, the budget is tightened and repetition tolerance is reduced, making the triggering cause clearly attributable to the combination of "intent—attention zone—budget item," thereby improving the consistency and semantic alignment of the judgment and intervention scenarios.
[0076] Furthermore, the driving behavior recognition method of this invention can also realize a novel in-vehicle semantic attention area risk budget gating architecture. It can set basic attention area budget curves for attention areas such as the road ahead, instrument panel, left and right rearview mirrors, central control, and passenger / mobile phone areas. It employs two-channel gating with continuous dwell and repeated entry, superimposed with a minimum dwell constraint, and the cumulative judgment is strictly limited to the current attention area, without cross-area mixing. It replaces the generalized adjustment of global thresholds or uniform weights with task attribute-driven attention area budget curves, forming differentiated tolerance boundaries for confirmation, reading, operation, and non-driving occupancy, suppressing false triggers caused by scanning and boundary jitter. In typical scenarios such as lane change verification, instrument panel viewing while driving, and a short-term central control operation, it maintains gating behavior consistent with the semantics of the driving task, making the judgment basis clear and callable even in cases of unintentional modulation.
[0077] Furthermore, this invention also implements a closed-loop method of perception-judgment-interpretation-linkage. Under dynamic budgeting, dual-channel accumulation is performed, outputting the first time the limit is exceeded and the risk level, and recording the identifier of the attention zone, the intent label, and the type of the budget item exceeded. Subsequently, based on the triggering reason, linkage strategies such as low-interference prompts, operation restrictions, or strong prompts are executed in the cockpit. The overall technical effect of this method is to integrate intervention decisions with scenario semantics and compliance records, clarifying why and when the system intervenes, what kind of legal task's tolerance has been exhausted, or what type of non-driving occupation is intercepted, facilitating engineering integration and accountability auditing, and maintaining necessary tolerance for legal short-term behavior while ensuring timely interception of non-driving occupation.
[0078] Furthermore, based on the above embodiments, this invention also provides a driving behavior recognition device, such as... Figure 4 The diagram shows a structural schematic of a driving behavior recognition device, which includes: The response module 40 is used to generate an attention zone sequence in response to a change in the attention zone corresponding to the driver's line of sight; wherein, the attention zone is a cockpit partition based on the driver's perspective; the attention zone sequence is an identifier sequence of multiple target attention zones recorded in chronological order, and the target attention zone is the attention zone corresponding to the change in the line of sight within a preset continuous time period. The recognition module 42 is used to recognize the driver's intention based on the attention area sequence, and obtain the intention label corresponding to the intention and the reasonable duration expectation of the intention; Adjustment module 44 is used to dynamically adjust the attention budget curve of each pre-configured target attention area according to the intent label and the expected reasonable duration, so as to obtain the dynamic budget curve corresponding to each target attention area; The determination module 46 is used to determine the behavior of the target attention area based on the dynamic budget curve, so as to determine whether the behavior of the target attention area meets the preset intervention conditions. Intervention module 48 is used to generate an intervention strategy that includes risk level and intervention timing when the determination result of the determination module is yes, so as to provide prompts to the driver's driving behavior based on the intervention strategy.
[0079] The driving behavior recognition device provided in this embodiment of the invention has the same technical features as the driving behavior recognition method provided in the above embodiments, so it can also solve the same technical problems and achieve the same technical effects.
[0080] Furthermore, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above method.
[0081] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method.
[0082] Furthermore, embodiments of the present invention also provide a schematic diagram of the structure of an electronic device, such as... Figure 5 The diagram shows the structure of the electronic device, which includes a processor 51 and a memory 50. The memory 50 stores computer-executable instructions that can be executed by the processor 51, and the processor 51 executes the computer-executable instructions to implement the above-described method.
[0083] exist Figure 5In the illustrated embodiment, the electronic device further includes a bus 52 and a communication interface 53, wherein the processor 51, the communication interface 53, and the memory 50 are connected via the bus 52.
[0084] The memory 50 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 53 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 52 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 52 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0085] Processor 51 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 51 or by instructions in software form. Processor 51 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor 51 reads the information in the memory and uses its hardware to complete the aforementioned method.
[0086] Furthermore, the computer program product of the driving behavior recognition method and apparatus provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0087] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0088] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0089] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0090] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0091] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for recognizing driving behavior, characterized in that, The method includes: In response to a change in the attention zone corresponding to the driver's line of sight, an attention zone sequence is generated; wherein, the attention zone is a cockpit partition based on the driver's perspective; the attention zone sequence is an identifier sequence of multiple target attention zones recorded in chronological order, and the target attention zone is the attention zone corresponding to the change in the line of sight within a preset continuous time period; Based on the attention zone sequence, the driver's intention is identified, and the intention label corresponding to the intention and the expected reasonable duration of the intention are obtained. Based on the intent tag and the expected reasonable duration, the attention budget curve of each pre-configured target attention area is dynamically adjusted to obtain the dynamic budget curve corresponding to each target attention area. The behavior of the target attention area is determined based on the dynamic budget curve to determine whether the behavior of the target attention area meets the preset intervention conditions. If so, generate an intervention strategy that includes risk level and intervention timing, and provide prompts to the driver's driving behavior based on the intervention strategy.
2. The method according to claim 1, characterized in that, The steps for generating a sequence of attention zones in response to a change in the driver's line of sight include: Based on the point of view, determine the number of attention zones where the spatial boundary intersects with the direction of view; If there is only one attention region, then that attention region is the target attention region, and the identifier of the target attention region is output. If there are multiple attention zones, a target attention zone is determined from the multiple attention zones according to the preset attention zone priority, the hand-accessible objects contained in the attention zone, and the identifier of the target attention zone; The attention region sequence is generated based on the identifiers of the target attention region within the preset continuous time period.
3. The method according to claim 1, characterized in that, The steps of identifying the driver's intent based on the attention region sequence, obtaining the intent label corresponding to the intent, and the expected reasonable duration of the intent include: Obtain a pre-constructed semantic map of attention regions, which records the identifier of each attention region, the target object identifier of the hand-reachable object contained in each attention region, and the mapping relationship of the spatial boundary of each attention region; Based on the target object identifier and driving task semantics corresponding to each attention area recorded by the attention area semantic mapping, a candidate intent corresponding to each target attention area is obtained; Extract the timing information of the gaze point's dwell time in the target attention area from the attention area sequence, and generate an intent label representing the intent and a corresponding reasonable duration expectation based on the dwell time information and the candidate intent.
4. The method according to claim 1, characterized in that, The step of dynamically adjusting the attention budget curve for each pre-configured target attention zone based on the intent label and the expected reasonable duration includes: The adjustment direction of the attention zone budget curve for each target attention zone is determined based on the intent label and the expected reasonable duration. The attention budget curve for each target attention zone is dynamically adjusted based on the adjustment direction.
5. The method according to claim 4, characterized in that, The step of determining the adjustment direction of the attention zone budget curve for each target attention zone based on the intent label and the expected reasonable duration includes: If the intent label is the intent label corresponding to the driving task, and the reasonable duration expectation is a short-term expectation requirement, then the adjustment direction of the attention area budget curve is the widening direction. If the intent label is an intent label corresponding to a non-driving task, and the reasonable duration is expected to be occupied by a non-driving task, then the adjustment direction of the attention area budget curve is the tightening direction.
6. The method according to claim 5, characterized in that, The step of dynamically adjusting the attention budget curve of each target attention area based on the adjustment direction includes: Obtain the driver's hand occupancy status; The adjustment force of the attention zone budget curve in the adjustment direction is determined based on the hand occupancy status. The attention budget curve of each target attention zone is dynamically adjusted based on the adjustment direction and the adjustment intensity.
7. The method according to claim 6, characterized in that, The step of determining the adjustment force of the attention zone budget curve in the adjustment direction based on the hand occupancy state includes: When the hand occupancy state is single-hand occupancy driving state, and the intent label is the intent label corresponding to a non-driving task, the tightening amplitude is increased in the tightening direction according to the preset tightening strategy. When the hand occupancy state is "both hands are holding" while driving, and the intent label is the intent label corresponding to the driving task, the widening range is reduced in the widening direction according to a pre-set widening strategy.
8. The method according to claim 1, characterized in that, The dynamic budget curve is used to characterize the benchmark of the continuous dwell time of the gaze point in the target attention area, and the benchmark of the number of times the gaze point repeatedly enters the target attention area within a preset time range; The step of determining whether the behavior of the target attention area meets the preset intervention conditions based on the dynamic budget curve includes: Extract the behavioral sequence within the preset continuous time period corresponding to the attention area sequence; Determine whether the cumulative continuous dwell time in the behavior sequence reaches the continuous dwell time benchmark represented by the dynamic budget curve, or whether the cumulative number of repetitions in the behavior sequence reaches the number of repetitions benchmark represented by the dynamic budget curve; If any of the judgment results is yes, then it is determined that the behavior of the target attention area has reached the preset intervention condition.
9. The method according to claim 8, characterized in that, The steps for generating an intervention strategy that includes risk level and timing of intervention include: The intervention timing is defined as the moment when the continuous dwell time benchmark or the repetition count benchmark is first reached in the behavioral sequence; and... Based on the task attributes of the target attention area, the intervention rules for the target attention area are determined, and the risk level is determined according to the intervention rules and the continuous dwell time benchmark reached by the behavior sequence, and / or the repetition frequency benchmark. The intervention strategy is determined based on the timing of intervention and the risk level.
10. A driving behavior recognition device, characterized in that, The device includes: A response module is used to generate a sequence of attention zones in response to a change in the attention zone corresponding to the driver's line of sight; wherein, the attention zone is a cockpit partition based on the driver's perspective; the attention zone sequence is an identifier sequence of multiple target attention zones recorded in chronological order, and the target attention zone is the attention zone corresponding to the change in the line of sight within a preset continuous time period; The recognition module is used to recognize the driver's intention based on the attention area sequence, and obtain the intention label corresponding to the intention and the reasonable duration expectation of the intention; The adjustment module is used to dynamically adjust the attention budget curve of each pre-configured target attention area according to the intent tag and the expected reasonable duration, so as to obtain the dynamic budget curve corresponding to each target attention area. The determination module is used to determine the behavior of the target attention area based on the dynamic budget curve, so as to determine whether the behavior of the target attention area meets the preset intervention conditions; An intervention module is used to generate an intervention strategy that includes a risk level and an intervention timing when the determination result of the determination module is yes, so as to provide prompts to the driver's driving behavior based on the intervention strategy.