A driver abnormality monitoring system based on large model collaborative decision making

Through multimodal data enhancement and collaborative decision-making mechanism, combined with visual pre-alignment and lightweight language reasoning modules, deep scene understanding of the driver abnormality monitoring system is achieved, solving the problems of complex situational understanding and weak response mechanism of the existing system, and improving the safety of assisted driving and data privacy protection.

CN120635869BActive Publication Date: 2025-10-17NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511129602.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-17
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing driver monitoring systems lack the ability to understand complex situations, have different optimization goals and evaluation criteria for perception and reasoning tasks, are unable to effectively interact with vehicle-mounted systems, and have single recognition results and weak response mechanisms.

Method used

It adopts multi-channel cameras, voice interaction module, DTOPA visual pre-alignment module, Qwen-Apan lightweight language reasoning module and collaborative decision-making module, and realizes driver abnormality monitoring, including visual pre-alignment, semantic analysis, risk assessment and safety response, through multimodal data enhancement, lightweight model and collaborative decision-making mechanism.

Benefits of technology

It has achieved a leap from traditional feature monitoring to deep scene understanding, accurately identifying complex driving behaviors, building a lightweight large model system, reducing delayed responses, ensuring user data privacy and security, and performing safety interventions through intelligent closed loops to improve the safety of assisted driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635869B_ABST
    Figure CN120635869B_ABST
Patent Text Reader

Abstract

The application discloses a driver abnormality monitoring system based on large model cooperative decision, comprising: a multi-path camera, a voice interaction module, a DTOPA visual pre-alignment module, a Qwen-Apan lightweight language reasoning module, a cooperative decision module and a safety response control module; wherein the multi-path camera is used for real-time extraction of in-vehicle environment video stream; the voice interaction module is used for reminding the driver and interacting with the driver; the DTOPA visual pre-alignment module is used for converting the video stream into a structured text description; the Qwen-Apan lightweight language reasoning module is used for semantic analysis and abnormality risk assessment of the text description; the cooperative decision module is used for triggering voice interaction and vehicle control instructions based on the risk assessment result; and the safety response control module is used for executing side parking, emergency contact person notification or multi-level early warning operation; and the application significantly improves the safety of auxiliary driving.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of auxiliary driving safety processing, and particularly relates to a driver abnormality monitoring system based on large model collaborative decision-making. BACKGROUND

[0002] Traditional driver monitoring systems mostly use a simple physiological index-based judgment method for fatigue detection, lacking the ability to understand complex situations, especially auxiliary driving situations. Current mainstream large language models tend to focus on perception or reasoning, with perception tasks focusing on accurate visual feature extraction and scene recognition, and reasoning tasks requiring logical analysis and decision-making based on understanding. The optimization objectives and evaluation standards of the two are fundamentally different, and neither has true intelligent analysis capabilities. Existing DMS systems have relatively single recognition results and can only provide simple beeping or vibration warnings. At the same time, it lacks a reliable coordination method with the current latest generation of electronic car machine systems, so it cannot effectively interact with the car machine system and function linkage. SUMMARY

[0003] The purpose of the present application is to provide a driver abnormality monitoring system based on large model collaborative decision-making, which solves the problem of single perception dimension, lack of scene understanding and reasoning ability, and weak response mechanism of existing driver monitoring systems.

[0004] Technical solution: The driver abnormality monitoring system based on large model collaborative decision-making comprises a multi-camera, a voice interaction module, a DTOPA visual pre-alignment module, a Qwen-Apan lightweight language reasoning module, a collaborative decision-making module, and a safety response control module. The multi-camera is used to extract real-time in-vehicle environment video streams. The voice interaction module is used to remind and interact with the driver, and finally feeds back the results to the entire system as part of the decision. The DTOPA visual pre-alignment module is used to convert video streams into structured text descriptions. The Qwen-Apan lightweight language reasoning module is used for semantic analysis and abnormal risk assessment of text descriptions. The collaborative decision-making module is used to trigger voice interaction and vehicle control instructions based on risk assessment results. The safety response control module is used to perform side parking, emergency contact notification, or multi-level warning operations.

[0005] Further, the DTOPA visual pre-alignment module specifically comprises: adopting a multi-modal data enhancement strategy to generate a driving text video (DTV) dataset to simulate real video timing characteristics; using a CLIP-ViT-L model to align image and text features; using a multi-scale convolution embedding and modal adaptive normalization to eliminate CLIP dependence through visual ability internalization technology.

[0006] Further, the multi-modal data augmentation strategy is as follows: LLM is used to automatically generate multi-modal pre-training data, and a large-scale driving text-video DTV dataset containing text, video and associated annotations is created.

[0007] Further, the Qwen-Apan lightweight language reasoning module is as follows: the GPTQ-Int8 quantization scheme is used to compress the model weight; and the KV cache is dynamically allocated based on attention entropy.

[0008] Further, the collaborative decision module is as follows: a driver state understanding index DSU is defined; a three-level response is triggered according to the abnormal level: voice reminder, emergency contact notification, and auxiliary parking; wherein the DSU formula is as follows:

[0009] ;

[0010] wherein w risk is a risk weighting function; groundtruth i is the ground truth of the ith sample, containing in-vehicle scene, driver state, risk level information and other facts, N is the total number of test samples (positive integer), α is the weight coefficient of description accuracy, controlling the importance of description accuracy in the total evaluation, β is the weight coefficient of state classification accuracy, controlling the importance of classification accuracy in the total evaluation, i is the sample index from 1 to N; represents the driver state description accuracy, represents the driver state classification accuracy, and the calculation formulas are as follows:

[0011] ;

[0012] ;

[0013] wherein, represents the ith driver state description generated by DTOPA, represents the standard description annotated by the dataset; ROUGE-L represents the longest common subsequence ratio, represents whether the corresponding label is mentioned in the jth driver state description; represents the true driver state label, 、 both represent the total number of test samples; j is the classification sample index from 1 to M.

[0014] Further, by simulating the diversified behavior patterns of the Liangfeng Angle Toad, the AdamH optimizer is designed to realize parameter optimization of the multi-strategy collaborative module.

[0015] Further, the CH chaotic mapping initialization parameter is introduced, and the gradient direction is updated according to the behavior mode of the Feihuang chenshania, i.e., foraging, defense, escape and social interaction.

[0016] ;

[0017] In the formula, alpha is greater than or equal to 0, beta is greater than or equal to 0, and alpha plus beta is less than or equal to 1, u is in [0, 4], γ And beta is in [2.5, 3.0], both are mapping control parameters, z i Is a current chaotic state value, z i+1 Is a next state value output by the CH chaotic mapping, and satisfies z i+1 Is in [0, 1]; w i Is an intermediate weight value for the i-th iteration for generating a chaotic sequence; lambda is an interpolation parameter, controlling the balance between deterministic components and random components; r is a random factor, introducing random disturbance to enhance ergodicity; pi is a circular constant, approximately equal to 3.14159; sin is a sine function, cos is a cosine function, and || is an absolute value function, ensuring that the output is non-negative.

[0018] The driver abnormality monitoring method disclosed in the application is realized by a driver abnormality monitoring system based on large model collaborative decision, and comprises the following steps:

[0019] (1) acquiring in-vehicle video streams at 25 FPS through multiple cameras;

[0020] (2) generating scene text descriptions in real time by using a DTOPA model;

[0021] (3) analyzing the text and outputting risk assessment by using a Qwen-Apan model;

[0022] (4) if an abnormality is detected, starting voice interaction to confirm the state;

[0023] (5) executing a safety response operation according to the interaction result.

[0024] Further, in step (1), a multi-modal data enhancement strategy is adopted to generate a driving text video DTV data set.

[0025] Further, in step (2), a 21-second sliding window mechanism is adopted, 20 frames are randomly extracted to form an analysis batch, and compressed frame data is transmitted through WebSocket.

[0026] Beneficial effects: Compared with the prior art, the present application has the following remarkable advantages: (1) realizing the technical leap from traditional "feature monitoring" to deep "scene understanding", accurately identifying complex driving behaviors such as fatigue, distraction, abnormal posture, etc. (2) Building a set of lightweight large model system that can run efficiently on vehicle-mounted edge computing devices, getting rid of cloud dependence, ensuring low-latency response and user data privacy and security. (3) Establishing an intelligent closed loop of "perception-understanding decision-making-interaction-execution", when risks are detected, it can confirm the state through voice interaction, and through the linkage of vehicle auxiliary systems, it can safely intervene by methods such as pulling over and parking, significantly improving the safety of assisted driving. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 is a flowchart of the present application;

[0028] Figure 2 is a schematic diagram of the visual pre-alignment module of the present application. DETAILED DESCRIPTION

[0029] The technical solutions of the present application will be further described below in combination with the drawings.

[0030] As Figure 1 shown, the embodiment of the present application provides a driver abnormality monitoring system based on large model collaborative decision-making, multiple cameras continuously collect in-vehicle video streams, transmit the DTOPA visual pre-alignment module, which is responsible for "translating" video frames into structured text descriptions, then the Qwen-Apan lightweight language reasoning module performs semantic analysis on the text to determine whether there is an abnormality. Once an abnormality is detected, the system will trigger the man-machine interaction module to communicate with the driver through voice, and according to the response result, the safety response control module decides whether to take measures such as alarm, notification of emergency contact person or auxiliary vehicle pulling over, etc.; including multiple cameras, voice interaction module, DTOPA visual pre-alignment module, Qwen-Apan lightweight language reasoning module, collaborative decision-making module, safety response control module; wherein the multiple cameras are used to extract in-vehicle environment video streams in real time; the voice interaction module is used to remind the driver and interact with the driver; the DTOPA visual pre-alignment module is used to convert video streams into structured text descriptions; the Qwen-Apan lightweight language reasoning module is used for semantic analysis and abnormal risk assessment of the text description; the collaborative decision-making module is used to trigger voice interaction and vehicle control instructions based on the risk assessment result; the safety response control module is used to perform the operations of pulling over, notifying emergency contact person or multi-level warning.

[0031] Among them, the DTOPA visual pre-alignment module is as follows:

[0032] Firstly, a text pre-alignment framework is adopted to simulate real video-text pairs by serializing text descriptions and their associated annotations through a multi-modal data augmentation strategy, which preliminarily aligns the concept level of the pure text LLM with the video modality. To bridge the gap between the synthesized text-video and real videos, the CLIP model is used as a feature extractor to effectively fuse image and text features during preliminary pre-training, thereby aligning the video image and text modalities. Then, through the internalization of visual ability, the CLIP model is removed, which greatly reduces the requirement for the performance of the end-side platform during deployment while retaining the alignment function of the video image and text modalities. At the same time, three new pre-alignment tasks are trained through a unified autoregressive language modeling objective, achieving the pre-alignment of text-video and LLM. The introduction of the support memory mechanism can bridge the gap between video images and text modalities through a training-free projection process.

[0033] As shown in Figure 2 , the multi-modal data augmentation strategy: using LLM to automatically generate multi-modal pre-training data, creating a large-scale driver-oriented text-video (DTV) dataset containing text-video and its associated annotations, thereby solving the problem of data scarcity in visual understanding tasks. Specifically, text videos composed of consecutive text frames are automatically generated by LLM, each containing 100-200 sequential text frames, each frame containing frame descriptions and multiple object descriptions, thereby achieving the time sequence characteristics of real videos and ensuring that the generated text videos can capture dynamic changes and event development processes. To ensure the diversity and domain coverage of the DTV dataset, different conditional prompts are used as generation guides, including video titles and video descriptions of driver anomaly datasets such as driver anomaly detection and driver behaviors, ensuring the generality and practicality of the generated data through a multi-source conditional prompt strategy.

[0034] CLIP model: to bridge the gap between text-video and real videos, CLIP-ViT-L is used as a feature extractor to align image and text modalities during preliminary pre-training.

[0035] For text-video, given the text frame , the CLIP text encoder extracts frame representations from the frame description and detailed object description , as follows:

[0036] ; (1)

[0037] wherein, denotes a SWAF (Scene-Aware Weighted Fusion, SAWF) function, denotes a CLIP text encoder. Thus, a text video containing n text frames is represented as ; where w c and w d are preset global-local information balance hyperparameters responsible for dynamic attention weights.

[0038] For real videos, given an image frame , the CLIP image encoder extracts frame-level visual features , which is formulated as follows:

[0039] ; (2)

[0040] where is projected into the LLM space via a linear layer, denoted as , where denotes a language model, denotes a projection layer that projects CLIP features into the LLM space, PreProcess is an image preprocessing function, S env represents the environmental state (such as whether the image needs to be denoised, whether the light is balanced, etc.);

[0041] Pre-alignment tasks: Three tasks are introduced in the DTOPA model to achieve the pre-alignment of text videos and LLMs. The text video summary task generates a detailed description to summarize the content of the entire text video according to the given text video, the text video question answering task predicts the answer according to the given text video and question, and the multiple-choice text video question answering task is responsible for selecting the correct answer from the candidate answers according to the given text video, question, and multiple options. Autoregressive language modeling objectives are used to train these three tasks in the model, as shown below:

[0042] ; (3)

[0043] where is the total loss function of the DTOPA model, N is the sequence length, λ is a regularization hyperparameter used to balance the loss, denotes the i-th target label, log is the logarithmic function, p represents the probability that the large model can accurately predict the next correct word unit at each step when generating the target text according to the input, t <i denotes all t before the i-th; R is a task-oriented regularization term (Task-Oriented Regularization). When the generated text V' lacks a description of the key behavior, the value of this term will increase.

[0044] As the CLIP model pre-training achieves image-text feature alignment, this alignment representation enables zero-shot reasoning without additional fine-tuning. However, the inter-modal gap phenomenon (i.e., CLIP image features and CLIP text features are located in completely different regions of the feature space) prevents direct mapping of visual features as text features using. To bridge this inter-modal gap, the DTOPA model employs a support memory mechanism to project CLIP visual features into the CLIP text feature space. This training-free projection process is represented as:

[0045] ; (4)

[0046] where, is the final feature projected from visual features to text space, represents the CLIP text features in the pre-constructed memory of size N, sim represents the similarity computation function between vectors, and τ is the temperature coefficient used to control the smoothness of the attention distribution. is the original visual feature. In zero-shot reasoning, the is used as the representation of the real video.

[0047] Using the visual capability internalization technique, the functionality of the CLIP module itself is internalized into DTOPA, so that the CLIP module can be removed. The core is to integrate the visual perception capability directly into the language model internally, while maintaining the video understanding capability learned in the first stage. This stage adopts low-rank adapters by introducing them into the first few layers of the LLM to encode visual information. Specifically, for the first Nvit layers of the LLM, DTOPA will apply:

[0048] ; (5)

[0049] where W represents the pre-trained weight matrix, ΔW = BA represents the low-rank adaptive update, B ∈ R d×r and A ∈ R r×d are low-rank decomposition matrices. h is the output of the final linear layer, x is the input, d is the original dimension, and r is the low-rank dimension. After training is complete, the parameters can be seamlessly merged into the base LLM after training of 1) to 4), so that zero additional computational overhead is achieved during inference.

[0050] ; (6)

[0051] where, merged W is the merged weight matrix for zero additional computational overhead during inference.

[0052] To replace the external CLIP visual encoder, DTOPA designs a direct visual embedding module, which first divides the video frame into patches through convolutional layers, and proposes a "multi-scale convolutional embedding and modal adaptive normalization" scheme. Convolution is no longer single-scale, but multi-scale convolution in parallel to capture visual features of different granularities:

[0053] ; (7)

[0054] where I represents the input video frame image, and p represents the patch size. Position encoding is then added and projected into the LLM space. patches is the patch feature after multi-scale convolution, and Concat is the feature concatenation operation; the feature map after convolution operation is flattened and transposed to obtain a sequence representation, and then RMS normalization is applied to stabilize the training process, where x norm is the normalized sequence feature, and RMSNorm represents RMS normalization, x seq is the sequence feature after sequence, that is, the feature representation of the one-dimensional sequence obtained by flattening and expanding the two-dimensional image patch:

[0055] ; (8)

[0056] where the calculation formula of RMS normalization is:

[0057] ; (9)

[0058] where d is the feature dimension, x i is the i-th feature, ε is a numerical stability constant (which can be set to 1e-8), and γ is a learnable scaling parameter.

[0059] The final visual embedding is obtained by linear projection, and this linear layer maps the intermediate-dimensional feature to the hidden dimension of DTOPA, ensuring that the visual representation can be processed in the same space as the text representation. This design allows visual information to be directly processed in a form that DTOPA can understand, without the need for an external intermediary. Compared to other visual models such as LLAVA, which have ViT + high training cost + traditional architecture, this design allows the added parameters to be seamlessly integrated into the LLM during inference, eliminating structural complexity and minimizing training and computing overhead, with inference speed per token increased by more than 1.6 times, and memory usage reduced by 15%, making DTOPA more suitable for deployment on devices with limited computing power.

[0060] The Qwen-Apan lightweight language reasoning module is based on a Qwen3-1.7B model, adopts a GPTQ-Int8 training post-quantization scheme, compresses model weights into 8-bit integers, greatly reduces GPU / memory occupancy, enables stable deployment on low-power platforms such as Nvidia Jetson Orin NX, shortens reasoning time, and is more suitable for driver abnormal state detection tasks.

[0061] Although the Qwen3-1.7B model is lightweight compared to larger models, when it is deployed on embedded devices, there are still obstacles such as high reasoning delay, long startup time, and easy triggering of memory overflow, so other technologies need to be introduced to compress the model size, reduce computing requirements, and improve reasoning efficiency.

[0062] In addition, to further optimize the memory usage efficiency of the KV cache, the application introduces a KV cache allocation strategy based on the attention entropy between the decision reasoning text and the video description generated by DTOPA. Traditional KV cache management methods use a uniform cache size allocation for all Transformer layers, ignoring the differences in attention density when different layers process different information. In the driver abnormal state monitoring application, there are significant differences in the attention levels of different attention layers to different information.

[0063] The KV cache allocation strategy quantifies the attention distribution characteristics of each layer through attention entropy, thereby using a dynamic method to determine the optimal KV cache allocation scheme. Specifically, the attention entropy calculation formula is as follows:

[0064] ; (10)

[0065] wherein, is the cross-modal attention entropy of the lth layer; represents the attention probability distribution of m1 to m2 in the i th attention head of the l th layer, m1 and m2 represent the two different modalities at present; Nhead is the number of attention heads. Based on the attention entropy, the cache allocation strategy is:

[0066] ; (11)

[0067] wherein, K l is the KV cache size of the lth layer, K base is the basic cache size, and are the mean and standard deviation of the attention entropy of all layers, respectively, η is a regulation factor, Cm(l) is a contribution factor, representing the contribution of modality m to the lth layer decision; M represents the set of all modalities.

[0068] GPTQ is a gradient-based post-training quantization technique that can significantly compress model size and speed up inference while minimizing accuracy loss. For each layer, the algorithm uses second-order information (Hessian matrix) to guide the quantization process, ensuring that the quantization error is minimized.

[0069] Here, a solution algorithm variant DSQ that preserves key weights is used. During quantization, not all weights are treated equally. Instead, a small subset of weights that are critical to "dangerous driving behavior judgment" is identified and quantized with higher precision. Specifically, the following optimization problem is solved:

[0070] ; (12)

[0071] where, denotes the original weight, denotes the quantized weight, denotes the input activation of the layer, Msaliency is a "weight saliency mask" with the same dimension as W, where the weights critical to safety judgment correspond to a value of 1, and the rest are 0. This mask can be obtained by analyzing the gradient or feature contribution, and λ is a hyperparameter that balances the overall quantization error and the protection of key weights; is the L2 norm, and argmin denotes the quantized weight that minimizes the expression.

[0072] Although the GPTQ technique is widely used in the deployment of LLMs due to its excellent compression capability and fast inference advantage, it is still difficult to meet the running and storage requirements in resource-constrained edge devices or vehicle platforms, and still relies on partial high-performance GPUs or hardware platforms with floating-point calculation capability in inference, limiting its practical value in lightweight terminals. Therefore, the present invention introduces an Int8 integer quantization strategy within the GPTQ framework, forming the GPTQ-Int8 scheme with distance compression and high efficiency. By compressing the weights to 8-bit integer representation and preserving the high-precision expression of key structures, the model size is significantly reduced while maintaining the stability of inference accuracy. Compared with the traditional FP16-GPTQ model, the Int8 version of the model size can be compressed by more than 50%, significantly reducing the memory, video memory and bandwidth requirements, making it more suitable for low-power environments. The memory access and cache hit rate of GPTQ-Int8 are better, and the overall execution delay is significantly reduced, achieving the best balance between model compression, inference acceleration and deployment cost while maintaining language understanding and semantic reasoning capabilities.

[0073] Co-decision module: As the core technology in the current artificial intelligence field, the large language model (LLM) can serve as a cross-modal semantic bridge to convert visual features into structured text semantics, effectively solving the gap problem between visual signals and semantic reasoning in video understanding. However, the current video-text dataset based on network collection has low language labeling efficiency, and only simple label annotation, which is difficult to use for training. Therefore, the present application proposes a co-decision mechanism algorithm of fusion of visual driving pre-alignment and lightweight language reasoning double model, realizing the technical leap from 'feature monitoring' to'scene understanding'.

[0074] Driver state understanding evaluation index DSU: The system proposes a new evaluation index DSU, and the calculation formula is as follows:

[0075] ; (13)

[0076] Among them, w risk is a risk weighting function that maps groundtruth to a risk weight, which makes DSU not only evaluate accuracy, but also evaluate risk perception accuracy. DSU represents the driver state description accuracy, DSU represents the driver state classification accuracy, and the calculation formulas are as follows:

[0077] ; (14)

[0078] ; (15)

[0079] Among them, represents the i-th driver state description generated by DTOPA, represents the standard description annotated by the data set, ROUGE-L represents the longest common subsequence ratio, represents whether the corresponding label is mentioned in the j-th driver state description, represents the true driver state label, If is an indicator function that returns 1 when the condition is true and 0 when the condition is false, M and N both represent the total number of test samples, and j is the classification sample index from 1 to M.

[0080] The traditional Adam and AdamW optimizers adopt a monotonically decreasing learning rate adjustment strategy. Although this design helps to ensure convergence stability, it is easy to cause the algorithm to fall into a local optimal solution too early in a complex loss terrain; in order to solve these fundamental defects of traditional optimizers, by simulating the diversified behavior patterns of the lotus peak horned toad, the present application proposes an AdamH optimizer, which realizes a multi-strategy collaborative parameter optimization mechanism.

[0081] In the training of large models, the parameter space of the neural network needs to be mapped to the position space of the Leucostigma ilishanense first. Each individual represents a complete set of network parameter configurations, and each parameter is initialized by cubic chaotic mapping. .

[0082] It should be noted that the conventional methods currently used in the initialization stage of training large models are often zero initialization, uniform distribution initialization, normal distribution initialization, etc.

[0083] At present, in order to solve these problems, chaotic mapping is often used to enhance initialization, but traditional single chaotic mapping (such as Logistic or Sine) may also lead to chaotic degeneration, for example, the output of Logistic mapping is easy to fall into periodic orbit when μ=4, and the point set is densely distributed near the diagonal line of the phase space; Sine mapping produces high-density aggregation in the boundary region, significantly reducing ergodicity.

[0084] To improve the performance of AdamH algorithm in the initialization stage, the present application introduces a CH chaotic mapping to perform initialization, and its mathematical expression is:

[0085] ; (16)

[0086] In the formula, α≥0, β≥0 and α+β≤1, u∈[0,4], γ ∈[2.5,3.0], these are mapping control parameters, zi is the current chaotic state value, z i+1 is the next state value output by the CH chaotic mapping, and satisfies z i+1 ∈[0,1]; w i is the intermediate weight value of the i-th iteration for generating chaotic sequence; λ is an interpolation parameter, which controls the balance between deterministic component and random component; r is a random factor, which introduces random disturbance to enhance ergodicity; π is the circular constant, which is approximately equal to 3.14159; sin is the sine function, cos is the cosine function, and || is the absolute value function, which ensures that the output is non-negative.

[0087] The CH chaotic mapping is integrated through a specific two-layer processing architecture and other structures, aiming to generate complex and uniformly distributed chaotic sequences. In phase space behavior analysis, the point sets generated by logistic mapping and sine mapping tend to concentrate in a limited area, showing obvious chaotic degradation phenomenon. In contrast, the CH chaotic mapping exhibits highly uniform phase space distribution characteristics, which significantly improves the randomness and robustness of its output sequence.

[0088] After initialization is completed, the population is sorted and ranked based on the fitness value, and the behavior probability is assigned to each individual:

[0089] ; (17)

[0090] ; (18)

[0091] In the above formula, sorted_indices is the index of the individual sorted by fitness value, f i is the fitness value (loss function value) of the ith individual, N is the number of population individuals, argsort is a function that returns the sorted index, P foraging,i is the foraging behavior probability of the ith individual, P defense,i is the defense behavior probability of the ith individual, P escape,i is the escape behavior probability of the ith individual, P social,i is the social behavior probability of the ith individual, rank i represents the rank, e is a natural constant, approximately equal to 2.71828.

[0092] In order to make full use of gradient information, AdamH integrates gradient direction in the update process, calculates the gradient of the current individual, and integrates gradient information into behavior:

[0093] ; (19)

[0094] In the formula, X i gradient represents the individual position after integrating gradient information, X i current represents the current individual position, as a learning rate decaying over time t, g(·,·) is an adaptive function that adjusts the step size according to the gradient and Hessian matrix information, ∇f(θ i ) is the gradient of the fitness function f with respect to the parameter θ i , θ i is the parameter vector of the ith individual, H i is the Hessian matrix of the fitness function f at θ i .

[0095] According to the above formula, the adaptive learning rate parameter of the i-th individual not only decays over time, but also adapts according to the local gradient and the curvature (Hessian matrix Hi) which can be regarded as the "terrain":

[0096] ; (20)

[0097] Inheriting the momentum mechanism of the Adam optimizer, AdamH maintains first and second order momentum estimates, continuously preserving and updating these momentum information:

[0098] ; (21)

[0099] ; (22)

[0100] where β1, β2 are momentum decay parameters; is the first moment estimate (momentum) at time t, is the second moment estimate at time t, is the updated individual position.

[0101] The final position update formula combining all behavior patterns is:

[0102] ; (23)

[0103] ; (24)

[0104] The coefficients in the formula are dynamically adjusted according to the alignment degree of the first order momentum and the second order momentum after bias correction. When the gradient direction is stable (high alignment degree), α H,t becomes small, trusting the gradient; when the gradient oscillates (low alignment degree), α H,t becomes large, reducing the dependence on the gradient. is the final updated individual position, is the base value of the dynamic coefficient, is the numerical stability coefficient, is the behavior weight vector, exp(·) is the exponential function, E is a small constant for numerical stability, and softmax(...) is the Softmax function that converts a numerical vector into a probability distribution with each item summing to 1.

[0105] The convergence of the algorithm is monitored by the following indicators:

[0106] ; (25)

[0107] where Convergence is the convergence criterion, X best is the current best individual position, is the convergence threshold, below which the algorithm is considered to have converged.

[0108] Repeat the above steps until a termination condition (maximum number of iterations or convergence) is met, and finally output the trained parameters.

Claims

1. A driver abnormality monitoring system based on large model collaborative decision-making, characterized by: include: Multi-channel cameras, voice interaction module, DTOPA visual pre-alignment module, Qwen-Apan lightweight language reasoning module, collaborative decision-making module, safety response control module; among them, the multi-channel cameras are used to extract the in-vehicle environment video stream in real time; the voice interaction module is used to remind the driver and interact with the driver, and finally feed back the results to the entire system as part of the decision-making; the DTOPA visual pre-alignment module is used to convert the video stream into a structured text description; the Qwen-Apan lightweight language reasoning module is used to perform semantic analysis and abnormal risk assessment on the text description; the collaborative decision-making module is used to trigger voice interaction and vehicle control instructions based on the risk assessment results; the safety response control module is used to execute pull-over parking, emergency contact notification or multi-level warning operations; among them, the collaborative decision-making module is as follows: define the driver status understanding index DSU; trigger three levels of response according to the abnormality level: voice reminder, emergency contact notification, and assisted pull-over parking; among them, the DSU formula is as follows: ; Among them, w risk is the risk-weighted function; Indicates the accuracy of driver status description, groundtruth i is the basic facts of the i-th sample, including the in-vehicle scene, driver status, and risk level information facts. N is the total number of test samples, which is a positive integer. α is the description accuracy weight coefficient, which controls the importance of description accuracy in the overall evaluation. β is the state classification accuracy weight coefficient, which controls the importance of classification accuracy in the overall evaluation. i is the sample index from 1 to N. It represents the classification accuracy of driver status, and its calculation formula is as follows: ; ; in, represents the i-th driver state description generated by DTOPA, represents the standard description of the dataset annotation; ROUGE-L represents the longest common subsequence ratio, Indicates whether the corresponding label is mentioned in the j-th driver state description; represents the true driver state label, 、 Both represent the total number of test samples; j is the classification sample index from 1 to M.

2. The driver abnormality monitoring system based on large model collaborative decision-making according to claim 1 is characterized in that: The DTOPA visual pre-alignment module is specifically designed as follows: a multimodal data augmentation strategy is used to generate a driving text video (DTV) dataset that simulates the temporal characteristics of real videos; the CLIP-ViT-L model is used to align image and text features; The CLIP dependency is eliminated through visual ability internalization technology, and multi-scale convolutional embedding and modality adaptive normalization are adopted.

3. The driver abnormality monitoring system based on large model collaborative decision-making according to claim 2 is characterized in that: The specific multimodal data enhancement strategy is as follows: LLM is used to automatically generate multimodal pre-training data and create a large-scale driving text video DTV dataset containing text videos and their associated annotations.

4. The driver abnormality monitoring system based on large model collaborative decision-making according to claim 1 is characterized in that: The Qwen-Apan lightweight language inference module is specifically as follows: the GPTQ-Int8 quantization scheme is used to compress model weights; the KV cache is dynamically allocated based on attention entropy.

5. The driver abnormality monitoring system based on large model collaborative decision-making according to claim 4 is characterized in that: By simulating the diverse behavioral patterns of the Lianfeng Horned Toad, the AdamH optimizer is designed to achieve parameter optimization of the multi-strategy collaborative module.

6. The driver abnormality monitoring system based on large model collaborative decision-making according to claim 5 is characterized in that: The CH chaotic map initialization parameters are introduced to integrate the behavior patterns of the Lianfeng horned toad, namely foraging, defense, escape, and social update gradient directions. The CH chaotic map initialization parameter formula is as follows: ; Where α≥0, β≥0 and α+β≤1, u∈[0,4], γ ∈[2.5,3.0], are mapping control parameters, z i is the current chaotic state value, z i+1 is the next state value output by the CH chaotic map, and satisfies z i+1 ∈[0,1]; w i is the intermediate weight value of the i-th iteration used to generate the chaotic sequence; λ is the interpolation parameter, which controls the balance between deterministic and random components; r is the random factor, which introduces random perturbations to enhance ergodicity; π is pi, which is approximately equal to 3.14159; sin is the sine function, cos is the cosine function, and || is the absolute value function, which ensures that the output is non-negative.

7. A driver abnormality monitoring method, implemented by the driver abnormality monitoring system based on large model collaborative decision-making according to claim 1, characterized in that: The following steps are involved: (1) Capture in-car video streams at 25 FPS using multiple cameras; (2) Generate scene text descriptions in real time using the DTOPA model; (3) Use the Qwen-Apan model to parse the text and output a risk assessment; (4) If the detection is abnormal, start the voice interaction confirmation state; (5) Perform security response operations based on the interaction results.

8. A driver abnormality monitoring method according to claim 7, characterized in that: In step (1), a multimodal data enhancement strategy is adopted to generate a driving text video DTV dataset.

9. A driver abnormality monitoring method according to claim 7, characterized in that: In step (2), a 21-second sliding window mechanism is used to randomly select 20 frames to form an analysis batch, and the compressed frame data is transmitted through WebSocket.

Citation Information

Patent Citations

  • Driving state monitoring and feedback method and system based on multi-mode human factor intelligent data analysis

    CN117909810A

  • Automatic driving large model framework based on 3D space-time perception and human-like decision-making reasoning

    CN120182938A