Vehicle self-adaptive control method, device and equipment, vehicle and medium
By collecting the driver's image, physiological and voice information, and utilizing multimodal data fusion and cross-attention mechanisms, the problem of low fatigue detection accuracy in existing technologies has been solved, achieving accurate fatigue state recognition and vehicle assisted control, thereby improving driving safety and intelligence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEXINLI INTELLIGENT CONTROL TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, driver fatigue detection relies on single eye image recognition, which is easily affected by lighting and differences in facial features, resulting in low detection accuracy, frequent false alarms and missed alarms, and an inability to accurately distinguish the degree of fatigue.
The system collects image, physiological, and audio information from drivers, identifies fatigue states through multimodal data fusion, and utilizes a cross-attention mechanism to fuse features from different modalities to generate accurate fatigue probabilities and control strategies.
It improves the accuracy and reliability of fatigue state recognition, can accurately distinguish different levels of fatigue, realize graded recognition and early warning, reduce the incidence of traffic accidents, and improve driving safety and intelligence.
Smart Images

Figure CN122035009A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle control technology, and in particular to a vehicle adaptive control method, device, equipment, vehicle, and medium. Background Technology
[0002] Fatigue driving is a significant hidden danger causing road traffic accidents and poses a serious threat to road traffic safety. Therefore, accurate detection of driver fatigue is of great practical significance. Currently, most vehicles use cameras to capture images of the driver's eyes and determine driver fatigue by recognizing the state of the eyes. This method is one of the mainstream non-contact detection techniques.
[0003] However, this type of detection method has significant drawbacks. It relies solely on eye features for judgment, resulting in a lack of detection dimensions and susceptibility to interference from factors such as lighting conditions and differences in driver facial features. This leads to low detection accuracy and a high likelihood of false alarms and missed detections. Therefore, there is a need to improve the precision and differentiation of safety warning capabilities. Summary of the Invention
[0004] This application provides a vehicle adaptive control method, device, equipment, vehicle, and medium to improve the accuracy of driver fatigue detection and ensure driving safety.
[0005] According to one aspect of this application, a vehicle adaptive control method is provided, comprising: Acquire driver status information collected from the driver of the current vehicle; the driver status information includes the driver's image information, physiological information and voice information; Determine the driver's fatigue state based on at least one driver status information; Determine the current vehicle control strategy based on fatigue status.
[0006] According to another aspect of this application, a vehicle adaptive control device is provided, comprising: The status information acquisition module is used to acquire driver status information collected for the current vehicle driver; the driver status information includes the driver's image information, physiological information and voice information; The fatigue state determination module is used to determine the driver's fatigue state based on at least one driver state information. The control strategy determination module is used to determine the current vehicle control strategy based on the fatigue state.
[0007] According to another aspect of this application, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the vehicle adaptive control method according to any embodiment of this application.
[0008] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the vehicle adaptive control method according to any embodiment of this application.
[0009] According to another aspect of this application, a computer program product is provided, the computer program product including a computer program that, when executed by a processor, implements the vehicle adaptive control method according to any embodiment of this application.
[0010] The technical solution of this application embodiment acquires driver status information collected from the driver of the current vehicle; wherein, the driver status information includes the driver's image information, physiological information, and voice information; based on at least one driver status information, the driver's fatigue state is determined; and based on the fatigue state, the control strategy of the current vehicle is determined. By collecting the driver's image information, physiological information, and voice information, and using a multimodal data fusion approach to identify the driver's fatigue state, and performing vehicle auxiliary control based on the identification results, it has significant technical advantages and safety value. Compared with the existing technology that relies solely on eye images for detection, multimodal data covers the driver's appearance, physiological signs, and voice characteristics, providing a more comprehensive detection dimension. This effectively avoids the problem of false alarms and missed alarms caused by external interference with single data, significantly improving the accuracy and reliability of fatigue state identification. At the same time, it can accurately distinguish different levels of fatigue, achieving graded identification and early warning of fatigue states. Furthermore, vehicle assisted control based on accurate fatigue recognition results can intervene in fatigued driving behavior in a timely manner, reminding drivers to pay attention to their own condition, effectively reducing the incidence of road traffic accidents caused by fatigued driving, ensuring the personal safety of drivers and road traffic safety, meeting the precise and differentiated early warning and control needs in the field of vehicle driving safety, and improving the safety and intelligence level of vehicle driving.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a vehicle adaptive control method according to Embodiment 1 of this application; Figure 2 This is a flowchart of a vehicle adaptive control method according to Embodiment 2 of this application; Figure 3 This is a schematic diagram of a vehicle adaptive control device according to Embodiment 3 of this application; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the vehicle adaptive control method of the embodiments of this application. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 Figure 1This application provides a flowchart of a vehicle adaptive control method according to Embodiment 1. This embodiment is applicable to situations where vehicle control is adjusted based on the driver's fatigue state. This method can be executed by a vehicle adaptive control device, which can be implemented in hardware and / or software. The vehicle adaptive control device can be configured in an electronic device, which can be deployed in a vehicle. Figure 1 As shown, the method includes: S110. Obtain driver status information collected for the current vehicle's driver; wherein, driver status information includes the driver's image information, physiological information and voice information.
[0017] The vehicle in question can be any vehicle requiring driver fatigue detection, especially one that is currently being driven. Driver status information can be data on the driver's physical or mental state during driving, including image information, physiological information, and audio information. Image information can be images of the driver captured by cameras installed inside the vehicle. The image composition can include the driver's head, face, and upper torso, and this image information can be used to identify the driver's facial expressions, posture, and movements. Physiological information can include the driver's heart rate and respiratory rate, which can be acquired through contact-type physiological sensors mounted on the steering wheel or non-contact sensors based on millimeter-wave radar. Audio information can be any sound emitted by the driver, including speaking, laughter, yawning, and other non-verbal sounds. This application does not limit the scope of the image, physiological, and audio information described above.
[0018] S120. Determine the driver's fatigue state based on at least one driver status information.
[0019] The driver's fatigue state can be represented by data information that characterizes the degree of driver fatigue. For example, different state grading methods can be used to distinguish the degree of driver fatigue. It is understood that the driver's state information includes image information, physiological information, and voice information. This information is analyzed to determine the degree of the driver's current fatigue state. For example, a pre-trained image processing model can be used to classify and recognize image information to initially determine whether the driver's expressions and movements conform to the characteristics of fatigue. Alternatively, information such as the driver's heart rate and respiratory rate can be used to initially analyze whether the driver's physical state conforms to the physical characteristics of fatigue. Furthermore, based on the driver's voice information, a pre-trained audio frame analysis model can be used to initially identify the driver's verbal and non-verbal information, determining whether there are voice features related to fatigue. Finally, the results obtained from the analysis of information from different dimensions and modalities are summarized to determine the driver's fatigue state. For example, the features of the analysis results from different modalities can be converted into vector form using a vectorization model, and weighted vector calculations can be performed. The calculated results are compared with a pre-set standard corresponding to the fatigue level to determine the driver's current fatigue level. Of course, other methods can also be used for analysis, which will not be elaborated upon in this embodiment.
[0020] S130. Determine the current vehicle control strategy based on the fatigue state.
[0021] The control strategy can be a variety of strategies for controlling the working components inside the vehicle. Different control strategies are used to control the vehicle based on the driver's different fatigue levels. Since there are various fatigue levels, for example, there can be multiple different fatigue levels. The control strategy can be differentiated according to different fatigue levels. For example, for mild fatigue, the vehicle can be controlled to initiate comfort interventions, such as changing the ambient lighting color, playing soothing music, releasing fragrances to refresh the driver, or providing gentle voice reminders; for moderate fatigue (distraction, tension), the vehicle can be controlled to initiate warning interventions, such as appropriately tightening the seat belt, lowering the music volume, increasing the air conditioning fan speed, highlighting road information on the central control screen, or issuing voice warnings; for severe fatigue (or health abnormalities), the vehicle can be controlled to initiate safety interventions, such as flashing red lights in the cabin, vibrating the seats, issuing clear voice warnings, and simultaneously sending a takeover request to the vehicle dynamic control system to further take over the vehicle's control domain actuators, controlling the power system, braking system, and steering system, enabling the vehicle to brake in time and ensure safety.
[0022] The technical solution of this application embodiment acquires driver status information collected from the driver of the current vehicle; wherein, the driver status information includes the driver's image information, physiological information, and voice information; based on at least one driver status information, the driver's fatigue state is determined; and based on the fatigue state, the control strategy of the current vehicle is determined. By collecting the driver's image information, physiological information, and voice information, and using a multimodal data fusion approach to identify the driver's fatigue state, and performing vehicle auxiliary control based on the identification results, it has significant technical advantages and safety value. Compared with the existing technology that relies solely on eye images for detection, multimodal data covers the driver's appearance, physiological signs, and voice characteristics, providing a more comprehensive detection dimension. This effectively avoids the problem of false alarms and missed alarms caused by external interference with single data, significantly improving the accuracy and reliability of fatigue state identification. At the same time, it can accurately distinguish different levels of fatigue, achieving graded identification and early warning of fatigue states. Furthermore, vehicle assisted control based on accurate fatigue recognition results can intervene in fatigued driving behavior in a timely manner, reminding drivers to pay attention to their own condition, effectively reducing the incidence of road traffic accidents caused by fatigued driving, ensuring the personal safety of drivers and road traffic safety, meeting the precise and differentiated early warning and control needs in the field of vehicle driving safety, and improving the safety and intelligence level of vehicle driving.
[0023] Example 2 Figure 2 This is a flowchart of a vehicle adaptive control method provided in Embodiment 2 of this application. This embodiment further refines the method for determining the driver's fatigue state based on the foregoing embodiments. Figure 2 As shown, the method includes: S210. Obtain driver status information collected for the current vehicle's driver; wherein, driver status information includes the driver's image information, physiological information, and voice information.
[0024] S220. Extract the corresponding modal features from the state information of each driver to obtain the modal feature sequence corresponding to each modal feature.
[0025] Since the driver's state information includes data from different modalities, such as image information, physiological information, and audio information, feature extraction is performed on these different modalities to obtain corresponding modal features. These features are then analyzed and processed in the time domain to obtain a time-series-based modal feature sequence. Of course, feature extraction and time-domain analysis can employ any algorithm or technique from related technologies, and this application embodiment does not limit the specific methods used.
[0026] When processing a single modality, temporal attention can be used for feature processing. For example, when processing image sequences (such as consecutive video frames) or speech signals, not every frame or every syllable is equally important for judging fatigue. For instance, a long yawn is much more important than a normal breathing sound. A self-attention layer can be added after an LSTM (Long Short-Term Memory) or Transformer layer, allowing the model to automatically learn and assign higher weights to key temporal segments, thus enabling it to distinguish segments that are beneficial for judging fatigue states.
[0027] S230. Calculate the cross-attention of each modal feature based on the modal feature sequence.
[0028] To enable attention computation across different modalities in the same space, features need to be projected onto the same dimension. Simultaneously, since subsequent attention requires perceiving temporal order, positional encoding is added. A fully connected layer (or 1x1 convolution) is used to linearly transform each modality feature, resulting in vectors for different modalities in the same dimension. Due to the real-time streaming processing, absolute positional encoding cannot be used (because future positions are unknown); relative positional encoding or learnable positional embeddings are typically employed. For example, a fixed learnable vector is assigned to each baseline time step, and the projected features are then added to the positional encoding, thus giving each modality's features a time stamp. Based on this time stamp, cross-attention across all modalities can be computed in parallel. The method for computing cross-attention can employ any cross-attention mechanism from the relevant field; this application does not limit this approach.
[0029] S240. Generate fused features based on each cross-attention.
[0030] The fused features are obtained by concatenating the weighted features based on the weights of the different modal features determined in the cross-attention mechanism.
[0031] S250. Based on the fusion characteristics, determine the fatigue probability corresponding to at least one type of fatigue of the driver.
[0032] The spliced and fused features are passed through a pre-trained neural network classifier to obtain possible fatigue types and corresponding fatigue probabilities. For example, the probability of the current driver being mildly fatigued is 97%, the probability of being moderately fatigued is 2%, and the probability of being severely fatigued is 1%.
[0033] S260. Determine the fatigue state based on each fatigue probability.
[0034] The driver's current fatigue level is determined based on the probability of fatigue. Continuing the previous example, if the classifier outputs "the probability of the driver being mildly fatigued is 97%", which outweighs all other possibilities, then the driver's current fatigue level is determined to be mild fatigue.
[0035] S270. Determine the current vehicle control strategy based on fatigue status.
[0036] It's important to note that information from different modalities (such as images with closed eyes, low heart rate, and fatigued speech) needs to be analyzed jointly. Cross-attention mechanisms can calculate the correlations between features from different modalities. For example, when the model detects a yawn (a speech feature) in speech, it will use cross-attention to focus on mouth-opening movements in image features; and vice versa. This mechanism effectively captures the intrinsic connections between modalities, making the combined judgment of "closed eyes + yawn + low heart rate" more reliable than any single signal.
[0037] In the technical solution of this application embodiment, a cross-attention mechanism is used to determine and fuse the cross-attention of different modal features. This can accurately capture the correlation between various modal features, strengthen key features related to fatigue state, and weaken irrelevant interference features. Determining the fatigue type and corresponding probability based on the fused features can improve the accuracy and reliability of fatigue identification, achieve accurate determination of fatigue type and probability quantification, provide a more scientific basis for fatigue classification and early warning, and further ensure driving safety.
[0038] In one optional implementation, if the driver state information is audio information, step S220 involves extracting corresponding modal features from each driver state information to obtain a modal feature sequence corresponding to each modal feature, which may include: A1. Based on a preset threshold, energy threshold detection and zero-crossing rate analysis are performed on the sound information to determine the sound segments.
[0039] The audio segment can be a data segment that emits a specific sound, rather than noise such as wind noise or tire noise. The preset threshold can be a threshold line for energy threshold detection or zero-crossing rate analysis. For example, an energy threshold is set, and when the sound information exceeds this energy threshold, it is determined to be an audio segment. The zero-crossing rate refers to the number of times a signal crosses zero per unit time, which is an important indicator for measuring the time-domain characteristics of a signal and reflecting its frequency. Therefore, a zero value can be set, and the number of times the signal crosses zero can be recorded to calculate the zero-crossing rate of each segment, thereby determining the audio segment. Of course, the energy threshold and zero value can be preset by those skilled in the art according to the actual situation, and this application embodiment does not limit this.
[0040] A2. Perform feature analysis on audio segments to obtain linguistic and non-linguistic features.
[0041] Among them, linguistic features can be natural language features, corresponding to the features of the driver's speech in the audio segment; non-linguistic features can be the features of non-linguistic sounds made by the driver, such as the sound features of yawning, sighing or coughing.
[0042] Language segments and non-language segments can be classified first by a pre-trained classifier, and then language features and non-language features can be extracted separately; or features can be extracted first from spoken segments, and then language features and non-language features can be distinguished by a pre-trained feature recognition model. This application does not limit this approach.
[0043] A3. Determine the feature sequences for each modality based on linguistic and non-linguistic features.
[0044] After classifying linguistic and non-linguistic features, temporal analysis is performed on each to obtain time-based modal feature sequences. That is, each linguistic feature corresponds to a linguistic feature sequence, and each non-linguistic feature corresponds to a non-linguistic feature sequence.
[0045] In the above embodiments, by performing feature analysis on audio segments, linguistic features and non-linguistic features are determined respectively. Further classification of sound features provides more dimensions for identifying the driver's fatigue level. We can understand that drivers not only express their fatigue through language, but also reflect their fatigue state through non-linguistic content, which helps to improve the efficiency and accuracy of driver fatigue level identification.
[0046] In a further optional implementation, the feature analysis of the audio segment described in A2 to obtain linguistic and non-linguistic features may include: B1. Based on a pre-trained neural network model classifier, identify the language segments and non-language segments in the audio clips.
[0047] The audio segments can be segments of the driver's speech within an audio clip; understandably, the language features are also derived from the analysis of natural language segments within the audio clip. Non-verbal segments are segments of non-speaking sounds emitted by the driver, and the non-verbal features can be characteristic information corresponding to these non-verbal segments, such as, but not limited to, yawning, sighing, and coughing. A pre-trained neural network model is used to classify the audio clips, distinguishing between language and non-verbal segments. For example, audio clips can be classified using a neural network classifier, such as a pre-trained deep neural network model for recognizing non-verbal sounds like yawning, sighing, and coughing. These models are typically trained on large amounts of audio data labeled with specific events and can output results such as "the probability of it being a yawn is 95%." Correspondingly, other audio clips not classified as non-verbal segments by the classifier can be categorized as language segments.
[0048] B2. Determine the linguistic features and non-linguistic features respectively based on the linguistic fragments and non-linguistic fragments.
[0049] Feature extraction is performed separately for linguistic and non-linguistic segments to obtain linguistic and non-linguistic features. Any feature extraction method from the relevant field can be used, and this application does not impose any limitations on it.
[0050] In the above embodiments, by analyzing sound information, linguistic and non-linguistic segments are classified, and then linguistic and non-linguistic features are extracted. Classifying sound segments first and then extracting features separately enables precise separation and targeted extraction of sound features. Linguistic features can capture changes in the driver's tone of voice, while non-linguistic features (such as yawning and breathing sounds) can directly reflect fatigue status. The two complement each other, avoiding the one-sidedness of single feature extraction, improving the correlation between sound features and fatigue status, helping to improve the accuracy of multimodal fatigue recognition, providing more reliable sound dimension support for fatigue grading recognition, and contributing to improving the accuracy of vehicle driving safety warnings.
[0051] In one optional implementation, the determination of language features based on language segments as described in B2 may include: identifying primary language features and secondary language features in language segments according to a pre-set algorithm; wherein, primary language features include semantic features; and secondary language features include intonation, speech rate, and loudness.
[0052] The primary language features can be the features corresponding to the driver's spoken sentences and text, such as semantic features, which represent the meaning expressed by the driver. The secondary language features, on the other hand, represent how the driver expresses language features, that is, what tone or intonation they use. In specific feature data, these can be categorized as intonation features, speech rate features, and loudness features, etc.
[0053] For semantic features, a pre-trained large language model can be used to perform semantic analysis on the driver's output to determine whether there is semantic expression of fatigue in the discourse.
[0054] For intonation features, fundamental frequency can be used for description. The fundamental frequency value of each frame of speech can be extracted, and its statistics, such as mean, standard deviation, and range, can be calculated. When fatigued, the intonation usually becomes flat and monotonous, manifested by a narrowing of the fundamental frequency range and a decrease in the standard deviation.
[0055] Speech rate characteristics can be quantified by calculating the number of syllables or words per unit time. When fatigued, a driver's speech rate typically decreases significantly, and pauses become longer. Speech rate (words per second) can be calculated in real-time using text obtained through Automatic Speech Recognition (ASR) and its corresponding timestamps.
[0056] For loudness features, loudness refers to the intensity or energy of sound. The logarithmic energy of each frame of speech can be extracted. When fatigued, muscle relaxation and shallow breathing often lead to a decrease in volume. This can be characterized by calculating the average energy and energy variance over a period of time.
[0057] In the above embodiments, determining the semantic and non-semantic features in the driver's voice segments allows for multi-dimensional exploration of the correlation between sound and fatigue. Semantic features can capture the logic and emotional changes in the driver's language expression, while non-semantic features can directly reflect the physiological changes caused by fatigue. The two complement each other, avoiding the limitations of single feature extraction, improving the effectiveness of voice features, further enhancing the accuracy of multimodal fatigue recognition, providing reliable support for fatigue grading recognition, and helping to strengthen the pertinence and reliability of vehicle driving safety warnings.
[0058] In another alternative implementation, the determination of non-verbal features based on non-verbal fragments as described in B2 may include: converting non-verbal fragments into acoustic feature maps to identify non-verbal events; and converting each non-verbal event into a non-verbal feature.
[0059] Since nonverbal segments may include different nonverbal events (such as yawning, sighing, breathing, and coughing), it is necessary to distinguish these different nonverbal events for subsequent fatigue analysis. Acoustic feature maps can be a type of graph that transforms audio signals into visual representations, such as Mel spectrograms. Acoustic feature maps can intuitively show the changes in sound frequency over time, providing unique visual patterns for observation and analysis of sounds such as the slow, gradual descent of yawn frequency, coughing (sudden, broadband impact), and sighing (slowly descending tone). For example, Mel spectrograms can be input into a deep neural network for classification. For instance, convolutional neural networks excel at capturing local patterns in spectrograms and can effectively identify sounds with specific time-frequency morphologies, such as yawning and coughing; or Transformer-based audio models, such as AST (Audio Spectrogram Transformer), can be used. These models segment the spectrogram into patches and perform sequence modeling, capturing longer-range dependencies and performing excellently in complex audio event detection. These models output the confidence score for each event category. For example, for a 2-second audio clip, the model might output: yawn: 0.95; cough: 0.02; other: 0.03. This is used to identify non-verbal events.
[0060] In the above embodiments, by identifying the types of non-verbal events in non-verbal segments through acoustic feature maps, different fatigue-related events such as yawning and abnormal breathing can be accurately distinguished, the fatigue features of the sound dimension can be refined, the correlation between non-verbal events and fatigue state can be improved, and the accuracy of fatigue recognition can be further improved, providing more accurate support for driver fatigue classification analysis.
[0061] Example 3 Figure 3 This is a schematic diagram of a vehicle adaptive control device provided in Embodiment 3 of this application. Figure 3 As shown, the device 300 includes: The status information acquisition module 310 is used to acquire driver status information collected for the current vehicle driver; wherein, the driver status information includes the driver's image information, physiological information and voice information; The fatigue state determination module 320 is used to determine the driver's fatigue state based on at least one driver state information. The control strategy determination module 330 is used to determine the control strategy of the current vehicle based on the fatigue state.
[0062] The technical solution of this application embodiment acquires driver status information collected from the driver of the current vehicle; wherein, the driver status information includes the driver's image information, physiological information, and voice information; based on at least one driver status information, the driver's fatigue state is determined; and based on the fatigue state, the control strategy of the current vehicle is determined. By collecting the driver's image information, physiological information, and voice information, and using a multimodal data fusion approach to identify the driver's fatigue state, and performing vehicle auxiliary control based on the identification results, it has significant technical advantages and safety value. Compared with the existing technology that relies solely on eye images for detection, multimodal data covers the driver's appearance, physiological signs, and voice characteristics, providing a more comprehensive detection dimension. This effectively avoids the problem of false alarms and missed alarms caused by external interference with single data, significantly improving the accuracy and reliability of fatigue state identification. At the same time, it can accurately distinguish different levels of fatigue, achieving graded identification and early warning of fatigue states. Furthermore, vehicle assisted control based on accurate fatigue recognition results can intervene in fatigued driving behavior in a timely manner, reminding drivers to pay attention to their own condition, effectively reducing the incidence of road traffic accidents caused by fatigued driving, ensuring the personal safety of drivers and road traffic safety, meeting the precise and differentiated early warning and control needs in the field of vehicle driving safety, and improving the safety and intelligence level of vehicle driving.
[0063] In one optional embodiment, the fatigue state determination module 320 may include: The feature sequence determination unit is used to extract the corresponding modal features from each driver's state information to obtain the modal feature sequence corresponding to each modal feature; The attention determination unit is used to calculate the cross-attention of each modality feature based on the modality feature sequence; The fusion feature determination unit is used to generate fusion features based on each cross attention. The fatigue probability determination unit is used to determine the fatigue probability corresponding to at least one type of fatigue of the driver based on the fused features. The fatigue state determination unit is used to determine the fatigue state based on various fatigue probabilities.
[0064] In one optional implementation, if the driver status information is audible information, the feature sequence determination unit may include: Based on preset thresholds, energy threshold detection and zero-crossing rate analysis are performed on sound information to determine sound segments; The audio segment analysis subunit is used to perform feature analysis on audio segments to obtain linguistic and non-linguistic features; The feature sequence determination subunit is used to determine the feature sequence of each modality based on linguistic and non-linguistic features.
[0065] In one alternative implementation, the audio segment analysis subunit may include: Segment classification is a unit used by a pre-trained neural network model classifier to identify linguistic and non-linguistic segments within spoken segments; Feature determination units are used to determine linguistic features and non-linguistic features based on linguistic segments and non-linguistic segments, respectively.
[0066] In one alternative implementation, the feature determination from the unit may include: The language feature determination unit is used to identify primary and secondary language features in a language segment according to a pre-defined algorithm; the primary language features include semantic features; and the secondary language features include intonation, speech rate, and loudness.
[0067] In one alternative implementation, the feature determination from the unit may further include: The non-verbal event determination sub-unit is used to convert non-verbal segments into acoustic feature maps to identify non-verbal events; Non-linguistic feature determination sub-units are used to transform each non-linguistic event into a non-linguistic feature.
[0068] The vehicle adaptive control device provided in this application embodiment can execute the vehicle adaptive control method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing each vehicle adaptive control method.
[0069] Example 4 Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0070] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0071] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0072] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as vehicle adaptive control methods.
[0073] This application also provides a vehicle that can be equipped with or deployed electronic devices as provided in this application, in order to implement a vehicle adaptive control method provided in the foregoing embodiments and implementations of this application.
[0074] In some embodiments, the vehicle adaptive control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the vehicle adaptive control method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the vehicle adaptive control method by any other suitable means (e.g., by means of firmware).
[0075] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0076] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0077] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0078] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0079] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0080] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0081] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the vehicle adaptive control method provided in any embodiment of this application. This program product shares the same inventive concept as the vehicle adaptive control methods disclosed in the embodiments of this application, and therefore will not be described in detail here.
[0082] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0083] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A vehicle adaptive control method, characterized in that, include: Acquire driver status information collected from the driver of the current vehicle; wherein, the driver status information includes the driver's image information, physiological information, and voice information; The driver's fatigue state is determined based on at least one of the driver status information. Based on the fatigue state, determine the control strategy for the current vehicle.
2. The method according to claim 1, characterized in that, Determining the driver's fatigue state based on at least one of the driver state information includes: Extract the corresponding modal features from each of the driver state information to obtain the modal feature sequence corresponding to each modal feature; Based on the modal feature sequences, calculate the cross-attention of each modal feature; Based on the aforementioned cross-attention, fused features are generated; Based on the fusion features, determine the fatigue probability corresponding to at least one type of fatigue of the driver; The fatigue state is determined based on the fatigue probabilities described.
3. The method according to claim 2, characterized in that, If the driver state information is audio information, the step of extracting corresponding modal features from each driver state information to obtain a modal feature sequence corresponding to each modal feature includes: Based on a preset threshold, energy threshold detection and zero-crossing rate analysis are performed on the sound information to determine the sound segments; Feature analysis is performed on the audio segments to obtain linguistic and non-linguistic features; Based on the linguistic features and the non-linguistic features, each modality feature sequence is determined.
4. The method according to claim 3, characterized in that, The feature analysis of the spoken segment yields linguistic and non-linguistic features, including: Based on a pre-trained neural network model classifier, the speech segments and non-speech segments in the spoken segment are determined; Based on the linguistic fragments and the non-linguistic fragments, linguistic features and non-linguistic features are determined respectively.
5. The method according to claim 4, characterized in that, The step of determining language features based on the language fragment includes: According to a pre-set algorithm, the main language features and secondary language features in the language segment are identified; wherein, the main language features include semantic features; and the secondary language features include intonation, speech rate and loudness.
6. The method according to claim 4, characterized in that, The step of determining non-linguistic features based on the non-linguistic fragments includes: The non-verbal segments are converted into acoustic feature maps to identify non-verbal events; Each of the aforementioned non-linguistic events is transformed into a non-linguistic feature.
7. A vehicle adaptive control device, characterized in that, include: The status information acquisition module is used to acquire driver status information collected for the current vehicle driver; wherein, the driver status information includes the driver's image information, physiological information and voice information; A fatigue state determination module is used to determine the driver's fatigue state based on at least one of the driver state information. The control strategy determination module is used to determine the control strategy of the current vehicle based on the fatigue state.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the vehicle adaptive control method according to any one of claims 1-6.
9. A vehicle, characterized in that, The vehicle is equipped with the electronic equipment as described in claim 8, for implementing the vehicle adaptive control method as described in any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the vehicle adaptive control method according to any one of claims 1-6.