Multi-screen voice interaction system and method applied to automobile cabin

Through the combination of the voice acquisition unit, multimodal perception unit and collaborative dispatching unit, the problem of noise interference and insufficient adaptability of the multi-screen voice interaction system in the vehicle is solved, high-precision voice recognition and dynamic interaction are achieved, and user experience and security are improved.

CN120299457APending Publication Date: 2025-07-11RIVOTEK TECH (JIANGSU) CO LTD
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202510439450.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing multi-screen voice interaction system is disturbed by noise in the interior environment, has low recognition accuracy, lacks multimodal perception capabilities, and cannot adaptively adjust under different driving conditions, affecting user experience and safety.

Method used

The voice acquisition unit, multimodal perception unit, human-computer interaction unit and collaborative scheduling unit are adopted, combined with adaptive filtering, deep learning and multi-sensor fusion technology to realize noise cancellation, multimodal information acquisition and adaptive interaction strategy optimization, and dynamically adjust the display content and interaction mode.

Benefits of technology

It improves the accuracy of voice recognition, fully understands the needs of drivers and passengers, improves the intelligence and safety of the system, dynamically adapts to different driving environments, and enhances user experience and operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299457A_ABST
    Figure CN120299457A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-screen voice interaction system and method applied in an automobile cabin, and relates to the technical field of vehicle-mounted intelligent interaction, and the system comprises a voice collection unit which is used for carrying out sound source positioning and collection, carrying out the noise reduction processing of a collected voice signal, extracting features from the processed voice signal, and generating a voice feature vector; the multi-modal sensing unit is used for collecting behavior information and physiological state data of a driver and passengers through multi-sensor fusion, so as to extract multi-modal information; the man-machine interaction unit is used for carrying out space-time modeling on the multi-modal information to generate a scene state vector; and the cooperative scheduling unit is used for realizing multi-screen intelligent distribution and cooperative control. According to the invention, voice instructions of a driver and passengers can be accurately identified, the safety and experience of the driver are improved, changes in different driving environments are dynamically adapted, the operation efficiency and user experience of a vehicle-mounted system are improved, and the intelligence and adaptability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of in-vehicle intelligent interaction, and particularly to a multi-screen voice interaction system and method applied to an automotive cockpit. Background Art

[0002] With the rapid development of intelligent technologies, the automotive industry is gradually moving towards intelligent cockpits, and the interaction mode between drivers and in-vehicle systems is gradually changing from traditional button operations, touchscreens, etc. to more natural and intelligent voice interactions. As an important part of this trend, the multi-screen voice interaction system has broad application prospects. It can interact with multi-screen display devices in the vehicle through voice commands to meet various needs such as information display, navigation control, and entertainment functions. Currently, in-vehicle interaction systems based on speech recognition and multi-modal perception technologies have been initially applied in the market, including the integrated application of voice control systems, gesture recognition systems, and biometric technologies. The progress of these technologies not only improves the user experience but also optimizes the safety and convenience to a certain extent during the driving process.

[0003] However, despite the certain progress made by existing multi-screen voice interaction systems, they still face a series of technical challenges. Traditional speech recognition systems are often interfered by background noise. Especially in the in-vehicle environment, factors such as engine noise and road noise will affect the acquisition and recognition accuracy of speech signals. Most existing systems rely only on single-sensor data or voice input for interaction and lack the ability of multi-modal perception, resulting in insufficient understanding of the actual needs and states of drivers and passengers in complex driving scenarios. Existing multi-screen display control systems usually lack the ability to adaptively adjust to changing driving environments and cannot dynamically optimize the display content and interaction methods under different driving states, vehicle speeds, and road conditions, which not only reduces the intelligence level of the system but also affects the driver's experience and safety. Summary of the Invention

[0004] In view of the problems existing in the existing multi-screen voice interaction system applied to an automotive cockpit, the present invention is proposed. Therefore, the problem to be solved by the present invention is how to provide a multi-screen voice interaction system and method applied to an automotive cockpit.

[0005] To solve the above technical problems, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a multi-screen voice interaction system applied to an automotive cockpit, which includes: a voice acquisition unit, including a voice acquisition module, a noise cancellation module, and a feature extraction module, for performing sound source localization and acquisition, and performing noise reduction processing on the acquired voice signal based on an adaptive filtering algorithm, extracting features from the processed voice signal, and generating a voice feature vector;

[0007] The multi-modal perception unit, including a visual monitoring module, an action recognition module, and a physiological feature analysis module, is used to collect the behavior information and physiological state data of the driver and passengers through multi-sensor fusion, and generate multi-modal information;

[0008] The human-computer interaction unit, including a scene understanding module, a decision-making generation module, and a feedback optimization module, where the scene understanding module performs spatio-temporal modeling on the multi-modal information to generate a scene state vector, the decision-making generation module constructs an interaction strategy sample library, and the feedback optimization module is used to adaptively optimize the interaction strategy;

[0009] The collaborative scheduling unit, including a deep learning module and a display control module, is used to achieve multi-screen intelligent distribution and collaborative control.

[0010] As a preferred solution of the multi-screen voice interaction system applied to the automotive cockpit according to the present invention, wherein: the feature extraction module includes:

[0011] The voice preprocessing component is used to perform frame processing on the voice signal;

[0012] The acoustic feature component is used to extract acoustic features from the preprocessed voice signal to form an acoustic feature matrix;

[0013] The semantic understanding component is used to perform semantic encoding and intention recognition on the text converted by automatic speech recognition to obtain semantic features and form a semantic feature matrix;

[0014] The emotion analysis component combines the acoustic features and semantic features to perform multi-dimensional emotion recognition to obtain emotion features and form an emotion feature matrix;

[0015] The features extracted by each component are fused through linear weighting to form a voice feature vector, expressed as:

[0016] V S = W1 × F a + W2 × F s + W3 × F p

[0017] Wherein, V S is the voice feature vector, F a is the acoustic feature matrix, F s is the semantic feature matrix, F p is the emotion feature matrix, and W1 to W3 are the corresponding feature weight matrices, and the feature weight matrices are optimized through the backpropagation algorithm.

[0018] As a preferred solution of the multi-screen voice interaction system applied to the automotive cockpit according to the present invention, wherein: the visual monitoring module includes:

[0019] Pupil tracking component, which uses an infrared camera to capture pupil size and gaze direction in real time;

[0020] Facial expression component, extracts facial micro-expression features based on deep residual network;

[0021] The action recognition module comprises:

[0022] The gesture recognition component obtains the three-dimensional trajectory sequence of the gesture through the depth camera;

[0023] The posture estimation component combines multi-view images to detect key points of the human body and obtain the coordinates of key points of the human body;

[0024] The physiological characteristics analysis module comprises:

[0025] The fatigue detection component is used to integrate multi-dimensional physiological characteristics to comprehensively evaluate the fatigue level of the subject and conduct fatigue level assessment.

[0026] As a preferred solution of the multi-screen voice interaction system applied to the car cabin of the present invention, the scene understanding module uses a spatiotemporal graph convolutional network to perform scene modeling on multimodal information, and the formula is:

[0027] V t =σ(W v ×X t +b v )

[0028]

[0029] S t =f(V t ,A t ,E t )

[0030] Among them, S t represents the scene state vector at time t, V t is the node feature matrix at time t, A t is the attention matrix at time t, E t is the edge feature matrix at time t, σ is the sigmoid activation function, W v is the learnable weight matrix, b v is the bias term, X t is the input feature at time t, Q t , K t is the query and key value matrix at time t, T represents the matrix transpose; d is the feature dimension.

[0031] As a preferred solution of the multi-screen voice interaction system applied to the car cabin of the present invention, the deep learning module adopts a hierarchical training strategy:

[0032] The feature extraction layer uses a pre-trained model for transfer learning to extract multi-modal features, and inputs the speech feature vector and the scene state vector into the pre-trained model;

[0033] The feature fusion layer designs an attention gating mechanism to achieve feature adaptive fusion of the speech feature vector and the scene state vector;

[0034] The temporal modeling layer obtains the dynamic changes of the fused features in time series;

[0035] The decision output layer optimizes multiple objectives simultaneously based on a multi-task learning framework for joint prediction of multiple tasks.

[0036] As a preferred solution of the multi-screen voice interaction system applied to the automotive cockpit described in the present invention, wherein: the display control module includes:

[0037] Perform priority sorting according to the importance and urgency of the information content to generate a content priority vector;

[0038] Calculate the real-time attention distribution according to the pupil feature vector, and dynamically allocate display resources according to the user's real-time attention distribution;

[0039] Real-time monitor the display information of each screen to generate a display status matrix;

[0040] Detect the display device to determine the display format and layout of the content, and output a display parameter matrix.

[0041] As a preferred solution of the multi-screen voice interaction system applied to the automotive cockpit described in the present invention, wherein: it further includes an interactive security guarantee mechanism:

[0042] Calculate the attention dispersion degree to evaluate the driver's attention state;

[0043] Dynamically adjust the task complexity according to the current vehicle speed and road conditions, trigger multi-modal warnings when dangerous conditions are detected, comprehensively calculate the safety score, and adjust the interaction system according to the safety score. The formula for the safety score Safe sco is:

[0044]

[0045] D att =∑w i ·x i

[0046] where D att is the attention dispersion degree, V speed is the vehicle speed factor, R road is the road condition factor, T task is the task complexity, wi is the attention feature weight, and x i is the i-th attention feature index;

[0047] When the safety score Safe sco is lower than the preset threshold, the display control module adjusts the display parameter matrix.

[0048] In a second aspect, the present invention provides a multi-screen voice interaction method applied to an automotive cockpit, which includes:

[0049] Collect cockpit voice signals and eliminate background noise to generate voice feature vectors;

[0050] Collect behavior information and physiological state data to extract multi-modal information;

[0051] Perform spatio-temporal modeling on the multi-modal information to generate a scene state vector;

[0052] Input the generated voice feature vectors and scene state vectors into the collaborative scheduling unit for display content allocation.

[0053] The beneficial effects of the present invention are as follows: It can accurately identify the voice commands of drivers and passengers, ensure the efficient operation of the in-vehicle voice system in complex environments, can perform more personalized and intelligent interactions according to the driver's state, improve the safety and experience of the driver, dynamically adapt to changes in different driving environments, ensure the long-term learning and adaptation ability of the system, improve the operation efficiency and user experience of the in-vehicle system, and enhance the intelligence and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0055] Figure 1 It is a structural diagram of a multi-screen voice interaction system applied to an automotive cockpit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the above objects, features, and advantages of the present invention more understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0057] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Persons skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0058] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separate or mutually exclusive of other embodiments.

[0059] Referring to Figure 1 , which is the first embodiment of the present invention. This embodiment provides a multi-screen voice interaction system applied to an automotive cockpit, including:

[0060] A voice acquisition unit, including a voice acquisition module, a noise cancellation module, and a feature extraction module. This unit uses a double-layer microphone array for omnidirectional sound source localization and acquisition, and eliminates in-vehicle noise based on an adaptive filtering algorithm;

[0061] Specifically, 8 omnidirectional microphones are provided in the voice acquisition module, which are distributed in a circular array. The radius of the circular array is 15 cm. The sampling frequency of the voice signal is 16 kHz, and it is segmented into several frames according to a length of 25 ms and an overlap of 10 ms; the noise cancellation module performs noise reduction processing on the acquired signal based on the adaptive least mean square algorithm;

[0062] The feature extraction module includes a voice preprocessing component that performs frame segmentation on the original voice signal;

[0063] The voice signal processed by the noise cancellation module is decomposed into sub-signals of multiple scales through db4 wavelet transform. Each sub-signal corresponds to a different frequency band. Threshold processing is performed on each layer of wavelet coefficients to filter out the coefficients with relatively small amplitudes (considered as noise), and the coefficients after threshold processing are reconstructed to obtain the voice signal after noise suppression.

[0064] An acoustic feature component for extracting Mel frequency cepstral coefficients, fundamental frequency contours, and sound intensity curves;

[0065] Extract acoustic features reflecting the physical properties of the voice from the preprocessed signal. Use the autocorrelation function or the short-time Fourier transform (STFT) algorithm to perform fundamental frequency detection on each frame to obtain the fundamental frequency curve. Perform Fourier transform on each frame of the signal, and obtain Mel frequency cepstral coefficients after processing by the Mel filter bank. Calculate the square average of the amplitudes of each frame of the signal as a quantization index of the sound intensity. Integrate the features extracted from each frame to form an acoustic feature matrix.

[0066] The semantic understanding component performs semantic encoding and intent recognition on the text converted by automatic speech recognition (ASR), captures semantic information, and performs word segmentation, stop word removal, etc. on the text output by automatic speech recognition.

[0067] Use the pre-trained BERT model to encode the text, extract high-dimensional semantic vectors related to the context, and use the output of the [CLS] token in BERT as the semantic representation of the entire sentence to form a semantic feature vector.

[0068] Add a fully connected layer or a classifier on the basis of BERT to fine-tune the semantic feature vector, recognize the specific intent of the user, and update the model parameters through supervised learning and backpropagation algorithm to form a semantic feature matrix, which reflects the semantic information and intent content of the text.

[0069] The sentiment analysis component splices and fuses the acoustic features and semantic features, constructs a multi-layer neural network to process the fused features, outputs the probability distribution values of each sentiment dimension, and uses a labeled sentiment dataset to train through the mean squared error loss function to obtain a sentiment feature matrix.

[0070] Fuse the features extracted by each component through linear weighting to form the final speech feature vector model. The model of the speech feature vector is as follows:

[0071] V S = W1×F a + W2×F s + W3×F p

[0072] where, V S is the speech feature vector, F a is the acoustic feature matrix, F s is the semantic feature matrix, F p is the sentiment feature matrix, and W1 to W3 are the corresponding feature weight matrices, which are optimized through the backpropagation algorithm.

[0073] The multimodal perception unit includes a visual monitoring module, an action recognition module, and a physiological feature analysis module. This unit collects the behavior information and physiological state data of the driver and passengers through multi-sensor fusion to generate multimodal information;

[0074] The visual monitoring module includes a pupil tracking component and a facial expression component; a binocular infrared camera and a depth camera are set up, and the sampling frequency of the binocular infrared camera is 60Hz;

[0075] The pupil tracking component obtains the eye information of the measured person in real time, monitors the changes in pupil size and line of sight direction, and reflects physiological signals such as attention, emotional fluctuations, and fatigue status.

[0076] An infrared camera is used for image acquisition, and infrared light can penetrate ambient light interference to ensure clear eye images can be obtained in low-light or strong-light environments.

[0077] Set a fixed sampling frequency to ensure real-time capture of eye movement data.

[0078] The collected original image is preprocessed by grayscale conversion, noise filtering (such as Gaussian filtering), and other preprocessing operations.

[0079] The eye area is extracted through binarization, edge detection and morphological operations.

[0080] Circular Hough transform or ellipse fitting algorithm is used to detect pupil edge and calculate pupil diameter;

[0081] The estimation of the gaze direction can be achieved by modeling the geometric relationship between the eye corner and the pupil center, and combining it with the calibration process to convert the image coordinates into the actual gaze direction.

[0082] Perform Kalman filtering or mean filtering on continuous frame data to eliminate short-term jitter caused by acquisition noise and ensure the stability of tracking results.

[0083] The facial expression component extracts facial micro-expression features to capture subtle physiological changes such as emotions, psychological states, and stress levels.

[0084] The traditional Haar feature is used to detect the facial area, key points of the face image are detected (such as eyes, nose, mouth corners, etc.), and geometric alignment is performed to ensure that the subsequent feature extraction has consistent standardized input.

[0085] A pre-trained deep residual network is used as a feature extractor to capture subtle facial changes at different scales and levels of abstraction through a multi-layer structure.

[0086] Fine-tuning is performed on the micro-expression dataset to make the model more adaptable to the low-intensity and short-term changing characteristics of facial expressions, and a set of high-dimensional feature vectors are output to represent expression information at all levels for subsequent emotion or state analysis.

[0087] The extracted features are subjected to PCA dimensionality reduction, and the expression change trend is analyzed in combination with time series.

[0088] The action recognition module includes a gesture recognition component and a posture estimation component;

[0089] The gesture recognition component obtains the subject's hand movement information and reflects the intention, interaction status and physiological response (such as fatigue or distraction) by capturing the three-dimensional gesture trajectory.

[0090] The depth camera is used to set the appropriate frame rate and depth resolution to collect image sequences containing depth information, thereby obtaining the three-dimensional position data of the hand.

[0091] Preprocess the depth image, including noise filtering and background segmentation, extract the hand area, and use convolutional neural network to further locate the key areas of the hand.

[0092] The hand key points detected in continuous frames are tracked to construct the three-dimensional trajectory sequence of gestures, and the gesture movement patterns are recognized and classified using the temporal modeling method long short-term memory.

[0093] Determine the current action category and intention based on the gesture's trajectory, speed, acceleration and other features.

[0094] The posture estimation component captures information about the main joints of the human body through multi-view images, achieves accurate detection of human posture and movement, and reflects the overall physiological state and behavior patterns.

[0095] Deploy multiple cameras to synchronously capture human body images from different angles to ensure that complete information can be obtained regardless of how the subject moves, and perform geometric correction and time synchronization on each camera;

[0096] The deep learning model is used to detect the key points of the human body in each perspective image and obtain the coordinates of the joints (such as shoulders, elbows, knees, ankles, etc.). The key point information extracted independently by each camera needs to take into account the perspective difference and occlusion problems.

[0097] Through projection transformation and triangulation, the key point data of each perspective are fused to reconstruct the precise posture of the human body in three-dimensional space.

[0098] For the dynamic changes of key points, the timing information can be combined to detect posture stability and movement patterns.

[0099] The physiological characteristic analysis module includes a fatigue detection component, which integrates multi-dimensional physiological characteristics (such as pupil changes, facial micro-expressions, gestures and posture data) to comprehensively evaluate the fatigue level of the subject and provide objective physiological status indicators.

[0100] The feature data extracted from each component are normalized, feature selected, and dimension reduced to ensure the comparability of the data.

[0101] Feature concatenation, weighted fusion or attention mechanism is used to merge pupil, face, gesture and posture information into a unified physiological feature vector.

[0102] The fatigue assessment model is constructed based on supervised learning, and traditional machine learning methods (such as SVM, random forest) or deep learning methods (such as fully connected networks, convolutional temporal networks) can be used for modeling.

[0103] Train a model using the labeled fatigue data (such as scenario data for driving fatigue, work fatigue, etc.) so that it can accurately predict the degree of fatigue.

[0104] In practical applications, the model predicts the real-time collected fused feature vectors and outputs the fatigue score or classification result.

[0105] By setting thresholds or trigger mechanisms, corresponding reminder or intervention measures can be initiated when the degree of fatigue reaches the warning value.

[0106] The human-computer interaction unit includes a scenario understanding module, a decision-making generation module, and a feedback optimization module. This unit constructs an interaction strategy model based on multimodal information and uses the user's intention and interaction behavior as training samples;

[0107] The scenario understanding module performs spatio-temporal modeling on multimodal information. The decision-making generation module constructs an interaction strategy sample library. The scenario understanding module uses a spatio-temporal graph convolutional network for scenario modeling. The formula is as follows:

[0108] V t =σ(W v ×X t +b v )

[0109]

[0110] S t =f(V t ,A t ,E t )

[0111] Among them, S t represents the scenario state vector at time t, V t is the node feature matrix, A t is the attention matrix, E t is the edge feature matrix, σ is the sigmoid activation function, W v is the learnable weight matrix, b v is the bias term, X t is the input feature, Q t ,K t are the query and key value matrices, T represents matrix transpose; d is the feature dimension, set to 256.

[0112] The scenario understanding module jointly forms an input feature matrix with the pupil feature vector, facial feature vector, three-dimensional hand coordinate sequence, human key point coordinate vector, and fatigue degree index;

[0113] The feedback optimization module is used to continuously monitor and learn the user feedback after the system executes the interaction strategy, so as to continuously adjust and optimize the interaction strategy, and collect the indirect feedback of the user on the system response (such as whether the operation is interrupted, whether the user repeats the instruction, the change of facial satisfaction, etc.). Online fine-tuning of the policy model, dynamically updating the interaction strategy sample library and modeling parameters, so that the system gradually adapts to the behavior habits and preferences of specific users.

[0114] The collaborative scheduling unit includes a deep learning module and a display control module. The deep learning module uses an improved Transformer-XL structure to process the interaction sequence data;

[0115] The deep learning module adopts a hierarchical training strategy: using a pre-trained model for transfer learning to extract multi-modal features. The feature extraction layer inputs the speech feature vector and the scene state vector into the pre-trained model;

[0116] Extract key semantic information and environmental context information from the speech feature vector and the scene state vector.

[0117] Select a suitable pre-trained model according to the task characteristics (such as pre-trained networks in the speech field, image / scene feature extraction models, etc.), and load the pre-trained weights.

[0118] Jointly input the speech feature vector and the scene state vector, and after preprocessing such as normalization and alignment, send them into the pre-trained network.

[0119] Use the pre-trained model to perform forward propagation on the input to obtain high-level feature representations, which can better capture the semantic connotations of speech and scene information.

[0120] The obtained output features provide a reliable representation basis for the subsequent feature fusion layer, which helps to improve the generalization ability of the overall model.

[0121] Feature fusion layer: Achieve adaptive fusion of multi-modal features through an attention gating mechanism, enabling the model to dynamically adjust the degree of attention to each modal information according to different scenarios and interaction contexts. Concatenate the speech feature vector and the scene state vector obtained by the feature extraction layer to form a joint feature representation, and calculate the attention gating unit T g The formula for

[0122] T g is: g T S = σ(W t · [V g ; S

[0123] where W g is the gating weight matrix, and B g is the gating bias term;

[0124] The original features are weighted by element-wise multiplication using the attention gating unit to obtain the feature representation after adaptive fusion.

[0125] Temporal modeling layer: The temporal modeling layer adopts a bidirectional long short-term memory network structure. The hidden layer dimension of the bidirectional long short-term memory network is 512. The feature sequence after attention gating fusion is used as the input to the bidirectional long short-term memory network. After processing the entire sequence, the hidden state at each time step is output.

[0126] Decision output layer: Based on the multi-task learning framework, multiple objectives are optimized simultaneously. Multiple objectives are optimized through the shared feature representation to achieve the joint prediction of multiple tasks such as interaction categories, continuous values, and temporal stability.

[0127] Multiple output branches are constructed. Each branch corresponds to a task. Classification branch: The interaction category prediction is output using a fully connected layer and a Softmax activation function. Regression branch: The continuous value prediction is output, such as estimating the sentiment intensity or other continuous metrics. Temporal consistency branch: The output is designed through temporal constraints or smoothing terms to ensure the continuity of the prediction over time.

[0128] The weights between different tasks are balanced through a multi-objective loss function to ensure that multiple tasks can be optimized. At the same time, the model complexity is controlled to prevent overfitting. The loss function L to is designed as follows:

[0129] L to = α·L cl + β·L re + γ·L te + λ·R(θ)

[0130] where L cl is the interaction category loss, L re is the continuous value regression loss, L te is the temporal consistency loss, R(θ) is the regularization term, and α, β, γ, λ are the balance coefficients.

[0131] The interaction category loss uses cross-entropy loss to measure the prediction accuracy of the model in the classification task. The continuous value regression loss selects mean squared error loss to evaluate the prediction effect of the model on continuous metrics. The temporal consistency loss uses the difference method to constrain the smoothness of the model output between adjacent time steps and reduce prediction fluctuations. The regularization term uses L2 regularization to impose constraints on the model parameters to prevent overfitting. The balance coefficients are the weight coefficients of each loss and are determined by empirical tuning of parameters.

[0132] The display control module includes: prioritizing interactive content according to the content importance and urgency of information, ensuring that high-priority information can obtain display resources preferentially during display, and generating a content priority vector.

[0133] Specifically, perform semantic and metadata analysis on interactive content (such as text, images, videos, etc.), extract key information (such as urgency, topic relevance, historical priority tags, etc.), and use a machine learning-based classifier to pre-judge and score the content.

[0134] Divide the interactive content into multiple levels according to preset criteria (for example: urgent, important, general, low priority). Each content item corresponds to a vector representing the priority, forming a content priority vector.

[0135] Calculate the attention distribution vector based on the pupil feature vector, and dynamically allocate display resources according to the user's real-time attention state, so that the areas or devices that the user pays more attention to can obtain more display content.

[0136] Furthermore, input the pupil feature vector collected by the pupil tracking unit into the resource allocation algorithm. After data normalization and feature mapping, calculate the attention distribution of the user on each display area or device.

[0137] Real-time monitor the display status of each screen, including the current display content, refresh status, and user interaction information, and update and maintain the display status matrix, where each element represents the status information of a certain screen at a certain moment (such as whether it is idle, the type of display content, the current brightness or resolution, etc.).

[0138] Use the network communication protocol to update and broadcast the status information of each screen in real time, ensure data consistency and status synchronization of each terminal, and ensure that the distribution strategy matches the actual display conditions.

[0139] Detect the hardware parameters of each display device, and obtain information such as screen size, resolution, color capability, and current usage environment.

[0140] According to the device detection results and the type of interactive content, determine the best display format and layout of the content through preset rules or adaptive algorithms (such as a layout adjustment model based on deep learning).

[0141] Output the display parameter matrix, which includes display adjustment parameters such as font size, image resolution, typesetting style, background color, etc., and adjust the display parameter matrix in real time to ensure that the display content is always in the best adaptation state.

[0142] Based on the comprehensive consideration of content priority, user attention, display adaptability, and scene relevance, a distribution score is calculated for each content and each screen to determine the optimal distribution plan of the content in the multi-screen system. The distribution score function is as follows:

[0143] Score(i,j) = P co(i) ·A u(j) ·D sc(i,j) ·S con

[0144] Among them, Score(i,j) is the distribution score for content i on screen j, i is the content index, j is the screen index, P co(i) is the priority vector of content i, A u(j) is the j-th value in the user attention distribution, reflecting the degree of user attention to screen j, D sc(i,j) represents the adaptability degree of content i and screen j in terms of format, resolution, layout, etc., obtained by comparing the display status matrix with the actual screen parameter matrix, S con is the scene relevance, indicating the dynamic matching degree between the current environment or usage scenario and the content.

[0145] According to the score, the content is assigned to the screen with the highest corresponding score, and the information with a higher score is preferentially displayed.

[0146] As the user attention and environment change, A u(j) and S con are updated in real time to keep the distribution score consistent with the actual usage scenario, realizing dynamic allocation and adaptive display.

[0147] The attention state of the driver is collected and evaluated in real time. By monitoring key indicators such as eye movement, fixation duration, and blink frequency, it is judged whether there is a risk of attention distraction or fatigue.

[0148] The visual and behavioral data of the driver are collected in real time by using in-vehicle cameras, infrared eye trackers, or other physiological sensors.

[0149] After preprocessing (denoising, normalization) the collected data, multiple attention feature indicators are extracted, and attention feature weights are assigned to each indicator. The weights are determined by historical data statistics and reflect the influence of different indicators on the attention dispersion degree. The attention dispersion degree is calculated. The higher the value, the more distracted the driver's attention. The formula is:

[0150] D att = ∑w i ·x i

[0151] Among them, D att is the attention dispersion degree, w i is the attention feature weight, xi is the i-th attention feature index;

[0152] Dynamically adjust the complexity of the in-vehicle interaction system according to the current vehicle speed and road conditions, reducing the safety risks caused by information overload during driving.

[0153] Obtain the real-time vehicle speed through the vehicle's built-in sensors. When the vehicle speed is high, unnecessary interaction information should be reduced.

[0154] Collect road condition data, evaluate factors such as traffic flow and road conditions to obtain a road condition factor. When the road conditions are complex or dangerous, automatically simplify the interaction content or turn off some functions.

[0155] Define the complexity of the current interaction task (evaluated according to indicators such as the amount of information, the number of operation steps, and the interface complexity). When the complexity of the interaction task is high and the driving environment is relatively tense, the system actively reduces the complexity or delays non-urgent operations.

[0156] When detecting potential dangers (such as low driver attention, too high vehicle speed, poor road conditions, or task complexity mismatch), quickly trigger a multi-modal warning mechanism to remind the driver to take emergency measures, and set the warning threshold of the safety score according to historical data and safety standards;

[0157] Input the data obtained by each sensor in real time into the model to calculate the comprehensive safety score. The safety score model is:

[0158]

[0159] where D att is the attention distraction degree, V speed is the vehicle speed factor, R road is the road condition factor, T task is the task complexity.

[0160] Dynamically adjust the interaction system according to the safety score. When the safety score is high, more interaction operations are allowed. When the safety score drops below the warning threshold, immediately restrict the interaction operations and trigger an emergency warning. The warning methods include visual (such as a warning icon flashing on the instrument panel or the central control screen), auditory (sounding an alarm), and tactile (vibrations of the steering wheel or seat) feedback.

[0161] After the warning is activated, the system continues to monitor each indicator in real time to ensure that the driver quickly regains attention, and dynamically adjusts the warning level or prompt information according to the feedback.

[0162] The input end of the multi-modal perception unit is connected to the output end of the voice collection unit, the input end of the human-computer interaction unit is connected to the output end of the multi-modal perception unit, and the input end of the collaborative scheduling unit is connected to the output end of the human-computer interaction unit.

[0163] This embodiment also provides a multi-screen voice interaction method applied to an automotive cockpit, including the following steps: collecting cockpit voice signals and eliminating background noise to generate voice feature vectors;

[0164] Collecting behavior information and physiological state data to extract multimodal information;

[0165] Performing spatio-temporal modeling on the multimodal information to generate a scene state vector;

[0166] Inputting the generated voice feature vectors and scene state vectors into a collaborative scheduling unit for display content allocation.

[0167] This embodiment also provides a computer device applicable to the case of a multi-screen voice interaction system in an automotive cockpit, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement all or part of the steps of the system described in the embodiments of the present invention as proposed in the above embodiments.

[0168] This embodiment also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, it executes the system in any optional implementation manner of the above embodiments. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (abbreviated as SRAM), electrically erasable programmable read-only memory (abbreviated as EEPROM), erasable programmable read-only memory (abbreviated as EPROM), programmable read-only memory (abbreviated as PROM), read-only memory (abbreviated as ROM), magnetic memory, flash memory, a magnetic disk or an optical disc.

[0169] The storage medium proposed in this embodiment and the data storage system proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0170] In summary, this method can accurately identify the voice commands of drivers and passengers, ensure the efficient operation of the in-vehicle voice system in complex environments, enable more personalized and intelligent interactions based on the driver's state, enhance the safety and experience of the driver, dynamically adapt to changes in different driving environments, ensure the system's long-term learning and adaptation capabilities, improve the operation efficiency and user experience of the in-vehicle system, and enhance the intelligence and adaptability of the system.

Claims

1. A multi-screen voice interaction system applied to an automotive cockpit, characterized in that: include: The speech acquisition unit includes a speech acquisition module, a noise elimination module and a feature extraction module, which are used to locate and acquire sound sources, and to perform noise reduction processing on the acquired speech signals based on an adaptive filtering algorithm, extract features from the processed speech signals, and generate speech feature vectors; A multimodal perception unit, including a visual monitoring module, an action recognition module and a physiological characteristic analysis module, is used to collect the behavior information and physiological status data of the driver and passengers through multi-sensor fusion to generate multimodal information; The human-computer interaction unit includes a scene understanding module, a decision generation module, and a feedback optimization module. The scene understanding module performs spatiotemporal modeling on multimodal information and generates scene state vectors. The decision generation module builds an interaction strategy sample library. The feedback optimization module is used to adaptively optimize the interaction strategy. The collaborative scheduling unit includes a deep learning module and a display control module, which is used to realize multi-screen intelligent distribution and collaborative control.

2. The multi-screen voice interaction system applied to the automotive cockpit according to claim 1, wherein: The feature extraction module comprises: A speech preprocessing component, used for performing frame processing on speech signals; An acoustic feature component, used to extract acoustic features from the preprocessed speech signal to form an acoustic feature matrix; The semantic understanding component is used to perform semantic encoding and intent recognition on the text converted by automatic speech recognition, obtain semantic features, and form a semantic feature matrix; The sentiment analysis component is used to combine acoustic features and semantic features for multi-dimensional sentiment recognition, obtain sentiment features, and form a sentiment feature matrix; The features extracted from each component are fused through linear weighting to form a speech feature vector, which is expressed as: V S = W1 × F a + W2 × F s + W3 × F p Among them, V S is the speech feature vector, F a is the acoustic feature matrix, F s is the semantic feature matrix, F p is the emotion feature matrix, and W1 to W3 are the corresponding feature weight matrices, which are optimized by the backpropagation algorithm.

3. The multi-screen voice interaction system applied to an automotive cockpit according to claim 2, wherein: The visual monitoring module comprises: Pupil tracking component, which uses an infrared camera to capture pupil size and gaze direction in real time; Facial expression component, extracts facial micro-expression features based on deep residual network; The action recognition module comprises: The gesture recognition component obtains the three-dimensional trajectory sequence of the gesture through the depth camera; The posture estimation component combines multi-view images to detect key points of the human body and obtain the coordinates of key points of the human body; The physiological characteristics analysis module comprises: The fatigue detection component is used to integrate multi-dimensional physiological characteristics to comprehensively evaluate the fatigue level of the subject and conduct fatigue level assessment.

4. The multi-screen voice interaction system applied to the automotive cockpit according to claim 3, wherein: The scene understanding module uses a spatiotemporal graph convolutional network to model the multimodal information scene, and the formula is: V t = σ(W v × X t + b v ) S t = f(V t , A t , E t ) Among them, S t represents the scene state vector at time t, V t is the node feature matrix at time t, A t is the attention matrix at time t, E t is the edge feature matrix at time t, σ is the sigmoid activation function, W v is the learnable weight matrix, b v is the bias term, X t is the input feature at time t, Q t , K t are the query and key-value matrices at time t, T represents matrix transpose; d is the feature dimension.

5. The multi-screen voice interaction system applied to an automotive cockpit according to claim 4, characterized in that: The deep learning module adopts a layered training strategy: Feature extraction layer: Use the pre-trained model for transfer learning, extract multimodal features, and input the speech feature vector and scene state vector into the pre-trained model; Feature fusion layer: Design an attention gating mechanism to achieve adaptive fusion of speech feature vectors and scene state vectors; The time series modeling layer obtains the dynamic changes of fusion features in time series; The decision output layer optimizes multiple objectives simultaneously based on the multi-task learning framework and performs joint prediction of multiple tasks.

6. The multi-screen voice interaction system applied to an automotive cockpit according to claim 5, characterized in that: The display control module comprises: Prioritize information according to its importance and urgency, and generate a content priority vector; Calculate the real-time attention distribution according to the pupil feature vector, and dynamically allocate display resources according to the user's real-time attention distribution; Monitor the display information of each screen in real time to generate a display status matrix; Detect the display device, determine the display format and layout of the content, and output the display parameter matrix.

7. The multi-screen voice interaction system applied to the automotive cockpit according to claim 6, wherein: It also includes an interactive security guarantee mechanism: Calculate the attention dispersion degree to evaluate the driver's attention state; Dynamically adjust the task complexity according to the current vehicle speed and road conditions, trigger multimodal warnings when dangerous situations are detected, comprehensively calculate the safety score, adjust the interaction system according to the safety score, and the safety score is Safe sco The formula for is as follows: Among them, D att is the distraction degree, V speed is the vehicle speed factor, R road is the road condition factor, T task is the task complexity, w i is the attention feature weight, x i is the i-th attention feature index; When the safety score Safe sco is lower than the preset threshold, the display control module adjusts the display parameter matrix.

8. A multi-screen voice interaction method applied to an automotive cockpit, based on the multi-screen voice interaction system applied to an automotive cockpit according to any one of claims 1 to 7, characterized in that: It includes: Collect cockpit voice signals and eliminate background noise to generate voice feature vectors; Collect behavior information and physiological state data to extract multi-modal information; Perform spatio-temporal modeling on the multi-modal information to generate a scene state vector; Input the generated voice feature vectors and scene state vectors into the collaborative scheduling unit for display content allocation.

Citation Information

Cited By

  • Exhibition hall voice interaction method and system

    CN120823832A

  • Vehicle-mounted braille communication system and braille coding method

    CN120891946A

  • Intelligent network connection automobile information interaction method and device based on category personnel judgment

    CN120909432A

  • Self-adaptive lifting control method for self-service consignment equipment fused with multi-modal data analysis

    CN121091677A

  • Multi-mode Al-sensing self-adaptive brightness-adjusting low-blue-light head-up display optical system

    CN121167641A