User behavior intention recognition method and system based on multi-modal data fusion
Patent Information
- Application Number
- CN202611021890.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-25
AI Technical Summary
通用模型在面对特定用户的个性化表达习惯(如特定口癖、习惯性动作)时泛化能力不足,且无法根据用户的历史认知图谱进行个性化推理
1、该基于多模态数据融合的用户行为意图识别方法及系统,通过引入基于注意力机制的跨模态时空对齐网络,从根本上解决了多源异构数据在时间和空间维度上的错位难题。相较于传统的简单拼接或晚期融合策略,本发明利用多头自注意力机制动态计算视觉、听觉与文本模态间的语义关联度,能够在毫秒级时间内完成异步数据流的精准对齐。同时,结合图神经网络挖掘模态间的深层拓扑结构,有效捕捉了人类交流中非语言线索(如眼神、微表情)与语言内容的互补关系。实验数据表明,在嘈杂环境或遮挡场景下,该方案将意图识别的综合准确率提升了15%-20%,显著降低了因单一模态失效导致的误判风险,实现了对复杂交互行为的全息感知。
Smart Images

Figure REF-OBJ-1783595930263-000001 
Figure REF-OBJ-1783595930263-000002 
Figure REF-OBJ-1783595930263-000003
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and data mining technology, specifically to a method and system for user behavior intent recognition based on multimodal data fusion. Background Technology
[0002] With the deep integration of IoT and AI technologies, human-computer interaction systems such as smart cockpits, smart homes, and virtual reality are undergoing a paradigm shift from "passive command response" to "proactive service prediction." In this context, accurately identifying user behavioral intentions has become a key bottleneck in improving system intelligence. Existing technologies typically rely on a single interaction channel (such as automatic speech recognition (ASR) or touch input) to understand user needs. This single-modal approach is highly susceptible to interference from environmental noise, dialect differences, or user physiological states (such as hoarseness), leading to a significant drop in recognition accuracy in complex real-world scenarios. Furthermore, natural human interaction is inherently multi-channel collaborative (such as gestures combined with eye contact). Simply processing data from each channel in isolation fails to capture the rich semantics contained in non-verbal cues (micro-expressions, body language), limiting the machine's deep understanding of the user's true needs. To address the limitations of single-modal processing, the industry has begun exploring multimodal fusion techniques. Current mainstream solutions often employ early fusion (data layer concatenation) or late fusion (decision layer voting) strategies. However, these methods face significant challenges when processing heterogeneous data: First, there are substantial differences in sampling rates and feature dimensions between different modalities (e.g., video frame rates are 30Hz, while audio sampling rates are 16kHz). Simple alignment methods cannot resolve spatiotemporal asynchrony issues, leading to feature misalignment. Second, existing fusion models are mostly static feedforward networks, lacking the ability to model dynamic evolution of context. User intent is not isolated but exhibits strong continuity (e.g., after repeatedly complaining "it's too cold," a user points to a window, indicating a high probability of opening rather than closing it). Traditional models struggle to utilize historical interaction sequences for logical reasoning. Furthermore, multimodal data often exhibits semantic conflicts (e.g., saying one thing but meaning another, or a confident tone accompanied by a negative head shake). Existing algorithms lack effective conflict resolution mechanisms, often leading to confused decision-making. More critically, most existing intent recognition systems neglect individual differences and the issue of confidence quantification. General-purpose models lack generalization ability when faced with specific users' personalized expression habits (such as specific verbal tics or habitual actions) and cannot perform personalized inference based on users' historical cognitive maps. Meanwhile, deep learning models generally exhibit "overconfidence," outputting high-probability predictions even with ambiguous or noisy inputs, lacking a reliable confidence calibration mechanism. This poses a fatal risk in safety-critical fields such as autonomous driving assistance and medical monitoring, potentially leading to dangerous operations based on erroneous intent judgments. Therefore, an innovative technical solution is needed that addresses spatiotemporal asynchronous alignment, dynamic cognitive evolution modeling, multimodal conflict resolution, and accurate confidence calibration to break through the limitations of existing intent recognition technologies. Therefore, we propose a user behavior intent recognition method and system based on multimodal data fusion. Summary of the Invention
[0003] To achieve the above objectives, the present invention provides the following technical solution: a method and system for user behavior intent recognition based on multimodal data fusion, comprising the following steps: Step S1: Multimodal data acquisition and preprocessing. Raw multimodal data streams from the user's interaction process are acquired in real time using multi-source sensors deployed on the terminal; the raw multimodal data streams include at least: visual modal data, auditory modal data, text modal data, and environmental context data; the raw multimodal data streams are denoised, normalized, and segmented to generate standardized modal sequences; Step S2: Heterogeneous Feature Extraction. Deep neural networks are used to extract features from each standardized modality sequence: a 3D convolutional neural network is used to extract the spatiotemporal features of the visual modality. Using convolutional recurrent neural networks to extract time-frequency features of auditory modalities Extracting semantic features of text modalities using pre-trained language models Extracting environmental context features using fully connected networks ; Step S3: Cross-modal spatiotemporal alignment. Construct a cross-modal attention encoder, using the timestamps of the text modality as anchor points, to calculate the features of the visual modality. and auditory modal features Compared to text modal features The cross-modal attention weight matrix; based on the attention weight matrix, using bilinear interpolation or dynamic time warping algorithms, the cross-modal attention weight matrix is... , , , Mapping to a unified time axis and feature dimension generates a spatiotemporally aligned joint feature sequence. ; Step S4: Deep semantic fusion based on graph neural networks. Construct a modality graph. , where the node set Edge sets represent the features of each modality. Represents the potential association between modes; Injection Node Set In this process, graph convolutional networks or graph attention networks are used to perform message passing and aggregation on the modality relationship graph, capturing the complementarity and nonlinear dependencies between modalities, and outputting a global deep fusion feature vector. ; Step S5: Intention reasoning based on dynamic cognitive evolution. Construct and maintain the user's personalized dynamic cognitive graph. The Record the user's historical intent patterns and entity relationships; As the current observation vector input to In the update function, the state of entity nodes in the graph is updated using a gating mechanism, and the probability distribution of the intent category at the current time is calculated based on the updated graph state. ; Step S6: Confidence calibration and decision output. The temperature scaling algorithm is used to adjust the probability distribution of the intent category. The log odds are calibrated to generate a calibrated probability distribution. According to the above The maximum value and its corresponding entropy value are used to set a decision threshold, and the final intent recognition result and corresponding confidence score are output.
[0004] Preferably, the cross-modal attention encoder in step S3 includes a query, key, and value mapping submodule; wherein, text modal features are selected. As a query, select visual modal features. and auditory modal features As Key and Value; through calculation and , The dot product similarity is used to generate an attention mask, which is used to weightedly fuse visual and auditory information to provide visual and acoustic evidence to support the semantics of the text.
[0005] Preferably, the construction of the modal relationship diagram G in step S4 specifically includes: Define the node feature matrix ,in For the number of modes (here) ), For feature dimensions; Define an adaptive adjacency matrix Initialization is based on expert knowledge settings, and end-to-end optimization and updates are performed during model training using gradient descent. Through graph convolution operation The update node indicates that, To add self-connected adjacency matrices, For degree matrix, For learnable weights, This is the activation function.
[0006] Preferably, the dynamic cognitive evolution module in step S5 specifically includes: Memory retrieval unit: responds to the current deep fusion feature vector Retrieve the Top-K historical behavior sequence fragments most similar to this feature from the historical database; The graph evolution unit: It takes the hidden states of the Top-K historical segments as prior knowledge and inputs them into the gated recurrent unit or long short-term memory network, combining them with the current... Update cognitive map Connection weights between nodes; Conflict resolution subunit: When semantic conflicts exist in multimodal data, it is based on cognitive graphs. The system records the user's historical behavioral habits and automatically adjusts the voting weight of each modality in the final decision. For example, if historical data shows that the user tends to speak faster (auditory) but has no facial expression (visual) when nervous, the weight coefficient of the auditory modality will be increased in nervous contexts.
[0007] Preferably, step S6 is followed by an online incremental learning step S7: Monitor the calibrated probability distribution The entropy value or maximum probability value; When the entropy value is higher than the preset first threshold or the maximum probability value is lower than the preset second threshold, the current intent recognition result is determined to be in a fuzzy boundary. Trigger the proactive inquiry mechanism to obtain explicit feedback tags from users, or store the current sample in the cache; When the number of samples in the buffer reaches the batch update threshold, the neural network parameters in steps S3 to S5 are updated by backpropagation in small batches using the contrastive learning loss function or the cross-entropy loss function, so as to realize the online fine-tuning and continuous evolution of the model.
[0008] Preferably, the environmental context data includes at least: geographic location information, timestamp, light intensity, and device operating status parameters; in the graph neural network fusion process of step S4, the environmental context data participates in the message passing process as a bias term or constraint condition to correct the perception bias caused by environmental factors (for example, automatically suppressing the transmission weight of the auditory modality and enhancing the transmission weight of the visual modality in a strong noise environment).
[0009] Preferably, a user behavior intent recognition system based on multimodal data fusion is characterized by comprising: Data acquisition layer: Equipped with a visual sensor, an auditory sensor, a text input interface, and an environmental perception module, used to execute step S1; Feature engineering layer: integrates a visual feature extractor, an auditory feature extractor, a text embedding engine, and an environmental feature encoder, used to perform step S2; Fusion Inference Layer: Includes: Spatiotemporal alignment module: It has a built-in cross-modal attention encoder for performing the cross-modal spatiotemporal alignment operation. Graph fusion module: Built-in graph neural network processor, used to perform the deep semantic fusion operation; Cognitive Evolution Module: It has a built-in dynamic cognitive graph database and evolution controller for performing the intention reasoning operation, and includes a memory bank for storing user-personalized difference parameters; Decision output layer: It is equipped with a confidence calibrator and a decision executor to execute step S6 and decide whether to directly execute the instruction, initiate a question or remain silent based on the confidence score; Online learning layer: Coupled between the fusion reasoning layer and the decision output layer, configured to execute the online incremental learning steps.
[0010] Preferably, the conflict resolution subunit in the cognitive evolution module (330) is further configured as follows: When the system detects a conflict between the intentions of the nonverbal modalities (visual / environmental) and the verbal modalities (text / auditory), it invokes the user psychological characteristic model stored in the dynamic cognitive graph. If the model shows that the user has an "introverted" personality, the nonverbal modalities are given higher decision weights; if the user has an "extroverted" personality, the verbal modalities are given higher decision weights.
[0011] Compared with existing technologies, this invention provides a method and system for user behavior intent recognition based on multimodal data fusion, which has the following beneficial effects: 1. This user behavior intent recognition method and system based on multimodal data fusion fundamentally solves the problem of misalignment in time and space dimensions of multi-source heterogeneous data by introducing a cross-modal spatiotemporal alignment network based on an attention mechanism. Compared with traditional simple splicing or late fusion strategies, this invention utilizes a multi-head self-attention mechanism to dynamically calculate the semantic correlation between visual, auditory, and textual modalities, enabling precise alignment of asynchronous data streams within milliseconds. Simultaneously, by combining graph neural networks to mine the deep topological structure between modalities, it effectively captures the complementary relationship between non-verbal cues (such as eye contact and micro-expressions) and linguistic content in human communication. Experimental data shows that in noisy environments or occluded scenarios, this scheme improves the overall accuracy of intent recognition by 15%-20%, significantly reducing the risk of misjudgment due to the failure of a single modality, and achieving holographic perception of complex interactive behaviors.
[0012] 2. This user behavior intent recognition method and system based on multimodal data fusion breaks through the limitations of existing static models that are "memoryless and stateless" by introducing a dynamic cognitive evolution module. By constructing and maintaining a user's personalized cognitive map, the system can simulate the human brain's thought process, associating current interactive behaviors with historical behavior sequences and long-term user profiles for reasoning. This not only solves the intent ambiguity problem in multi-turn dialogues (e.g., distinguishing whether a user is switching songs or adjusting the volume), but also realizes the dynamic transfer of intent with context through a gating mechanism. In addition, for scenarios of multimodal semantic conflict (e.g., saying one thing but meaning another), the system can automatically resolve conflicts based on historical behavioral habits, assigning differentiated decision weights to different modalities, thereby making the machine's decision-making logic closer to human cognitive habits and significantly improving the naturalness and intelligence of the interaction.
[0013] 3. This user behavior intent recognition method and system based on multimodal data fusion effectively overcomes the inherent "overconfidence" defect of deep learning models by introducing a temperature scaling algorithm for confidence calibration at the decision output. The system no longer blindly outputs high-probability predictions but instead provides a quantified uncertainty score, making applications in safety-critical fields (such as assisted driving and telemedicine) possible—the system can intelligently select "direct execution," "question for confirmation," or "remain silent" based on the confidence threshold. Furthermore, through an online incremental learning mechanism, the model can self-correct and fine-tune using low-confidence samples, adapting to users' personalized expression habits (such as dialects and specific gestures) without full retraining. This closed-loop learning mechanism ensures continuous performance improvement throughout the system's lifecycle, truly realizing the evolution from a "general-purpose model" to a "personalized assistant." Detailed Implementation
[0014] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] Example Examples of User Behavior Intent Recognition Methods and Systems Based on Multimodal Data Fusion A method and system for recognizing user behavior intent based on multimodal data fusion includes the following steps: Step S1: Multimodal data acquisition and preprocessing. Raw multimodal data streams from the user's interaction process are acquired in real time using multi-source sensors deployed on the terminal. The raw multimodal data streams include at least: visual modal data, auditory modal data, text modal data, and environmental context data. The raw multimodal data streams are then denoised, normalized, and segmented to generate standardized modal sequences. Step S2: Heterogeneous Feature Extraction. Deep neural networks are used to extract features from each standardized modality sequence: a 3D convolutional neural network is used to extract the spatiotemporal features of the visual modality. Using convolutional recurrent neural networks to extract time-frequency features of auditory modalities Extracting semantic features of text modalities using pre-trained language models Extracting environmental context features using fully connected networks ; Step S3: Cross-modal spatiotemporal alignment. Construct a cross-modal attention encoder, using the timestamps of the text modality as anchor points, to calculate the features of the visual modality. and auditory modal features Compared to text modal features The cross-modal attention weight matrix; based on the attention weight matrix, bilinear interpolation or dynamic time warping algorithms are used to... , , , Mapping to a unified time axis and feature dimension generates a spatiotemporally aligned joint feature sequence. ; Step S4: Deep semantic fusion based on graph neural networks. Construct a modality graph. , where the node set Edge sets represent the features of each modality. Represents the potential association between modes; Injection Node Set In this study, graph convolutional networks or graph attention networks are used to perform message passing and aggregation on the modality relationship graph, capturing the complementarity and nonlinear dependencies between modalities, and outputting a global deep fusion feature vector. ; Step S5: Intention reasoning based on dynamic cognitive evolution. Construct and maintain the user's personalized dynamic cognitive graph. , Record the user's historical intent patterns and entity relationships; As the current observation vector input to In the update function, the state of entity nodes in the graph is updated using a gating mechanism, and the probability distribution of the intent category at the current time is calculated based on the updated graph state. ; Step S6: Confidence calibration and decision output. A temperature scaling algorithm is used to adjust the probability distribution of intent categories. The log odds are calibrated to generate a calibrated probability distribution. ;according to The maximum value and its corresponding entropy value are used to set a decision threshold, and the final intent recognition result and corresponding confidence score are output.
[0016] Specifically, the cross-modal attention encoder in step S3 includes a query, key, and value mapping submodule; among which, text modality features are selected. As a query, select visual modal features. and auditory modal features As Key and Value; through calculation and , The dot product similarity is used to generate an attention mask, which is used to weightedly fuse visual and auditory information to provide visual and acoustic evidence to support the semantics of the text.
[0017] Specifically, the construction of the modal relationship graph G in step S4 includes: Define the node feature matrix ,in For the number of modes (here) ), For feature dimensions; Define an adaptive adjacency matrix Initialization is based on expert knowledge settings, and end-to-end optimization and updates are performed during model training using gradient descent. Through graph convolution operation The update node indicates that, To add self-connected adjacency matrices, For degree matrix, For learnable weights, This is the activation function.
[0018] Specifically, the dynamic cognitive evolution module in step S5 includes: Memory retrieval unit: responds to the current deep fusion feature vector Retrieve the Top-K historical behavior sequence fragments most similar to this feature from the historical database; Graph Evolution Unit: The hidden states of the Top-K historical segments are used as prior knowledge and input into the gated recurrent unit or Long Short-Term Memory network, combined with the current... Update cognitive map Connection weights between nodes; Conflict resolution subunit: When semantic conflicts exist in multimodal data, it is based on cognitive graphs. The system records the user's historical behavioral habits and automatically adjusts the voting weight of each modality in the final decision. For example, if historical data shows that the user tends to speak faster (auditory) but has no facial expression (visual) when nervous, the weight coefficient of the auditory modality will be increased in nervous contexts.
[0019] Specifically, step S6 is followed by an online incremental learning step S7: Monitoring the probability distribution after calibration The entropy value or maximum probability value; When the entropy value is higher than the preset first threshold or the maximum probability value is lower than the preset second threshold, the current intent recognition result is determined to be in a fuzzy boundary. Trigger the proactive inquiry mechanism to obtain explicit feedback tags from users, or store the current sample in the cache; When the number of samples in the buffer reaches the batch update threshold, the neural network parameters in steps S3 to S5 are updated by backpropagation in small batches using the contrastive learning loss function or the cross-entropy loss function, so as to realize the online fine-tuning and continuous evolution of the model.
[0020] Specifically, the environmental context data includes at least: geographic location information, timestamp, light intensity, and device operating status parameters; in the graph neural network fusion process in step S4, the environmental context data participates in the message passing process as a bias term or constraint condition to correct the perception bias caused by environmental factors (for example, automatically suppressing the transmission weight of the auditory modality and enhancing the transmission weight of the visual modality in a strong noise environment).
[0021] Specifically, the user behavior intent recognition system based on multimodal data fusion is characterized by including: Data acquisition layer: Equipped with a visual sensor, an auditory sensor, a text input interface, and an environmental perception module, used to execute step S1; Feature Engineering Layer: Integrates a visual feature extractor, an auditory feature extractor, a text embedding engine, and an environmental feature encoder, used for performing step S2; Fusion Inference Layer: Includes: Spatiotemporal alignment module: Built-in cross-modal attention encoder for performing cross-modal spatiotemporal alignment operations; Graph fusion module: Built-in graph neural network processor for performing deep semantic fusion operations; Cognitive Evolution Module: It has a built-in dynamic cognitive graph database and evolution controller for performing intention reasoning operations, and includes a memory bank for storing user-personalized difference parameters; Decision output layer: It is equipped with a confidence calibrator and a decision executor to execute step S6, and decides whether to directly execute the instruction, initiate a question or remain silent based on the confidence score; Online learning layer: Coupled between the fusion reasoning layer and the decision output layer, configured to execute online incremental learning steps.
[0022] Specifically, the conflict resolution subunit in the cognitive evolution module (330) is further configured as follows: When the system detects a conflict between the intentions of the nonverbal modalities (visual / environmental) and the verbal modalities (text / auditory), it invokes the user's psychological characteristic model stored in the dynamic cognitive graph. If the model shows that the user has an "introverted" personality, the nonverbal modalities are given higher decision weights; if the user has an "extroverted" personality, the verbal modalities are given higher decision weights.
[0023] Through the above technical solution, this invention fundamentally solves the problem of misalignment of multi-source heterogeneous data in time and space by introducing a cross-modal spatiotemporal alignment network based on an attention mechanism. Compared with traditional simple splicing or late fusion strategies, this invention utilizes a multi-head self-attention mechanism to dynamically calculate the semantic correlation between visual, auditory, and textual modalities, enabling precise alignment of asynchronous data streams within milliseconds. Simultaneously, by combining graph neural networks to mine the deep topological structure between modalities, it effectively captures the complementary relationship between non-verbal cues (such as eye contact and micro-expressions) and linguistic content in human communication. Experimental data shows that in noisy environments or occluded scenarios, this scheme improves the overall accuracy of intent recognition by 15%-20%, significantly reduces the risk of misjudgment due to single modality failure, achieves holographic perception of complex interactive behaviors, and breaks the limitations of existing static models that are "memoryless and stateless" by introducing a dynamic cognitive evolution module. By constructing and maintaining the user's personalized cognitive graph, the system can simulate the human brain's thought process, associating current interactive behaviors with historical behavior sequences and long-term user profiles for reasoning. This not only resolves the issue of intent ambiguity in multi-turn dialogues (e.g., distinguishing whether a user is switching songs or adjusting volume), but also enables dynamic transfer of intent based on context through a gating mechanism. Furthermore, for scenarios involving multimodal semantic conflicts (e.g., saying one thing but meaning another), the system can automatically resolve conflicts based on historical behavioral habits, assigning differentiated decision weights to different modalities. This makes the machine's decision-making logic closer to human cognitive habits, significantly improving the naturalness and intelligence of the interaction. By introducing a temperature scaling algorithm for confidence calibration at the decision output, it effectively overcomes the inherent "overconfidence" defect of deep learning models. The system no longer blindly outputs high-probability predictions but instead provides quantified uncertainty scores, making applications in safety-critical fields (e.g., assisted driving, telemedicine) possible—the system can intelligently choose "direct execution," "question for confirmation," or "remain silent" based on the confidence threshold. Further, through an online incremental learning mechanism, the model can self-correct and fine-tune using low-confidence samples, adapting to users' personalized expression habits (e.g., dialects, specific gestures) without full retraining. This closed-loop learning mechanism ensures that the system's performance continues to improve throughout its lifecycle, truly realizing the evolution from a "general model" to a "personalized steward".
[0024] This embodiment uses a smart cockpit scenario as an example to illustrate in detail the user behavior intent recognition method based on multimodal data fusion and dynamic cognitive evolution proposed in this invention. In this scenario, drivers may have various intents during driving, such as "adjusting the air conditioning," "changing songs," "answering a phone call," or "driving while fatigued." The system needs to accurately identify these intents to ensure driving safety.
[0025] Step S1: Multimodal data acquisition and preprocessing The system collects raw data streams through multi-source sensors deployed inside the vehicle. Specifically, this includes: Visual modal acquisition: A near-infrared camera is used to acquire a video stream of the driver's face at a frame rate of 30fps. The preprocessing steps include: using a multi-task cascaded convolutional neural network for face detection and alignment, extracting 68 facial key points, and cropping and normalizing the face region to 112×112 pixels based on affine transformation, with pixel values normalized to the [0,1] range. To further enhance robustness, random brightness and contrast perturbations are applied to the image.
[0026] Auditory modal acquisition: An in-vehicle microphone array was used to acquire in-vehicle audio signals at a sampling rate of 16kHz. The preprocessing steps include: removing silent segments using a WebRTC-based speech activity detection algorithm, performing pre-emphasis processing (with a coefficient typically of 0.97) to improve high-frequency resolution, using a Hamming window for frame segmentation (frame length 25ms, frame shift 10ms), and generating a time-frequency map with a dimension of 257 × number of frames through short-time Fourier transform.
[0027] Text and environmental modal acquisition: Audio signals are converted into text sequences via an in-vehicle automatic speech recognition (ASR) module. Simultaneously, the environmental context is read via the Controller Area Network (CAN) bus. Including vehicle speed (km / h), outside temperature (°C), steering wheel angle, and GPS location information.
[0028] Step S2: Heterogeneous Feature Extraction The system uses a deep neural network to extract features from the preprocessed data of each modality, generating fixed-dimensional feature vectors: Visual feature extraction: The normalized facial image sequence is input into a 3D ResNet-18 network. The input dimension is... (Number of channels × Number of time frames × Height × Width). The output of the global average pooling layer is extracted to obtain a 512-dimensional visual feature vector. This vector encodes the driver's facial expressions, gaze direction, and head posture information.
[0029] Auditory feature extraction: The time-frequency map is input into the VGGish network (pre-trained on a large-scale audio dataset), and the activation values of the fully connected layer FC1 are extracted to obtain a 128-dimensional auditory feature vector. This vector contains timbre, pitch, and acoustic environment characteristics.
[0030] Text feature extraction: extracting ASR text sequences The input is fed into the BERT-Base model, and the final hidden state corresponding to its [CLS] special label is extracted. After passing through a linear projection layer, a 768-dimensional text semantic feature vector is obtained. .
[0031] Environmental feature extraction: Scalars such as vehicle speed and temperature are Z-score standardized and then input into a two-layer fully connected network (the first layer has dimensions from the input dimension to 32 and uses ReLU as the activation function; the second layer has dimensions from 32 to 64), outputting a 64-dimensional environmental feature vector. .
[0032] Step S3: Cross-modal spatiotemporal alignment Since the sampling rates and acquisition starting points of data differ across modalities, the system constructs a cross-modal attention encoder for spatiotemporal alignment: Time alignment: Using the timestamp of the text modality as the reference anchor point. The Dynamic Time Warping (DTW) algorithm is used to calculate the optimal matching path between the audio frame sequence and the text word sequence. For example, the time period corresponding to the text word "open" is... The DTW algorithm automatically finds the most matching acoustic feature segment in the audio sequence, thereby solving the problem of audio-visual desynchronization.
[0033] Spatial Alignment and Feature Enhancement: Constructing a Multi-head Cross-Modal Attention Mechanism. Setting Text Features. For Query, visual features and auditory characteristics spliced matrix For Key and Value. Through linear transformation. (in , , (For the learnable parameter matrix), calculate the attention weight matrix. This weight matrix quantifies the relevance of each word in the text to each segment of the visual / auditory features. The final output is an aligned joint feature sequence. This sequence integrates visual and auditory evidence supporting the semantics of the text, with the following dimensions: .
[0034] Step S4: Deep Semantic Fusion Based on Graph Neural Networks System construction modal relationship diagram To achieve deep semantic fusion: Graph construction: Defining node sets These correspond to visual feature nodes, auditory feature nodes, text feature nodes, and environmental feature nodes, respectively. An adjacency matrix is defined. This is used to represent the strength of physical connections or semantic associations between nodes. Initially, diagonal elements are set to 1 (self-connection), and off-diagonal elements are set according to expert experience (e.g., visual and auditory senses belong to the same perceptual modality, so the initial weight is set to 0.8).
[0035] Graph convolution operation: Message passing is performed using a Graph Attention Network (GAT). For nodes... Its neighbor node set is First, calculate the attention coefficient. :
[0036] in It is a node eigenvectors, It is a learnable weight matrix. It is a weight vector. This indicates a concatenation operation. The coefficients are then normalized to obtain the weights. .node The updated features are represented as follows:
[0037] in This is the ELU activation function.
[0038] Feature aggregation: After two layers of GAT convolution, the system concatenates the updated feature vectors from the four nodes and generates a global deep fusion feature vector through a global average pooling layer. This vector not only contains independent information about each modality, but also explicitly models the complex interaction relationships between modalities.
[0039] Step S5: Intention Reasoning Based on Dynamic Cognitive Evolution The system builds and maintains users' personalized dynamic cognitive maps. Reasoning: Cognitive Graph Initialization: The system establishes an independent cognitive graph for each user. The graph contains entity nodes (such as "air conditioner", "window", "music") and intent nodes (such as "cool down" and "entertainment"). Edge weights represent the strength of the association between the entity and the intent (e.g., "high temperature" has a higher weight when connected to "cool down").
[0040] Memory retrieval and state update: Deeply fusing feature vectors The observation at the current moment is input to the gated recurrent unit (GRU). The GRU updates the gates... and reset door Controlling the flow of information:
[0041]
[0042]
[0043]
[0044] in This represents the hidden state of the cognitive graph from the previous moment. This represents the updated state at the current moment. This process simulates the mechanism by which humans use short-term memory to understand current intentions.
[0045] Intent Classification: Update the hidden state The input is fed into a fully connected classification layer, followed by a Softmax function to calculate the probability distribution of the intent category. The preset intent category set Y = {0: turn down the air conditioner, 1: open the window, 2: change the song, 3: answer the phone, 4: other}.
[0046] Step S6: Confidence Calibration and Decision Output To prevent the model from becoming overconfident, the system calibrates the output probabilities: Temperature scaling: Introducing temperature parameters (For example The log-odds ratio of the raw output of the classification layer. After softening, the calibrated probability distribution is as follows: .
[0047] Hierarchical decision-making mechanism: Execute directly: If The system determines that the confidence level is extremely high and directly sends instructions to the vehicle domain controller, such as executing "lower the air conditioning temperature by 2 degrees".
[0048] Ask a question to confirm: If The system determined that there was some uncertainty and asked a follow-up question through the car's audio system: "Do you want to lower the air conditioning temperature?", waiting for the user's second confirmation.
[0049] Silence or prompt: If If the system determines that the input is unclear or invalid, it will remain silent or display the message "Sorry, I did not hear you, please repeat."
[0050] Step S7: Online Incremental Learning and Conflict Resolution The system possesses closed-loop learning capabilities and can handle multimodal conflicts. Conflict resolution example: Suppose the ASR-recognized text is "Don't turn it on" (intent leaning towards turning it off), but visual features show the driver frequently wiping away sweat (visual leaning towards heat). The system retrieves a dynamic cognitive map. If historical data shows that the user has a history of "saying one thing but meaning another" behavior (e.g., actually performing cooling actions in three similar scenarios in the past), the system will automatically increase the attention weight of the visual modality in GAT. Reduce text modal weights This corrects the misleading nature of a single modality.
[0051] Online fine-tuning: When a user gives a negative answer to a question (such as saying "no") or the system misjudges, the sample ( Samples marked as difficult are stored in the replay buffer. Once the buffer accumulates a predetermined number of samples (e.g., 100), the system uses a contrastive loss function to perform mini-batch gradient descent updates on the network parameters from steps S3 to S5 in the background. This online incremental learning mechanism allows the model to continuously adapt to the driver's personalized expression habits without requiring full retraining.
[0052] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A user behavior intent recognition method based on multimodal data fusion, characterized in that: Includes the following steps: Step S1: Multimodal data acquisition and preprocessing. Raw multimodal data streams from the user's interaction process are acquired in real time using multi-source sensors deployed on the terminal; the raw multimodal data streams include at least: visual modal data, auditory modal data, text modal data, and environmental context data; the raw multimodal data streams are denoised, normalized, and segmented to generate standardized modal sequences; Step S2: Heterogeneous Feature Extraction. Deep neural networks are used to extract features from each standardized modality sequence: a 3D convolutional neural network is used to extract the spatiotemporal features of the visual modality. Using convolutional recurrent neural networks to extract time-frequency features of auditory modalities Extracting semantic features of text modalities using pre-trained language models Extracting environmental context features using fully connected networks ; Step S3: Cross-modal spatiotemporal alignment. Construct a cross-modal attention encoder, using the timestamps of the text modality as anchor points, to calculate the features of the visual modality. and auditory modal features Compared to text modal features The cross-modal attention weight matrix; based on the attention weight matrix, using bilinear interpolation or dynamic time warping algorithms, the cross-modal attention weight matrix is... , , , Mapping to a unified time axis and feature dimension generates a spatiotemporally aligned joint feature sequence. ; Step S4: Deep semantic fusion based on graph neural networks. Construct a modality graph. , where the node set Edge sets represent the features of each modality. Represents the potential association between modes; Injection Node Set In this process, graph convolutional networks or graph attention networks are used to perform message passing and aggregation on the modality relationship graph, capturing the complementarity and nonlinear dependencies between modalities, and outputting a global deep fusion feature vector. ; Step S5: Intention reasoning based on dynamic cognitive evolution. Construct and maintain the user's personalized dynamic cognitive graph. The Record the user's historical intent patterns and entity relationships; As the current observation vector input to In the update function, the state of entity nodes in the graph is updated using a gating mechanism, and the probability distribution of the intent category at the current time is calculated based on the updated graph state. ; Step S6: Confidence calibration and decision output. The temperature scaling algorithm is used to adjust the probability distribution of the intent category. The log odds are calibrated to generate a calibrated probability distribution. According to the above The maximum value and its corresponding entropy value are used to set a decision threshold, and the final intent recognition result and corresponding confidence score are output.
2. The user behavior intent recognition method and system based on multimodal data fusion according to claim 1, characterized in that: The cross-modal attention encoder in step S3 includes a query, key, and value mapping submodule; wherein, text modal features are selected. As a query, select visual modal features. and auditory modal features As Key and Value; through calculation and , The dot product similarity is used to generate an attention mask, which is used to weightedly fuse visual and auditory information to provide visual and acoustic evidence to support the semantics of the text.
3. The user behavior intent recognition method based on multimodal data fusion according to claim 1, characterized in that: The construction of the modal relationship diagram G in step S4 specifically includes: Define the node feature matrix ,in For the number of modes (here) ), For feature dimensions; Define an adaptive adjacency matrix Initialization is based on expert knowledge settings, and end-to-end optimization and updates are performed during model training using gradient descent. Through graph convolution operation The update node indicates that, To add self-connected adjacency matrices, For degree matrix, For learnable weights, This is the activation function.
4. The user behavior intent recognition method based on multimodal data fusion according to claim 1, characterized in that: The dynamic cognitive evolution module in step S5 specifically includes: Memory retrieval unit: responds to the current deep fusion feature vector Retrieve the Top-K historical behavior sequence fragments most similar to this feature from the historical database; The graph evolution unit: It takes the hidden states of the Top-K historical segments as prior knowledge and inputs them into the gated recurrent unit or long short-term memory network, combining them with the current... Update cognitive map Connection weights between nodes; Conflict resolution subunit: When semantic conflicts exist in multimodal data, it is based on cognitive graphs. The system records the user's historical behavioral habits and automatically adjusts the voting weight of each modality in the final decision. For example, if historical data shows that the user tends to speak faster (auditory) but has no facial expression (visual) when nervous, the weight coefficient of the auditory modality will be increased in nervous contexts.
5. The user behavior intent recognition method based on multimodal data fusion according to claim 1, characterized in that: Following step S6, an online incremental learning step S7 is also included: Monitor the calibrated probability distribution The entropy value or maximum probability value; When the entropy value is higher than the preset first threshold or the maximum probability value is lower than the preset second threshold, the current intent recognition result is determined to be in a fuzzy boundary. Trigger the proactive inquiry mechanism to obtain explicit feedback tags from users, or store the current sample in the cache; When the number of samples in the buffer reaches the batch update threshold, the neural network parameters in steps S3 to S5 are updated by backpropagation in small batches using the contrastive learning loss function or the cross-entropy loss function, so as to realize the online fine-tuning and continuous evolution of the model.
6. The user behavior intent recognition method based on multimodal data fusion according to claim 1, characterized in that: The environmental context data includes at least: geographic location information, timestamp, light intensity, and device operating status parameters; in the graph neural network fusion process in step S4, the environmental context data participates in the message passing process as a bias term or constraint condition to correct the perception bias caused by environmental factors (for example, automatically suppressing the transmission weight of the auditory modality and enhancing the transmission weight of the visual modality in a strong noise environment).
7. A user behavior intent recognition system based on multimodal data fusion, characterized in that: include: Data acquisition layer: configured with a visual sensor, an auditory sensor, a text input interface and an environmental perception module, used to execute step S1 as described in claim 1; Feature engineering layer: integrates a visual feature extractor, an auditory feature extractor, a text embedding engine, and an environmental feature encoder, for performing step S2 as described in claim 1; Fusion Inference Layer: Includes: Spatiotemporal alignment module: Built-in cross-modal attention encoder, used to perform the cross-modal spatiotemporal alignment operation as described in claim 2; Graph fusion module: Built-in graph neural network processor, used to perform the deep semantic fusion operation as described in claim 3; Cognitive Evolution Module: It has a built-in dynamic cognitive graph database and evolution controller for performing the intention reasoning operation as described in claim 4, and includes a memory bank for storing user-personalized difference parameters; Decision output layer: It is equipped with a confidence calibrator and a decision executor to execute step S6 as described in claim 1, and to decide whether to directly execute the instruction, initiate a question or remain silent based on the confidence score; Online learning layer: Coupled between the fusion inference layer and the decision output layer, configured to perform the online incremental learning steps as described in claim 5.
8. The user behavior intent recognition system based on multimodal data fusion according to claim 7, characterized in that: The conflict resolution subunit in the cognitive evolution module (330) is further configured as follows: When the system detects a conflict between the intentions of the nonverbal modalities (visual / environmental) and the verbal modalities (text / auditory), it invokes the user psychological characteristic model stored in the dynamic cognitive graph. If the model shows that the user has an "introverted" personality, the nonverbal modalities are given higher decision weights; if the user has an "extroverted" personality, the verbal modalities are given higher decision weights.