User behavior prediction system and method based on multi-modal data fusion

The user behavior prediction system, which integrates multimodal data fusion, collects and processes various heterogeneous data. By utilizing attention mechanisms and time-series prediction models, it solves the problem of insufficient accuracy in user behavior prediction in traditional methods and achieves more accurate behavior prediction results.

CN120832498AInactive Publication Date: 2025-10-24BEIJING DATA100 INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511331644.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-10-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional single-modal data-driven user behavior prediction methods struggle to capture the underlying motivations behind user behavior and fail to fully characterize the complex relationships between multimodal information, resulting in insufficient prediction accuracy.

Method used

A multimodal data acquisition module is used to collect visual, auditory, textual, physiological signals and environmental context data in real time. High-dimensional features are extracted through the preprocessing and feature extraction modules. Feature fusion is performed through the cross-modal fusion module through the attention mechanism. The time series prediction model of Transformer and bidirectional LSTM is combined to output the probability distribution of behavioral intention. Finally, end-to-end adaptive learning is performed through the model optimization and feedback module.

Benefits of technology

It achieves more robust user behavior prediction, improves prediction accuracy, and can accurately capture the evolution of behavior and output reliable probability distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832498A_ABST
    Figure CN120832498A_ABST
Patent Text Reader

Abstract

The invention discloses a user behavior prediction system and method based on multi-modal data fusion, and particularly relates to the field of user behavior prediction, and the system comprises a multi-modal data collection module, a preprocessing and feature extraction module, a cross-modal fusion module, a user behavior prediction module, and a model optimization and feedback module. According to the system, multi-dimensional original data such as visual sense, auditory sense, text, physiological signals and environment context of a user are acquired in real time through a multi-modal data acquisition module; then, deep networks such as ResNet, VGGish and BERT are adopted to extract high-dimensional feature vectors of all modals, and contribution weights of features of different modals are dynamically learned through an attention mechanism; and finally, based on a time sequence model of Transform and LSTM, analyzing fusion features, and outputting probability distribution of future behavior intentions. And parameter joint optimization and continuous learning are realized through a multi-objective loss function and end-to-end back propagation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of user behavior prediction, more particularly, the present application relates to a user behavior prediction system and method based on multi-modal data fusion. BACKGROUND

[0002] With the rapid development of mobile internet, internet of things and intelligent terminal technology, the data generated by users in their daily production and life shows explosive growth and multi-modal characteristics. User behavior prediction, as a core technology for understanding user intent and optimizing service experience, has been widely used in personalized recommendation, intelligent interaction, risk warning and other fields - for example, e-commerce platforms predict user purchase behavior to achieve accurate product push, smart homes predict user operation habits to optimize device response logic, and public safety systems analyze crowd behavior patterns to warn potential risks. However, the complexity of user behavior and the diversity of data forms make traditional single-modal data-driven prediction methods face serious challenges.

[0003] Traditional user behavior prediction methods often rely on a single type of data for modeling, such as collaborative filtering algorithms based on user historical click sequences, sentiment prediction models based on text reviews, or preference analysis methods based on image interaction records; Such methods have significant limitations: on the one hand, single-modal data can only reflect the local characteristics of user behavior, making it difficult to capture the underlying motivations behind the behavior, on the other hand, user behavior is often the result of multiple modal information driving together, and single-modal data cannot fully capture the complex association of such user behavior. SUMMARY

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present application provide x to solve the problems raised in the background art.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions: A multi-modal data acquisition module for acquiring raw data of a user in real time from multiple heterogeneous data sources; A preprocessing and feature extraction module for performing data cleaning and standardization processing on the raw data of each modality, and extracting high-dimensional feature vectors using deep neural network models corresponding to the modality characteristics, forming visual features, auditory features, text features, physiological features and context features; A cross-modal fusion module for receiving multi-modal feature vectors and performing fusion through a feature fusion network based on an attention mechanism, the fusion network being trained to dynamically learn the contribution weights of different modal features to behavior prediction in a specific context, and outputting a fused feature representation; a user behavior prediction module configured to input the fused feature representation into a time-series prediction model, and output a future behavior intention probability distribution of the user, the behavior intention probability distribution including one or more predicted behaviors and corresponding occurrence probabilities thereof; a model optimization and feedback module configured to receive a real behavior feedback of the user, calculate a loss between a prediction result and a real result, and jointly optimize model parameters in the cross-modal fusion module and the user behavior prediction module through a back propagation algorithm based on the loss, to realize end-to-end adaptive learning.

[0006] Preferably, in the multi-modal data acquisition module, multi-dimensional original information of the user is dynamically captured by integrating heterogeneous data sources distributed in different scenes; the original data includes visual modal data, auditory modal data, text modal data, physiological signal modal data, and environmental context modal data; the visual modal data covers multi-dimensional visual information, including dynamic interactive images such as facial micro-expressions, body movement trajectories, and gaze focus areas of the user captured in real time through a terminal device camera, static visual content such as pictures, video frames, and avatars published by the user on a social platform, and three-dimensional action sequences and scene interaction records generated in a virtual reality environment; the auditory modal data includes active audio information such as instruction voice, emotional vocalization, and conversation content generated by the user through voice interaction, passive audio signals such as background noise and device operation sound in the environment, and paralanguage features such as intonation, speech rate, and pauses in the voice interaction process.

[0007] Preferably, in the preprocessing and feature extraction module, multi-modal original data is converted into structured features that can be directly used for fusion analysis through a targeted processing procedure; first, fine-grained preprocessing is performed on each modal data: for visual modal data, data is purified through image denoising, illumination normalization, and object detection, and then the image is uniformly scaled to a preset size through scale normalization; for auditory modal data, audio noise reduction, endpoint detection, and sampling rate unification are performed first, and then the data is converted into a mel-spectrogram for subsequent processing; after text modal data is subjected to word segmentation, stop word removal, and special symbol cleaning, stem extraction technology is used to standardize the word form, and word vector mapping is performed to achieve preliminary numerical value; Physiological signal modal data needs to be subjected to baseline drift correction, waveform rejection, and signal resampling to ensure time-series consistency; Environmental context modal data is preprocessed through format standardization, missing value filling, and discrete feature encoding; In the feature extraction stage, the module matches a dedicated deep neural network model according to the information characteristics of different modalities: the visual feature extraction adopts an improved ResNet-50 architecture, captures hierarchical features from local textures to global contours through multi-scale convolutional layers, and finally outputs a 2048-dimensional visual feature vector, while introducing an attention mechanism to strengthen the feature weights of key areas such as user expressions and actions.

[0008] Preferably, in the cross-modal fusion module, the visual, auditory, textual, physiological, and environmental context heterogeneous feature vectors are converted into unified fusion feature representations, the correlation weights of different modalities in a specific scene are dynamically modeled through a fusion network based on an attention mechanism, and multi-dimensional information is complementarily enhanced; first, the input modal feature vectors are dimensionally aligned and spatially mapped, the visual, auditory, textual, physiological, and context features are uniformly mapped to a shared feature space through linear transformation, and the calculation method is specifically: wherein, represents a learnable linear projection matrix designed for modality m, represents the bias term of modality m, is the mapped feature vector.

[0009] Preferably, in the user behavior prediction module, the unified feature representation generated by the cross-modal fusion module is converted into a user behavior intention prediction result, based on the multi-dimensional user information contained in the fusion feature, the time sequence evolution law of user behavior is captured through a time series prediction model, and the behavior intention probability distribution in the future period is output; First, the input fusion feature is expanded and modeled in the time sequence dimension, and an improved time series prediction model architecture is adopted for the sequence dependency of user behavior (such as the behavior chain of browsing-clicking-purchasing), which is based on a Transformer encoder as the basic skeleton, combined with a bidirectional LSTM network to capture the long and short term dependency of the behavior sequence, forming a double-layer prediction mechanism of “local time sequence modeling + global semantic association”; The fusion features are sorted according to the time stamp to construct the time sequence feature sequence of user behavior wherein, is the length of the historical window, is the current time, and the time sequence information is injected through position encoding to ensure that the model perceives the time sequence order of behavior occurrence; the bidirectional LSTM layer serves as a bottom-level time sequence encoder, which scans the sequence from both forward and backward directions to generate a hidden state sequence containing local time sequence patterns, capturing the dynamic change trend of recent user behavior; wherein, the forward direction is from the past to the present, and the backward direction is from the present to the past; The upper-layer Transformer encoder models the global correlation between features at any time in the sequence through a multi-head self-attention mechanism, for example, identifying the potential correlation between "browsing behavior a week ago" and "current adding-to-cart behavior", and the attention weight calculation incorporates a behavior interval decay factor to give a dynamically reduced correlation weight to behaviors with a long time interval. The calculation method is as follows: wherein, is represented as an attention weight value, is represented as a behavior feature vector at the nth time, is represented as a behavior feature vector at the qth time, is represented as a behavior feature vector at the kth time, is represented as a behavior interval decay coefficient, is represented as a similarity calculation function. This makes the model pay attention to the strong influence of recent behaviors and not ignore the cumulative effect of long-term behaviors.

[0010] Preferably, in the model optimization and feedback module, the model parameters of the cross-modal fusion module and the user behavior prediction module are dynamically adjusted by continuously receiving the deviation signal of the user real behavior data and the prediction result, to realize the end-to-end performance evolution. First, a multi-dimensional loss calculation system is constructed, and a hybrid loss function is designed according to the discrete and sequential characteristics of user behavior prediction: the cross-entropy loss is used to quantify the deviation between the predicted behavior intention probability distribution and the real behavior label, and the calculation method is as follows: wherein, is represented as a cross-entropy loss, is represented as a one-hot encoding label of a real behavior, is represented as a prediction probability; The timing consistency loss is introduced to constrain the rationality of the behavior sequence, and the transition probability of the predicted result and the real behavior sequence at consecutive times is compared to punish the prediction with a large jump (such as an abnormal prediction that jumps from "browsing" to "purchase" directly without the intermediate "adding-to-cart" link), and the calculation method is as follows: wherein, is represented as a timing consistency loss value, is a behavior transition probability matrix predicted at the tth time, is a real transition matrix; The modal feature reconstruction loss is also added, the original modal feature is inferred from the fused feature through the autoencoder, and the reconstruction error is calculated, and the calculation method is as follows: wherein, is expressed as a modal feature reconstruction loss value, is the reconstructed modal feature, is the original feature; The final total loss function is wherein, is expressed as a loss weight, which is dynamically adjusted according to the scene.

[0011] Technical effects and advantages of the present application: The present application acquires visual, auditory, text, physiological signal and environmental context data in real time from heterogeneous data sources through the acquisition module; after cleaning and standardization, the pre-processing and feature extraction module extracts corresponding high-dimensional features using ResNet-50, VGGish+LSTM, BERT and other models; the cross-modal fusion module dynamically learns modal weights through the attention mechanism and outputs unified fusion features; the prediction module outputs behavior intention probability distribution using the Transformer+bidirectional LSTM model; the optimization module combines multiple loss functions such as cross-entropy to jointly optimize parameters through hierarchical feedback and back propagation, realizing end-to-end adaptive learning. The present application integrates multi-dimensional data such as visual, auditory, text, physiological signal and environmental context, constructs an end-to-end feature fusion and prediction framework, fully excavates the complementary information of multi-modal data, realizes more robust user behavior prediction, and greatly improves the accuracy of user behavior prediction; the time series prediction model combining Transformer and bidirectional LSTM accurately captures the behavior evolution law and outputs reliable probability distribution. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a schematic diagram of the module connection of the present application.

[0013] Figure 2 is a schematic diagram of the method flow of the present application.

[0014] Figure 3 is a schematic diagram of the behavior prediction and model optimization flow of the present application. DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0016] Please refer to Figure 1As shown, the present application provides a user behavior prediction system based on multi-modal data fusion, including a multi-modal data acquisition module, a preprocessing and feature extraction module, a cross-modal fusion module, a user behavior prediction module, a model optimization and feedback module.

[0017] The multi-modal data acquisition module is used to acquire the original data of the user in real time from multiple heterogeneous data sources, and the original data includes visual modal data, auditory modal data, text modal data, physiological signal modal data and environmental context modal data. In the multi-modal data acquisition module, the dynamic capture of multi-dimensional original information of the user is realized by integrating heterogeneous data sources distributed in different scenes. The original data includes visual modal data, auditory modal data, text modal data, physiological signal modal data and environmental context modal data. The visual modal data covers multi-dimensional visual information, including dynamic interactive images such as user facial micro-expression, body movement trajectory and gaze focus area acquired in real time through the terminal device camera, static visual content such as pictures, video frames and avatars published by the user on the social platform, and three-dimensional action sequences and scene interaction records generated in the virtual reality environment. The auditory modal data includes active audio information such as instruction voice, emotional voice and dialogue content generated by the user through voice interaction, passive audio signals such as background noise and device running sound in the environment, and paralinguistic features such as intonation, speech rate and pause in the voice interaction process. The text modal data covers two types of explicit text and implicit text, the former includes direct text content such as user input comments, messages, search words and chat records, and the latter includes indirect text information such as document editing history, form filling traces and code writing process. The physiological signal modal data is acquired through wearable devices and special sensors, covering physiological indicators such as heart rate, blood pressure, skin electrical response, brain wave fluctuation and respiratory rate. These data can reflect the user's emotional state, attention concentration and physiological load in real time. The environmental context modal data includes spatial and temporal dimension information (such as real-time geographic location, place type, indoor and outdoor environment, time node), device state information (such as terminal model, operating system, network type, power level) and social relationship context (such as current interaction object, social circle dynamics) and other scenario data.

[0018] The preprocessing and feature extraction module is used to perform data cleaning and standardization processing on the original data of each modal respectively, and extract high-dimensional feature vectors using deep neural network models corresponding to the modal characteristics to form visual features, auditory features, text features, physiological features and context features. In the preprocessing and feature extraction module, multimodal raw data is converted into structured features that can be directly used for fusion analysis through targeted processing processes. First, refined preprocessing is performed on each modal data: for visual modal data, the data is purified through image denoising, illumination normalization, and target detection operations, and then the image is uniformly scaled to a preset size through scale normalization. For auditory modal data, audio noise reduction, endpoint detection and sampling rate unification are first performed, and then converted into Mel-spectrograms for subsequent processing. Text modal data is segmented, stop words are removed, and special symbols are cleaned. The word form is standardized using stemming technology, and preliminary digitization is achieved through word vector mapping. Physiological signal modal data requires baseline drift correction, waveform removal, and signal resampling to ensure timing consistency; Environmental context modal data is preprocessed through format standardization, missing value filling and discrete feature encoding; During the feature extraction phase, the module matches a dedicated deep neural network model to the information characteristics of different modalities. Visual feature extraction uses an improved ResNet-50 architecture, capturing hierarchical features from local texture to global contour through multi-scale convolutional layers, ultimately outputting a 2048-dimensional visual feature vector. At the same time, an attention mechanism is introduced to strengthen the feature weights of key areas such as user expressions and movements. Perform convolution calculations on local areas of the input image to extract low-level features such as edges and textures: in, Represents the specific position value of the output feature map of the lth layer, For the The input feature map of the layer, Expressed as The convolution kernel of the layer, is represented as a bias term, Expressed as activation function, Represented as the output feature map of the lth layer; The calculation method of residual connection is specifically as follows: in, Represented as the l-th layer output feature map, It is represented as a residual function, which preserves the underlying features by directly connecting the input and output; For the The input feature map of the layer, Expressed as The convolution kernel of the layer.

[0019] After multiple layers of convolution, the feature map is converted into a fixed-length vector through global average pooling: Final output 2048-dimensional visual feature vector Auditory feature extraction is based on the pre-trained VGGish model, which extracts audio spectral features through stacking convolution and pooling layers based on mel spectrogram, and combines LSTM network to capture timing features such as intonation and speech rate, generating 1280-dimensional auditory feature vector; In the mel spectrum conversion formula, the spectrum of the audio signal is mapped to the mel scale, and the calculation method is as follows: Where, is the frequency, and the mel spectrogram is obtained by the mel filter bank is the time frame, is the frequency dimension; Then extract the spectral features through multi-layer convolution, and the calculation method is as follows: Where, is the mel spectrum of the th frame, is the CNN feature of the current frame In the LSTM timing modeling, the dynamic characteristics of the audio sequence are captured: Where , and respectively represent the input gate, the forgetting gate, and the output gate, is the cell state, is the hidden state of the t frame, is the hidden state of the t-1 frame, , and represent the input feature weight matrix, , and represent the historical state weight matrix, , and represent the bias term; Finally, the hidden state of the last frame is taken as the auditory feature vector: is the auditory feature vector; ​​​​​​​​​​​​The text feature extraction adopts a BERT pre-training model, which understands the context semantics through a bidirectional Transformer structure, takes the hidden layer output corresponding to the [CLS] mark as a 768-dimensional text feature vector, and supports dynamic adjustment of the segmentation granularity to adapt to different text lengths. In the Token embedding formula, the embedding of each word is composed of word embedding, paragraph embedding and position embedding, and the calculation method is as follows: wherein, represents the final embedding vector of the a-th word, is the a-th word, is the paragraph mark, represents the word embedding matrix, represents the paragraph embedding matrix, represents the position embedding matrix.

[0020] Transformer self-attention calculation: capture the dependency between words through multi-head attention, and output multi-head attention combination through concatenating multiple single-head attention outputs; The text feature vector output takes the final hidden layer output of the [CLS] mark in the BERT input as the text feature: represents the text feature vector, represents the hidden state vector corresponding to the [CLS] position in the last layer; The physiological feature extraction constructs a multi-branch CNN-LSTM network, wherein the convolution branch extracts local features such as heart rate variability, the LSTM branch captures long-term dependencies of physiological signals, and the fusion forms a 512-dimensional physiological feature vector; The 1D convolution is performed on the physiological signal segment, and the LSTM long-term dependency modeling is consistent with the LSTM method of the auditory feature, the local feature sequence output by the CNN is input into the LSTM, and the hidden state at the last time is taken as the physiological feature vector; The environmental context feature is modeled by a graph neural network to model the spatio-temporal correlation, and maps the multi-dimensional context information into a 256-dimensional feature vector while preserving the interpretability of key attributes; The environmental context is modeled by a graph neural network to model the correlation between elements. In the graph structure definition, the node represents the context element (such as "location=office" and "time=14:00"), the edge represents the correlation between elements (such as the co-occurrence relationship between time and location), and the adjacency matrix represents the correlation strength; the feature of each node is updated by aggregating the neighbor node features, and finally the final features of all nodes are globally pooled to obtain the context feature vector.

[0021] The cross-modal fusion module is configured to receive the multi-modal feature vectors and fuse the multi-modal feature vectors through an attention mechanism-based feature fusion network. The fusion network is trained to dynamically learn contribution weights of different modal features to behavior prediction in a specific context environment and output a fused feature representation. In the cross-modal fusion module, the visual, auditory, text, physiological, and environmental context heterogeneous feature vectors are converted into a unified fused feature representation. The correlation weights of different modalities in a specific scenario are dynamically modeled through an attention mechanism-based fusion network, and multi-dimensional information is complementarily enhanced. First, the input modal feature vectors are dimensionally aligned and spatially mapped. The visual, auditory, text, physiological, and context features are uniformly mapped to a shared feature space through linear transformation. The calculation method is specifically as follows: wherein, is a learnable linear projection matrix designed for modality m, is a bias term for modality m, is the mapped feature vector; In the fusion network, the intra-modal attention layer strengthens the key information within the single-modal feature through a self-attention mechanism. The features of the user's expression region in the visual feature are given higher weights, and the features of the sentiment tendency words in the text feature are enhanced. The calculation formula is as follows: , is the mth modal feature vector enhanced by the self-attention mechanism, wherein the self-attention realizes importance weighting by calculating the correlation degree of the internal elements of the feature vector; the cross-modal attention layer as a core component dynamically learns the contribution weights of different modalities to the current behavior prediction task. It calculates the similarity scores of each modality feature and the query vector by introducing a learnable context query vector wherein, is the similarity score, is a new feature vector enhanced by the self-attention mechanism, and q is a context query vector; After softmax normalization, the modality attention weight is obtained. The calculation method is specifically as follows:

[0022] is the mth modality attention weight, is the similarity score, the sum of the exponential similarity of all modalities is calculated; ​The weight values are dynamically adjusted according to the scene, for example, in the user emotion prediction scene, the weight of the physiological signal and the facial expression feature is significantly higher than that of the environmental context feature; the feature aggregation layer weights and fuses each modality feature based on the attention weight, and performs nonlinear transformation and dimension compression through a multilayer perception machine, and finally generates a fixed-dimension fusion feature vector: wherein, is represented as a joint representation vector after multi-modal fusion, is represented as the mth modality attention weight, is represented as the mth modality feature vector after self-attention enhancement.

[0023] The user behavior prediction module is used to input the fusion feature representation into a time series prediction model, and output a future behavior intention probability distribution of the user, wherein the behavior intention probability distribution includes one or more predicted behaviors and their corresponding occurrence probabilities. In the user behavior prediction module, the unified feature representation generated by the cross-modal fusion module is converted into a user behavior intention prediction result, based on the multi-dimensional user information contained in the fusion feature, the time series evolution law of the user behavior is captured through a time series prediction model, and a behavior intention probability distribution in a future period of time is output. First, the input fusion feature is expanded and modeled in the time series dimension, and an improved time series prediction model architecture is used to capture the sequence dependence of the user behavior (such as the behavior chain of browsing-clicking-purchasing), which is based on a Transformer encoder as the basic skeleton, combined with a bidirectional LSTM network to capture the long-short term dependence of the behavior sequence, forming a double-layer prediction mechanism of “local time series modeling + global semantic association”; The fusion feature is sorted according to the timestamp to construct a time series feature sequence of the user behavior, wherein, is the length of the historical window, is the current time, and the time sequence information is injected through position encoding to ensure that the model perceives the time sequence order of the behavior; the bidirectional LSTM layer is used as a bottom layer time series encoder, which scans the sequence from the forward direction and the reverse direction respectively to generate a hidden state sequence containing local time series patterns, which captures the dynamic change trend of the user's recent behavior; wherein, the forward direction is from the past to the present, and the reverse direction is from the present to the past.The upper layer Transformer encoder models the global correlation between features at any time in the sequence through a multi-head self-attention mechanism, for example, to identify the potential correlation between "browsing behavior a week ago" and "current adding to shopping cart behavior", and the attention weight calculation incorporates a behavior interval decay factor to give dynamic reduced correlation weight to behaviors with longer time intervals. The calculation method is as follows: wherein, is the attention weight value, is the behavior feature vector at the nth time, is the behavior feature vector at the qth time, is the behavior feature vector at the kth time, is the behavior interval decay coefficient, is a similarity calculation function. This allows the model to focus on the strong influence of recent behavior and not ignore the cumulative effect of long-term behavior; Taking the "user shopping behavior prediction" scenario as an example, assuming that the historical window K = 20 (covering the past 100 minutes, with a sampling step of 5 minutes), the current time n = t (corresponding to "t time user clicks the settlement button"), q = t-5 (corresponding to "t-25 minutes user browses product reviews"), k traverses all time points from t-20 to t, and λ = 0.03: Step A1: First, calculate (cosine similarity between (settlement time feature) and (browsing review time feature) (semantically closely related, as browsing reviews are often key behaviors before settlement); Step A2: Calculate the time interval , and the decay factor ; Step A3: The numerator is ; Step A4: Calculate the similar numerator for all k time points and sum them up (such as corresponding to "t-5 minutes adding to shopping cart", similarity 0.92, interval 1, decay factor 0.97, numerator ; Step A5: Finally , is the sum of the numerators for all k time points. If the total is 15.2, then , i.e. the attention weight of the browsing review behavior on the current settlement behavior prediction is about 13.6%, and , if the total is 15.2, then , i.e. the weight of the adding to shopping cart behavior is higher (such as 0.158), which is consistent with the actual logic that recent key behaviors have a stronger influence; After the time sequence feature coding is completed, the module compresses the variable-length time sequence feature sequence into a fixed-dimension global time sequence feature vector through an adaptive pooling layer , and inputs it into a prediction head composed of a multi-layer perception and a softmax classifier, wherein the multi-layer perception further mines high-order interactions between features through a nonlinear transformation, and the softmax classifier maps the transformed features to a preset behavior intention category space (such as “click”, “click”, “purchase”, “share”, “exit”, etc.), and finally outputs the occurrence probability of each behavior category to form a complete behavior intention probability distribution (wherein is the total number of behavior categories, and ); and a few-shot learning mechanism is used to adapt to new behavior types (such as the behavior intention corresponding to the newly added “collection” function of the platform), and a confidence threshold is introduced to filter low-probability prediction results.

[0024] The model optimization and feedback module is used to receive user real behavior feedback, calculate the loss between the prediction result and the real result, and use the loss to jointly optimize the model parameters in the cross-modal fusion module and the user behavior prediction module through a back propagation algorithm to realize end-to-end adaptive learning.

[0025] In the model optimization and feedback module, the model parameters of the cross-modal fusion module and the user behavior prediction module are dynamically adjusted by continuously receiving the deviation signal of the user real behavior data and the prediction result to realize end-to-end performance evolution; first, a multi-dimensional loss calculation system is constructed, and a mixed loss function is designed according to the discrete and sequential characteristics of user behavior prediction: the cross-entropy loss is used to quantify the deviation between the predicted behavior intention probability distribution and the real behavior label, and the calculation method is as follows: wherein, represents the cross-entropy loss, represents the one-hot encoding label of the real behavior, represents the prediction probability; The temporal consistency loss is introduced to constrain the rationality of the behavior sequence, and the transition probability of the continuous time prediction result and the real behavior sequence is compared to punish the prediction with large jumps (such as the abnormal prediction of directly jumping from “browsing” to “purchase” without the intermediate “add to cart” link), and the calculation method is as follows: wherein, represents the temporal consistency loss value, is the behavior transition probability matrix predicted at t, is the real transition matrix; The modal feature reconstruction loss is added at the same time, the original modal feature is deduced from the fusion feature through the auto-encoder, the reconstruction error is calculated, and the calculation method is specifically as follows: Among them, is expressed as a modal feature reconstruction loss value, is the reconstructed modal feature, is the original feature; The final total loss function is , wherein, is expressed as a loss weight, which is dynamically adjusted according to the scene; In the parameter optimization stage, an adaptive gradient descent algorithm (such as AdamW) is used to realize end-to-end back propagation: first, the total loss is calculated The gradient of the output layer parameter of the user behavior prediction module is returned to the parameters of the Transformer encoder and the bidirectional LSTM layer by layer through the chain rule; the gradient is propagated to the cross-modal fusion module to update the modal mapping matrix, the attention query vector and the nonlinear transformation parameter of the MLP, so that the weight distribution of the fusion feature is more suitable for the driving factors of the real behavior; for example, if the user's actual purchase behavior depends more on the image of the goods than on the text description, the attention weight of the visual modal after optimization is improved; In dynamic adaptive learning, a hierarchical feedback mechanism is designed, short-term feedback is directed to the instantaneous behavior deviation of a single user (such as predicting "click" but actually "exit"); online gradient descent is used to quickly fine-tune the prediction head parameters, ensuring an immediate response to user behavior mutations; medium-term feedback is based on batch user data within a sliding window, and the attention parameters of the cross-modal fusion network are updated through small batch gradient descent to optimize the modal weight distribution strategy; long-term feedback is to retrain the full amount of historical data at a fixed period, and the learning rate decay strategy is used to adjust the deep network parameters, realizing the long-term performance iteration of the model; at the same time, a performance monitoring unit is built in the module to calculate the prediction accuracy, F1 score, confusion matrix and other indicators in real time, and when the indicators continuously decrease by more than a threshold value, the parameter reset and incremental training process is automatically triggered, avoiding the model falling into local optimum.

[0026] Please refer to Figure 2 The application also provides a user behavior prediction method based on multi-modal data fusion, comprising the following steps: S1: Real-time acquisition of original data of users from multiple heterogeneous data sources; S2: The original data of each modal is respectively subjected to data cleaning and standardization processing, and a deep neural network model corresponding to the modal characteristics is used to extract a high-dimensional feature vector, forming visual features, auditory features, text features, physiological features and context features; S3: receive the multi-modal feature vectors and fuse them through an attention mechanism based feature fusion network, the fusion network is trained to dynamically learn the contribution weight of different modal features to behavior prediction in a specific context environment, and output a fused feature representation; S4: input the fused feature representation into a time series prediction model, output a user future behavior intention probability distribution, the behavior intention probability distribution includes one or more predicted behaviors and their corresponding occurrence probabilities; S5: receive user real behavior feedback, calculate the loss between the prediction result and the real result, and use the loss to jointly optimize the model parameters in the cross-modal fusion module and the user behavior prediction module through a back propagation algorithm to realize end-to-end adaptive learning.

[0027] Finally: the above only for the preferred embodiments of the present application, and not for limiting the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A user behavior prediction system based on multi-modal data fusion, characterized in that, Comprise: A multi-modal data acquisition module for real-time acquisition of raw data of a user from multiple heterogeneous data sources; A preprocessing and feature extraction module for data cleaning and standardization processing of raw data of each modality respectively, and extracting high-dimensional feature vectors using deep neural network models corresponding to the characteristics of the modalities to form visual features, auditory features, text features, physiological features and context features; A cross-modal fusion module for receiving multi-modal feature vectors and performing fusion through a feature fusion network based on an attention mechanism, the fusion network being trained to dynamically learn the contribution weight of different modalities to behavior prediction in a specific context, and output a fused feature representation; A user behavior prediction module for inputting the fused feature representation into a time series prediction model to output a future behavior intention probability distribution of the user, the behavior intention probability distribution including one or more predicted behaviors and their corresponding occurrence probabilities; A model optimization and feedback module for receiving user real behavior feedback, calculating the loss between the predicted result and the real result, and using the loss to jointly optimize the model parameters in the cross-modal fusion module and the user behavior prediction module through a back propagation algorithm to realize end-to-end adaptive learning. 2.The user behavior prediction system based on multi-modal data fusion according to claim 1, characterized in that: In the multi-modal data acquisition module, the dynamic capture of multi-dimensional raw information of the user is realized by integrating heterogeneous data sources distributed in different scenes. The raw data includes visual modal data, auditory modal data, text modal data, physiological signal modal data, and environmental context modal data. 3.The user behavior prediction system based on multi-modal data fusion of claim 1, wherein: In the preprocessing and feature extraction module, multi-modal raw data is converted into structured features that can be directly used for fusion analysis through a targeted processing procedure; first, fine preprocessing is performed for each modality data: for visual modality data, image denoising, illumination normalization, and target detection operations are performed to purify the data, and then the image is uniformly scaled to a preset size through scale normalization; for auditory modality data, audio noise reduction, endpoint detection, and sampling rate unification are performed, and then converted into a mel spectrum graph for subsequent processing; text modality data is processed through word segmentation, stop word removal, and special symbol cleaning, then standardized using stem extraction technology, and then preliminary numericalization is achieved through word vector mapping; physiological signal modality data needs to be corrected for baseline drift, waveform rejection, and signal resampling; environmental context modality data is preprocessed through format standardization, missing value filling, and discrete feature encoding.

4. The user behavior prediction system based on multi-modal data fusion according to claim 3, characterized in that: In the feature extraction stage, the module matches deep neural network models according to the information characteristics of different modalities: the visual feature extraction uses an improved ResNet-50 architecture, which captures hierarchical features from local textures to global contours through multi-scale convolutional layers, and finally outputs a 2048-dimensional visual feature vector, while introducing an attention mechanism to strengthen the feature weight of key regions of user expressions and actions; Convolution calculation is performed on the local regions of the input image to extract low-level features such as edges and textures; The auditory feature extraction is based on a pre-trained VGGish model, which extracts audio spectral features from a mel-spectrogram through stacked convolution and pooling layers, and combines an LSTM network to capture the prosody and speech rate temporal features, generating a 1280-dimensional auditory feature vector; The text feature extraction adopts a BERT pre-training model to understand the context semantics through a bidirectional Transformer structure, taking the hidden layer output corresponding to the [CLS] marker as a 768-dimensional text feature vector, and supporting dynamic adjustment of the segmentation granularity to adapt to different text lengths; The physiological feature extraction constructs a multi-branch CNN-LSTM network, in which the convolution branch extracts local features such as heart rate variability, and the LSTM branch captures long-term dependencies of physiological signals, and the fusion forms a 512-dimensional physiological feature vector; The environmental context feature is modeled by a graph neural network to associate the spatial and temporal relationships, and is mapped to a 256-dimensional feature vector. 5.The user behavior prediction system based on multi-modal data fusion of claim 1, wherein: In the cross-modal fusion module, the heterogeneous feature vectors of vision, hearing, text, physiology, and environmental context are converted into unified fusion feature representations, the correlation weights of different modalities in a specific scene are dynamically modeled through a fusion network based on an attention mechanism, and multi-dimensional information is complementarily enhanced; first, the dimension alignment and spatial mapping of the input modal feature vectors are performed, the visual, auditory, textual, physiological, and contextual features are uniformly mapped to a shared feature space through linear transformation, and the calculation method is specifically: wherein, denotes a learnable linear projection matrix designed for modality m, denotes a bias term for modality m, is the mapped feature vector; In the fusion network, the intra-modal attention layer strengthens the key information within the single-modal feature through the self-attention mechanism, gives higher weight to the feature of the user expression area in the visual feature, and enhances the representation of the feature of the sentiment tendency words in the text feature. The calculation formula is , is the mth modal feature vector after self-attention enhancement, wherein the self-attention realizes importance weighting by calculating the correlation degree of the internal elements of the feature vector. The feature aggregation layer then performs weighted fusion of the modal features based on the attention weights, and performs nonlinear transformation and dimension compression through a multilayer perceptron, finally generating a fixed-dimensional fusion feature vector: wherein, denotes the joint representation vector after multimodal fusion, denotes the m-th modality attention weight, denotes the m-th modality feature vector after self-attention enhancement. 6.The user behavior prediction system based on multi-modal data fusion of claim 1, wherein: In the user behavior prediction module, the unified feature representation generated by the cross-modal fusion module is converted into a user behavior intention prediction result, based on the multi-dimensional user information contained in the fusion feature, the time sequence prediction model captures the time sequence evolution law of the user behavior, outputs the behavior intention probability distribution in the future period of time, and outputs the behavior intention probability distribution in the future period of time representing the expansion and modeling of the time sequence dimension; The fusion features are sorted according to timestamps to construct a time sequence feature sequence of user behaviors wherein is a historical window length, is a current time, and time sequence information is injected through position encoding to ensure that the model perceives the time sequence order of behaviors; a bidirectional LSTM layer is used as a bottom time sequence encoder to scan the sequence from two directions respectively to generate a hidden state sequence containing local time sequence patterns to capture the dynamic change trend of recent user behaviors.

7. The user behavior prediction system based on multi-modal data fusion according to claim 6, characterized in that: After the time sequence feature encoding is completed, the module compresses the variable-length time sequence feature sequence into a fixed-dimension global time sequence feature vector through an adaptive pooling layer , and inputs it into a prediction head composed of a multi-layer perception and a softmax classifier, wherein the multi-layer perception further mines high-order interactions between features through nonlinear transformation, and the softmax classifier maps the transformed features to a preset behavior intention category space, finally outputs the occurrence probability of each behavior category, and forms a complete behavior intention probability distribution , wherein, represents the total number of behavior categories, and ; and a few-shot learning mechanism is used to adapt to new behavior types, and a confidence threshold is introduced to filter low-probability prediction results. 8.The user behavior prediction system based on multi-modal data fusion of claim 1, wherein: In the model optimization and feedback module, the model parameters of the cross-modal fusion module and the user behavior prediction module are dynamically adjusted by continuously receiving the deviation signals of the user's real behavior data and the prediction results; Firstly, a multi-dimensional loss calculation system is constructed. According to the discrete and sequential characteristics of user behavior prediction, a hybrid loss function is designed: the cross-entropy loss is used to quantify the deviation between the predicted behavior intention probability distribution and the real behavior label, the time consistency loss is introduced to constrain the rationality of the behavior sequence, the transition probability between the predicted results at continuous time and the real behavior sequence is compared to punish the prediction with large jumps, and the modal feature reconstruction loss is added to calculate the reconstruction error by using the autoencoder to infer the original modal feature from the fused feature. Finally, the total loss function is wherein, is the loss weight, which is dynamically adjusted according to the scene.

9. The method of claim 1-8, wherein the method of predicting user behavior based on multi-modal data fusion is performed by using the system of predicting user behavior based on multi-modal data fusion according to any one of claims 1-8. Comprising the following steps: S1: Real-time acquisition of user's original data from multiple heterogeneous data sources; S2: The original data of each modality is respectively subjected to data cleaning and standardization processing, and a deep neural network model corresponding to the modality characteristics is used to extract a high-dimensional feature vector, forming visual features, auditory features, text features, physiological features, and context features; S3: Receive multi-modal feature vectors and fuse them through a feature fusion network based on an attention mechanism, the fusion network is trained to dynamically learn the contribution weights of different modal feature vectors to behavior prediction in a specific context, and output a fusion feature representation; S4: Input the fusion feature representation into a time series prediction model to output the user's future behavior intention probability distribution, the behavior intention probability distribution includes one or more predicted behaviors and their corresponding occurrence probabilities; S5: Receive user real behavior feedback, calculate the loss between the prediction result and the real result, and use the loss to jointly optimize the model parameters in the cross-modal fusion module and the user behavior prediction module through a back propagation algorithm, to realize end-to-end adaptive learning.

Citation Information

Patent Citations

  • Transformer substation safe operation monitoring method and system based on OpenCV

    CN118485973A

  • Intelligent financial product recommendation method and system based on multi-modal user portraits

    CN118710376A

  • Multi-element sales planning agent system and method

    CN120430836A

  • Multimodal data-based method and system for recognizing cognitive engagement in classroom

    US20250022314A1

Cited By

  • Multi-modal data multi-model combined training method and system

    CN121144858A

  • Fatigue state recognition method based on learnable filter bank and joint regularization

    CN121265057A

  • Automobile user behavior analysis system and method based on big data

    CN121302276A

  • Mineral prediction method and system based on multi-modal Transform architecture

    CN121456689A

  • Submarine optical cable disturbance detection method and device and computer readable storage medium

    CN121479530A