Dynamic user portrait generation method and device based on multi-modal feature fusion
By using multimodal feature fusion and deep learning, the system collects and analyzes users' image, behavioral, and emotional data in real time, solving the problems of single-dimensionality and insufficient dynamic updates in existing user profiling technologies, and generating accurate, dynamic, and personalized user profiles.
Patent Information
- Application Number
- CN202511441620.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-02-10
AI Technical Summary
Existing user profiling technologies suffer from limitations such as single-dimensionality, insufficient dynamic updates, inadequate multimodal fusion, and insufficient cognitive modeling, resulting in insufficient personalization and interpretability of the profiles.
By employing a multimodal feature fusion approach, we collect and integrate users' image, behavioral, emotional, and cognitive data in real time. Through deep learning and sentiment analysis, we dynamically characterize users' interests, emotions, and cognitive biases to generate accurate and dynamic user profiles.
It enables real-time updates of user profiles, improving timeliness and accuracy, comprehensively reflecting users' interests, emotions, and cognitive characteristics, and enhancing the personalization of the profiles.
Smart Images

Figure CN121502185A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of user portrait, and particularly relates to a dynamic user portrait generation method, system, device and medium based on multi-modal feature fusion. BACKGROUND
[0002] With the rapid development of the Internet and social media, user behavior data, social data and multi-modal information (text, image, video, etc.) are constantly emerging. How to use these data to accurately construct user portraits to support precision marketing, personalized recommendation and social governance has become an important issue in academia and industry. Traditional user portrait technology usually relies on static attributes of users, such as age, gender, geographic location, interests, etc. These static data can provide some reference value, but they cannot fully reflect the dynamic cognition, emotional state and cognitive bias of users.
[0003] In recent years, the rapid development of artificial intelligence and big data technology has provided a new breakthrough for user portrait technology. In particular, based on deep learning, large language models (such as GPT-4) and graph neural networks, researchers have begun to try to construct user portraits from a more dynamic and complex perspective, especially by capturing the emotional changes, cognitive biases and emotional fluctuations of user behavior to form more accurate portraits.
[0004] However, the existing user portrait technology has the following four problems: 1. Single dimension; mostly relying on static attributes and historical behavior data, lacking modeling of deep dynamic features such as user emotions and cognition.
[0005] 2. Insufficient dynamic update; most solutions use batch processing, which cannot reflect the changes in user interests and cognitive fluctuations in real time.
[0006] 3. Insufficient multi-modal fusion; there is a lack of effective fusion methods for heterogeneous data such as text, image, video, etc., and the information utilization rate is not high.
[0007] 4. Insufficient cognitive modeling; existing solutions ignore users' cognitive biases, values and psychological states, resulting in insufficient personalization and interpretability of the portrait. SUMMARY
[0008] The main purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, and to provide a dynamic user portrait generation method, system, device and medium based on multi-modal feature fusion, which can collect and fuse multi-modal data of users in real time, and through deep learning, sentiment analysis and cognitive modeling, etc. The changing trend of user interests, emotions and cognitive biases is dynamically described, thereby generating accurate, dynamic and personalized user portraits.
[0009] In order to achieve the above object, the present application adopts the following technical solutions: The dynamic user portrait generation method based on multi-modal feature fusion comprises the following steps: Collecting image, behavior, emotion and cognition data of the user, and pre-processing the data; Feature extraction and modeling are performed on the pre-processed data; Multi-modal feature fusion is used to fuse image, behavior, emotion and cognition features to generate a unified portrait representation; Dynamic adjustment and optimization are used to monitor and collect user behavior data, monitor user emotional state and cognitive state, and update the user portrait through a feedback mechanism when the user behavior, emotional state or cognitive state changes.
[0010] The present application also provides a dynamic user portrait generation system based on multi-modal feature fusion, which applies the dynamic user portrait generation method based on multi-modal feature fusion provided by the present application. The data acquisition module is used to collect image, behavior, emotion and cognition data of the user, and pre-process the data; The feature extraction module is used to extract features and model the pre-processed data of the data acquisition module; The multi-modal feature fusion module is used to fuse image, behavior, emotion and cognition features to generate a unified portrait representation; The dynamic adjustment and optimization module is used to monitor and collect user behavior data, monitor user emotional state and cognitive state, and update the user portrait through a feedback mechanism when the user behavior, emotional state or cognitive state changes.
[0011] The present application also provides an electronic device, which comprises: At least one processor; and, The memory is in communication connection with the at least one processor; wherein, The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to execute the dynamic user portrait generation method based on multi-modal feature fusion provided by the present application.
[0012] The present application also provides a computer readable storage medium storing a program, which is executed by a processor to implement the dynamic user portrait generation method based on multi-modal feature fusion provided by the present application.
[0013] Compared with the prior art, the present application has the following advantages and beneficial effects: 1、Compared with the static user portrait in the prior art, the application proposes a dynamic updating mechanism based on user behavior change and emotional fluctuation, which can capture the changes of user behavior and emotion in real time and quickly update the user portrait; the method can adjust the portrait according to the real-time interaction, emotional fluctuation and cognitive bias of the user, ensure that the user portrait always accurately reflects the current state of the user, and thus improve the timeliness and accuracy of the user portrait.
[0014] 2、The application proposes a multi-modal data fusion method, which can effectively integrate heterogeneous information from different data sources such as text, image, behavior data and emotional data. This fusion method overcomes the limitations of traditional methods that rely only on a single data source (such as text data or behavior data), thereby providing a more comprehensive and personalized user portrait that can more accurately capture the user's interests, emotions and cognitive characteristics.
[0015] 3、The application combines psychological theories (such as the Big Five Personality Model and the Cognitive Bias Model) to model the cognitive characteristics of users and identify and correct their cognitive biases, making the user portrait more comprehensive, covering not only the user's behavioral characteristics and emotional characteristics but also their cognitive biases, avoiding the portrait bias caused by ignoring the cognitive level in traditional methods, and enhancing the accuracy and individualization of the user portrait. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 is a flowchart of the method of the present application.
[0018] Figure 2 is a block diagram of the system of the embodiment of the present application.
[0019] Figure 3 is a structural diagram of the electronic device of the embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to make those skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0021] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0022] like Figure 1 As shown, in one embodiment of this application, a dynamic user profile generation method based on multimodal feature fusion is provided, including the following steps: Collect users' image, behavioral, emotional, and cognitive data, and preprocess the data; Feature extraction and modeling are performed on the preprocessed data; Multimodal feature fusion integrates image, behavioral, emotional, and cognitive features to generate a unified profile representation; Dynamically adjust and optimize, monitor and collect user behavior data, monitor user emotional and cognitive states, and update user profiles through feedback mechanisms when user behavior, emotional state, or cognitive state changes.
[0023] In this embodiment, user image, behavior, emotion, and cognitive data are collected, specifically as follows: Collect users' image, behavioral, emotional, and cognitive data from multiple data sources, such as social media, e-commerce platforms, and mobile applications. The collected data is categorized into text data, image data, behavioral data, and cognitive data. Text data, such as social media posts, comments, and chat logs, is represented as ,in This represents a specific piece of text data; Image data, such as images or videos from social media platforms, is represented as ,in This represents a single image data point; Behavioral data, such as user clicks, purchases, and search history, is represented as ,in This represents a specific piece of user behavior data; Cognitive data, such as users' emotional fluctuations and cognitive biases, is represented as .
[0024] In this embodiment, data preprocessing includes data cleaning, noise reduction, and standardization; for text data, text segmentation is also included, specifically: Text data Word segmentation is performed to obtain a vocabulary set. And convert it into word vector representation. .
[0025] In this embodiment, feature extraction and modeling are performed on the preprocessed data, specifically including: For each text data Sentiment analysis is performed using a pre-trained BERT model to extract sentiment information, resulting in a sentiment score for the text data, represented as follows: ; in, For text Sentiment score; The sentiment scores of multiple texts constitute the sentiment feature, represented as ; Time series analysis methods (such as LSTM, GRU, etc., LSTM in this embodiment) are used to predict sentiment trends and capture user sentiment fluctuations. The time series analysis method includes: inputting user sentiment scores at different times into the model in chronological order; the model processes the input step by step to capture the changing patterns of sentiment over time, and based on this, predicts the sentiment trend over a future period, thereby reflecting the dynamic evolution trend of user emotions; the sentiment trend prediction is expressed by the following formula: ; in, It refers to the emotional state at the current moment. It is user behavior data or sentiment data; Use models such as BERT or LDA to analyze user comments or posts. Topic modeling is performed; the topic modeling process includes: after the user's text is input into the model, the model automatically identifies and extracts the potential topic information and assigns a corresponding topic category or topic distribution to each text; by integrating all text results, a topic feature vector reflecting the user's main interests and preferences is obtained. ; In this embodiment, the BERT model is used for modeling, and the topic modeling result is expressed by the following formula: ; Image data feature extraction, specifically: For image data Visual feature extraction is performed using convolutional neural networks or The model extracts the visual features of each image, represented as... ,in For image Visual features; use The image data is processed to obtain the image feature vector: ; wherein, is a feature vector obtained after convolution processing on the image ; behavior data feature extraction, specifically: time series modeling on the user's behavior data , capturing the trend of user behavior. Use deep learning models (such as RNN, LSTM) for time series modeling to generate sequence features of user behavior ; In this embodiment, LSTM is used to model the behavior data, represented as: ; wherein, is the behavior hidden state generated at time step , is the behavior data of the user at time step ; cognitive feature extraction, specifically: extract cognitive features from user social interactions and behavior data , such as cognitive bias, emotional fluctuations, etc. Use psychological models (such as the Big Five Personality Model, Cognitive Bias Model, etc.) for modeling, and cognitive features are represented as , wherein, is the cognitive feature of each user (such as extroversion, neuroticism, etc.).
[0026] wherein, the modeling process using psychological models includes: establishing an index system according to the cognitive dimensions of the target image (for example, an index set composed of personality dimensions and common cognitive bias types); combining social interactions and behavior data to construct cognitive-related observation features (including but not limited to language and emotional cues, interaction and response patterns, preference consistency and switching features, temporal context and scene factors, etc.); based on the preset psychological model, the observation features are inferred to obtain the estimation results and confidence of each cognitive dimension; consistency calibration and stabilization processing of the estimation results of different data sources and different time periods are performed to suppress incidental fluctuations and retain trend information; output structured cognitive feature vectors and standardize them to facilitate fusion with other modal features and subsequent updates.
[0027] In this embodiment, multi-modal feature fusion is specifically: use attention mechanism or graph neural network (GNN) to fuse different modal data to generate a unified user image representation ; If the attention mechanism is used, it is specifically: for each modal data and weighted average: ; wherein, is the weight of the modal , is the feature representation of the modal .
[0028] If a graph neural network is used, it is specifically: for each modal data use a graph neural network for fusion, generate the final user portrait representation : ; wherein, GNN() represents the fusion operation of the graph neural network, is the feature representation of different modalities; Through multi-modal feature fusion, the multi-modal features are integrated into a unified user portrait representation , denoted as: ; wherein, is the emotional feature, is the image feature, is the behavior feature, is the cognitive feature.
[0029] In order to ensure the real-time and accuracy of the user portrait, the present application proposes a dynamic adjustment mechanism based on user behavior changes, through real-time monitoring and adjustment of user behavior, emotional fluctuations and cognitive bias, the present application can automatically update the user portrait when the user behavior changes, to ensure that it always accurately reflects the current state of the user. In this embodiment, dynamic adjustment and optimization specifically includes: continuously monitor the user's behavior changes, collect the latest user behavior data , use time series analysis method to detect the change of user behavior: ; wherein, represents the amount of change in user behavior, is the latest behavior data, is the last updated behavior data; monitor the user's emotional state, when a significant change in user behavior is detected, process the user's behavior data through the emotion analysis model BERT, analyze the user's emotional fluctuations, and obtain the user's new emotional state , The process is represented as: ; wherein, representing new behavior data performing sentiment analysis, outputting the user's emotional score; When the user's emotional fluctuation is large, adjust the emotional features according to the emotional changes, so that the portrait accurately reflects the current emotional state, represented as: ; wherein, is the adjustment amount of emotional features, is the adjusted emotional features, is the emotional features before adjustment; Monitor the user's cognitive state, combine cognitive psychology, analyze the user's behavior patterns and psychological characteristics, identify the user's cognitive bias (such as stubborn attitude towards certain topics); By comparing the differences between the user's current behavior and historical behavior, it is determined whether there is a cognitive bias: ; If exceeds a certain threshold, it means that the user's cognitive bias has changed, and the portrait needs to be corrected; After identifying the cognitive bias, adjust the user's cognitive features to correct the bias, represented as: ; wherein, is the cognitive bias correction amount, is the corrected cognitive features.
[0030] When detecting that the user's behavior, emotion, cognition and other changes are large, update the user portrait in real time according to the changes; In this embodiment, updating the user portrait specifically includes: Let the current user portrait be , the updated user portrait is , and the update formula is: ; wherein, is the change amount of the user portrait at the current time, reflecting the changes in user behavior and emotional state; Optimize the feedback of the user portrait historical data to improve the accuracy and personalization of the portrait generation, for example, according to the user's past behavior and emotional change pattern, adjust the model parameters, so that the portrait update is more accurate.
[0031] According to the user historical data Optimization, represented as: ; wherein, This refers to an algorithm that optimizes user profiles based on historical data, adjusting model parameters to enhance the personalization of the profiles. In this embodiment, The process of employing a personalized incremental learning algorithm based on historical data includes: 1) Data Construction and Weighting: Construct a scrolling window of user historical data by time and assign weights to data of different time periods and quality. 2) Personalized warm-up: Freeze the general feature layer and fine-tune only the individualized parameters to enable the model to quickly adapt to the user's recent behavioral and emotional changes; 3) Online incremental update: Employs small-step incremental optimization and maintains a replay buffer to avoid forgetting; 4) Modal gating and reliability integration: Dynamically allocate the contribution of different modalities based on data quality and modal coverage; 5) Stabilization and Calibration: Improve consistency and interpretability through parameter smoothing and result calibration; 6) Output and Write-back: Generate updated individualized profiles and write them back to the system for subsequent dynamic optimization.
[0032] This invention also designs a user profile update frequency adjustment mechanism, which dynamically adjusts the optimization frequency of the user profile based on the frequency of user behavior and emotional fluctuations; let the current optimization frequency be... The new optimized frequency is Then we have: ; in, The frequency adjustment was optimized to reflect the magnitude of user activity and emotional fluctuations.
[0033] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously.
[0034] Based on the same idea as the dynamic user profile generation method based on multimodal feature fusion in the above embodiments, the present invention also provides a dynamic user profile generation system based on multimodal feature fusion, which can be used to execute the above-described dynamic user profile generation method based on multimodal feature fusion. For ease of explanation, the structural diagram of the embodiment of the dynamic user profile generation system based on multimodal feature fusion only shows the parts related to the embodiments of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0035] like Figure 2As shown, the dynamic user profile generation system 100 based on multimodal feature fusion includes a data acquisition module 101, a feature extraction module 102, a multimodal feature fusion module 103, and a dynamic adjustment and optimization module 104. The data acquisition module 101 is used to collect users' image, behavior, emotion and cognitive data, and to preprocess the data; Feature extraction module 102 is used to extract features and model data from the data preprocessed by the data acquisition module; The multimodal feature fusion module 103 is used to fuse image, behavioral, emotional and cognitive features to generate a unified profile representation; The dynamic adjustment and optimization module 104 is used to monitor and collect user behavior data, monitor user emotional state and cognitive state, and update user profile when user behavior, emotional state or cognitive state changes through feedback mechanism.
[0036] It should be noted that the dynamic user profile generation system based on multimodal feature fusion of the present invention corresponds one-to-one with the dynamic user profile generation method based on multimodal feature fusion of the present invention. The technical features and beneficial effects described in the embodiments of the dynamic user profile generation method based on multimodal feature fusion are applicable to the embodiments of the dynamic user profile generation system based on multimodal feature fusion. For details, please refer to the description in the embodiments of the method of the present invention, which will not be repeated here.
[0037] Furthermore, in the above implementation of the dynamic user profile generation system based on multimodal feature fusion, the logical division of each program module is only an example. In actual applications, the above functions can be assigned to different program modules as needed, for example, for the sake of corresponding hardware configuration requirements or software implementation convenience. That is, the internal structure of the dynamic user profile generation system based on multimodal feature fusion is divided into different program modules to complete all or part of the functions described above.
[0038] like Figure 3 As shown, in another embodiment, an electronic device is provided for implementing a dynamic user profile generation method based on multimodal feature fusion. The electronic device 200 may include a first processor 201, a first memory 202 and a bus, and may also include a computer program stored in the first memory 202 and executable on the first processor 201, such as a dynamic user profile generation program 203 based on multimodal feature fusion.
[0039] The first memory 202 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the first memory 202 can be an internal storage unit of the electronic device 200, such as the portable hard drive of the electronic device 200. In other embodiments, the first memory 202 can also be an external storage device of the electronic device 200, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 200. Furthermore, the first memory 202 can include both internal and external storage units of the electronic device 200. The first memory 202 can be used not only to store application software and various types of data installed on the electronic device 200, such as the code of the dynamic user profile generation program 203 based on multimodal feature fusion, but also to temporarily store data that has been output or will be output.
[0040] In some embodiments, the first processor 201 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The first processor 201 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the first memory 202 and calls data stored in the first memory 202 to perform various functions of the electronic device 200 and process data.
[0041] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device 200, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0042] The dynamic user profile generation program 203 based on multimodal feature fusion, stored in the first memory 202 of the electronic device 200, is a combination of multiple instructions. When run in the first processor 201, it can achieve the following: Collect users' image, behavioral, emotional, and cognitive data, and preprocess the data; Feature extraction and modeling are performed on the preprocessed data; By integrating image, behavioral, emotional, and cognitive features, a unified profile representation can be generated. Monitor and collect user behavior data, monitor user emotional and cognitive states, and update user profiles through feedback mechanisms when user behavior, emotional state, or cognitive state changes.
[0043] Furthermore, if the modules / units integrated in the electronic device 200 are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0044] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0045] In another embodiment, a computer-readable storage medium is provided, storing a program that, when executed by a processor, implements the dynamic user profile generation method based on multimodal feature fusion of the present invention, specifically: Collect users' image, behavioral, emotional, and cognitive data, and preprocess the data; Feature extraction and modeling are performed on the preprocessed data; By integrating image, behavioral, emotional, and cognitive features, a unified profile representation can be generated. Monitor and collect user behavior data, monitor user emotional and cognitive states, and update user profiles through feedback mechanisms when user behavior, emotional state, or cognitive state changes.
[0046] The computer-readable storage medium may be transient or non-transient. Exemplary examples include, but are not limited to, various media capable of storing computer program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0047] For example, the processor may be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0048] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0049] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A dynamic user profile generation method based on multimodal feature fusion, characterized in that, Includes the following steps: Collect users' image, behavioral, emotional, and cognitive data, and preprocess the data; Feature extraction and modeling are performed on the preprocessed data; Multimodal feature fusion integrates image, behavioral, emotional, and cognitive features to generate a unified profile representation; Dynamically adjust and optimize, monitor and collect user behavior data, monitor user emotional and cognitive states, and update user profiles through feedback mechanisms when user behavior, emotional state, or cognitive state changes.
2. The dynamic user profile generation method based on multimodal feature fusion according to claim 1, characterized in that, Collect users' image, behavioral, emotional, and cognitive data, specifically: The system collects user image, behavioral, emotional, and cognitive data from multiple data sources, including social media, e-commerce platforms, and mobile applications. The collected data is categorized into text data, image data, behavioral data, and cognitive data. Text data includes social media posts, comments, and chat logs, represented as... ,in This represents a specific piece of text data; Image data includes photos or videos from social media platforms, represented as ,in This represents a single image data point; Behavioral data includes user clicks, purchases, and search records, represented as ; Cognitive data includes users' emotional fluctuations and cognitive biases, represented as .
3. The dynamic user profile generation method based on multimodal feature fusion according to claim 2, characterized in that, Data preprocessing includes data cleaning, noise reduction, and standardization. For text data, it also includes text segmentation, specifically: Text data Word segmentation is performed to obtain a vocabulary set. And convert it into word vector representation. .
4. The dynamic user profile generation method based on multimodal feature fusion according to claim 3, characterized in that, Feature extraction and modeling are performed on the preprocessed data, specifically including: For each text data Sentiment analysis is performed using a pre-trained BERT model to extract sentiment information, resulting in a sentiment score for the text data, represented as follows: ; in, For text Sentiment score; The sentiment scores of multiple texts constitute the sentiment feature, represented as ; Time series analysis is used to predict sentiment trends and capture user emotional fluctuations. The time series analysis method involves inputting user sentiment scores at different times into the model in chronological order. The model processes the input progressively to capture patterns of sentiment change over time and predicts future sentiment trends, thus reflecting the dynamic evolution of user emotions. Sentiment trend prediction is expressed by the following formula: ; in, It refers to the emotional state at the current moment. It is user behavior data or sentiment data; Use BERT or LDA models to analyze user comments or posts. Topic modeling is performed; the topic modeling process includes: after the user's text is input into the model, the model automatically identifies and extracts the potential topic information and assigns a corresponding topic category or topic distribution to each text; by integrating all text results, a topic feature vector reflecting the user's main interests and preferences is obtained. ; The topic modeling results are represented by the following formula: ; Image data feature extraction, specifically: For image data Visual feature extraction is performed using convolutional neural networks or The model extracts the visual features of each image, represented as... ,in For image Visual features; use The image data is processed to obtain the image feature vector: ; in, For the image The feature vector obtained after convolution processing; Behavioral data feature extraction, specifically: User behavior data Perform time-series modeling to capture trends in user behavior; use deep learning models for time-series modeling to generate sequential features of user behavior. ; Using LSTM to model behavioral data: ; in, In time step The generated behavior hides the state. It is the user at the time step Behavioral data; Cognitive feature extraction, specifically: Extracting cognitive features from users' social interactions and behavioral data This includes cognitive biases and emotional fluctuations, which are modeled using psychological models. Cognitive characteristics are represented as... ,in, Cognitive characteristics for each user; The process of modeling using psychological models includes: establishing an indicator system based on the cognitive dimensions of the target profile; constructing cognitive-related observational features by combining social interaction and behavioral data; inferring the observational features based on the pre-set psychological model to obtain the estimation results and confidence levels of each cognitive dimension; performing consistency calibration and stabilization processing on the estimation results from different data sources and different time periods to suppress occasional fluctuations and retain trend information; and outputting structured cognitive feature vectors and standardizing them for fusion with other modal features and subsequent updates. The psychological models used include the Big Five personality model and the cognitive bias model.
5. The dynamic user profile generation method based on multimodal feature fusion according to claim 4, characterized in that, Multimodal feature fusion specifically refers to: By employing attention mechanisms or graph neural networks to fuse data from different modalities, a unified user profile representation can be generated. ; The attention mechanism is specifically used for each modal data. Calculate its corresponding weight And perform a weighted average: ; in, It is modal The weight, It is modal Feature representation; The specific use of graph neural networks is as follows: For each modal data The final user profile representation is generated by fusing graph neural networks. : ; Where GNN() represents the fusion operation of a graph neural network. These are feature representations of different modalities; By fusing multimodal features, multimodal features are integrated into a unified user profile representation. , is represented as: ; in, As an emotional characteristic, For image features, For behavioral characteristics, These are cognitive characteristics.
6. The dynamic user profile generation method based on multimodal feature fusion according to claim 5, characterized in that, Dynamic adjustment and optimization specifically include: Continuously monitor changes in user behavior and collect the latest user behavior data. Time series analysis methods are used to detect changes in user behavior: ; in, A quantity representing changes in user behavior. For the latest behavioral data, This is the behavioral data from the last update; By monitoring users' emotional states, and when significant changes in user behavior are detected, the user's behavioral data is processed using the BERT sentiment analysis model to analyze emotional fluctuations and determine the user's new emotional state. , This process can be represented as: ; in, Indicates new behavioral data Perform sentiment analysis and output the user's sentiment score; When a user's emotions fluctuate, the emotional characteristics are adjusted according to the emotional changes so that the profile accurately reflects the current emotional state, represented as: ; in, The moderating factor for emotional characteristics. For the adjusted emotional characteristics, The emotional characteristics before adjustment; By monitoring users' cognitive states and combining cognitive psychology, we can identify cognitive biases by analyzing users' behavioral patterns and psychological characteristics. Furthermore, by comparing the differences between users' current and historical behaviors, we can determine whether cognitive biases exist. ; like If a certain threshold is exceeded, it indicates that the user's cognitive bias has changed, and the user profile needs to be revised. After identifying cognitive biases, these biases are corrected by adjusting the user's cognitive characteristics, as shown below: ; in, This is the correction factor for cognitive bias. These are the revised cognitive features.
7. The dynamic user profile generation method based on multimodal feature fusion according to claim 6, characterized in that, Updating user profiles specifically includes: When significant changes in a user's cognition or behavior are detected, the user profile is updated in real time based on these changes; let the current user profile be... The updated user profile is The updated formula is: ; in, The change in the user profile at the current moment reflects changes in user behavior and emotional state; By optimizing based on feedback from historical user profile data, the accuracy and personalization of user profile generation can be improved. Specifically, this involves using historical user data... Optimization is represented as: ; in, This refers to an algorithm that optimizes user profiles based on historical data, adjusts model parameters, and enhances the personalization of the profiles. The optimization frequency of user profiles is dynamically adjusted based on the frequency of user behavior and emotional fluctuations; let the current optimization frequency be... The new optimized frequency is ,but: ; in, The frequency adjustment was optimized to reflect the magnitude of user activity and emotional fluctuations.
8. A dynamic user profile generation system based on multimodal feature fusion, characterized in that, The dynamic user profile generation method based on multimodal feature fusion according to any one of claims 1-7 includes a data acquisition module, a feature extraction module, a multimodal feature fusion module, and a dynamic adjustment and optimization module. The data acquisition module is used to collect users' image, behavioral, emotional, and cognitive data, and to preprocess the data. The feature extraction module is used to extract features and model the data preprocessed by the data acquisition module. The multimodal feature fusion module is used to fuse image, behavioral, emotional, and cognitive features to generate a unified profile representation. The dynamic adjustment and optimization module is used to monitor and collect user behavior data, monitor user emotional and cognitive states, and update user profiles when user behavior, emotional state, or cognitive state changes through a feedback mechanism.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory that is communicatively connected to at least one processor; wherein, The memory stores computer program instructions that can be executed by at least one processor, which enables the at least one processor to perform the dynamic user profile generation method based on multimodal feature fusion as described in any one of claims 1-7.
10. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the dynamic user profile generation method based on multimodal feature fusion as described in any one of claims 1-7.