A Method for Analyzing User Preferences Based on a Memory-Augmented Multimodal Large Model
Through the user preference analysis method based on memory-enhanced multimodal large model, the problems of multi-level timing preference conflict, implicit preference extraction, cross-domain preference migration and multi-modal signal consistency processing are solved, and efficient personalized services of the intelligent system are realized.
Patent Information
- Application Number
- CN202510345153.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing technology is difficult to effectively distinguish and coordinate multi-level timing preference conflicts, cannot capture implicit preference signals, and lacks cross-domain preference migration and multimodal signal consistency processing mechanisms, which affects the personalized service quality of intelligent systems.
The user preference analysis method based on memory-enhanced multimodal large model is adopted. Through the preprocessing of multimodal interactive data and timing preference hierarchical recognition, multi-level timing preference classification results and transformation rules are generated, preference conflicts are coordinated, and preference understanding and response capabilities are improved through implicit preference extraction and verification, cross-domain preference boundary recognition and cross-modal preference consistency analysis.
It realizes the capture and understanding of explicit and implicit preferences, coordinates multi-level preference conflicts, handles cross-domain and multi-modal preferences, improves the personalized service quality of the intelligent system, and solves the core challenges of the current preference analysis system.
Smart Images

Figure CN119884981B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multimodal data processing, and in particular, to a method for parsing user preferences based on a memory-enhanced multimodal large model. Background Art
[0002] With the rapid development of artificial intelligence technology, the interaction methods between intelligent systems and users have become increasingly diverse, expanding from single text input to multimodal interactions such as images, voices, and gestures. In such a complex interaction environment, accurately understanding and responding to user preferences has become the key to improving the user experience. As the core technology of personalized intelligent services, user preference parsing directly affects the adaptability of the system and user satisfaction. This is particularly important for products such as intelligent assistants, recommendation systems, and creative design tools. Therefore, constructing a parsing system that can accurately capture and understand user multimodal preferences has great theoretical value and application significance.
[0003] Currently, the research on user preference parsing mainly focuses on capturing preferences in single modality and at specific time points. Traditional methods such as collaborative filtering, content-based filtering, and rule-based reasoning systems usually only consider explicit preference signals from a single source. Recent research has begun to attempt to fuse multimodal information. For example, the CLIP model uses contrastive learning to achieve joint representation of text and images, and the MultiSense system combines voice and facial expressions to understand user emotions. In terms of time series, some research has introduced time decay models. For example, BPR-MF uses an exponential decay function to weight historical behaviors. For preference transfer, methods such as TransRec have proposed knowledge transfer techniques based on domain mapping. In terms of conflict handling, existing systems mostly adopt simple priority rules or the strategy of covering old preferences with the most recent preferences.
[0004] Despite certain progress, the current preference parsing technology still has problems. First, in terms of coordinating multi-level time series preference conflicts, existing technologies regard all preferences as being at the same time level, and cannot effectively distinguish short-term temporary preferences, medium-term task preferences, and long-term stable preferences, resulting in an inability to determine whether the "high-saturation" temporary needs in a specific scenario should override the "natural-saturation" long-term preferences, and when the temporary preferences should expire. Second, in terms of implicit preference extraction and verification, existing systems mainly rely on explicit user feedback, and cannot capture implicit preference signals contained in behaviors such as users' slight hesitation and repeated adjustments, or affect the user experience by frequently asking for verification. These technical gaps seriously restrict the in-depth understanding and accurate response of intelligent systems to user preferences. Summary of the Invention
[0005] The object of the invention is to provide a method for parsing user preferences based on a memory-enhanced multimodal large model to solve the above problems existing in the prior art.
[0006] Technical solution: A method for parsing user preferences based on a memory-enhanced multi-modal large model, comprising the following steps:
[0007] Collect and preprocess multi-modal interaction data to generate a standardized feature vector set and a cleaned multi-modal interaction sequence;
[0008] Perform temporal preference hierarchy recognition and classification based on the cleaned multi-modal interaction sequence to generate a multi-level temporal preference classification result and a temporal preference conversion rule set;
[0009] Perform multi-level preference conflict coordination and decision-making based on the multi-level temporal preference classification result and the temporal preference conversion rule set to generate a unified preference decision result;
[0010] Perform memory-enhanced preference model update and application based on the unified preference decision result to generate an updated user preference model, and apply it to the interaction interface and content generation to obtain personalized recommendation results.
[0011] Beneficial effects: The present invention can capture explicit and implicit preferences, handle cross-modal and cross-domain preferences, coordinate multi-level preference conflicts, and has long-term memory and progressive learning capabilities; the present invention breaks through the limitations of traditional preference parsing, realizes a preference understanding mechanism closer to human cognition, improves the personalized service quality of intelligent systems, solves the core challenges faced by current preference parsing systems, and provides a technical foundation for building a truly user-centered artificial intelligence system. Description of the drawings
[0012] Figure 1 It is a flowchart of the steps of a method for parsing user preferences based on a memory-enhanced multi-modal large model provided by an embodiment of the present application.
[0013] Figure 2 It is a flowchart of the steps for constructing an adaptive time decay model provided by an embodiment of the present application.
[0014] Figure 3 It is a flowchart of the steps for constructing a semantic-visual alignment space provided by an embodiment of the present application.
[0015] Figure 4 It is a flowchart of the steps for constructing a low-interference verification strategy provided by an embodiment of the present application. Detailed implementation manners
[0016] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0017] It should be particularly noted that, for clearly showing the step flow of the present application, serial numbers are marked for each step in the specification. These serial numbers are only for the convenience of description and do not limit the execution order of the steps. In actual operation, according to the technical requirements of the specific implementation scenario, the steps can be executed in an order different from that shown in the specification, and in some cases, parallel processing between steps can also be achieved.
[0018] As Figure 1 shown, a user preference parsing method based on a memory-enhanced multi-modal large model includes the following steps:
[0019] S1. Collect and preprocess multi-modal interaction data to generate a standardized feature vector set and a cleaned multi-modal interaction sequence;
[0020] S2. Perform temporal preference hierarchy identification and classification based on the cleaned multi-modal interaction sequence to generate a multi-level temporal preference classification result and a temporal preference conversion rule set;
[0021] S3. Perform cross-modal preference consistency analysis based on the standardized feature vector set to generate a cross-modal preference consistency report;
[0022] S4. Perform implicit preference extraction and verification based on the cleaned multi-modal interaction sequence to generate an implicit preference mapping table;
[0023] S5. Perform cross-domain preference boundary identification and migration rule generation based on the multi-level temporal preference classification result to generate a cross-domain preference migration rule set;
[0024] S6. Perform multi-level preference conflict coordination and decision-making based on the multi-level temporal preference classification result and the temporal preference conversion rule set to generate a unified preference decision result;
[0025] S7. Perform memory-enhanced preference model update and application based on the unified preference decision result to generate an updated user preference model and apply it to the interaction interface and content generation to obtain personalized recommendation results.
[0026] The multi-modal interaction data in this embodiment refers to the information containing multiple modalities such as text, pictures, videos, audio, and expressions transmitted between the user and the system during a human-computer dialogue. The multi-modal interaction sequence is the historical multi-modal interaction data between the user and the system arranged in chronological order, reflecting the context history of the user's behavior and the system's response. For example, the user first sends a picture, then sends a descriptive text, and finally sends a voice opinion, forming a complete multi-modal interaction sequence. The multi-level temporal preference classification result is a hierarchical and phased classification of the user's implicit preferences based on the understanding of the multi-modal interaction sequence and in combination with the context relationship. For example, when the user inputs "I think this picture is too bright", the corresponding multi-level temporal preference classification result is "Picture adjustment → Brightness → Reduce brightness". The temporal preference conversion rule set is a set of rules that convert the preferences expressed by the user in different historical contexts into explicit image processing operation instructions based on the memory bank and the set rules. For example, for the text expressing the picture brightness problem, the rule "Involving picture brightness → Automatically match memory → Image processing instruction" is formulated, and finally it is converted into the "Reduce picture brightness" instruction that can directly call the relevant image processing tools. The unified preference decision result comprehensively analyzes the context information and preference memory, and uniformly outputs the user preference decision result to guide the system behavior. For example, for the user's expression "I think this picture is too bright", combined with the picture in the context and the mapping of the similar text processing process saved in the memory bank, the finally generated unified preference decision result is: "Automatically perform the brightness reduction process on the current picture".
[0027] According to one aspect of the present application, step S1 is further as follows:
[0028] S11. Capture multi-source interaction data: Read the user device interface data, and synchronously collect text instructions, image operations, interaction behaviors, and time information through a multi-channel listening mechanism to construct an original multi-source interaction data packet, ensuring the time synchronization and source integrity of the data.
[0029] S12. Identify signal separation and modality: Read the multi-source interaction data packet, classify the data using a modality recognition algorithm, separate the mixed signals into text modality data streams, image modality data streams, and behavior modality data streams, and retain the temporal correspondence relationship between the modalities.
[0030] S13. Noise filtering and anomaly detection: Read each modality data stream, identify and remove the noise signals generated by user misoperations or system interference through statistical outlier detection and fluctuation threshold filtering, and output the denoised modality data streams.
[0031] S14. Context Association and Conversation Segmentation: Read the denoised modal data stream, divide the continuous interaction data into meaningful conversation units based on time interval and semantic coherence analysis, establish intra-modal and inter-modal context associations, and output segmented conversation sequences and context association graphs.
[0032] S15. Feature Extraction and Standardization: Read the segmented conversation sequences, extract semantic features, visual features, and behavioral features for different modal data, and make the features of each modality comparable through normalization processing, and output a set of standardized feature vectors and a cleaned multi-modal interaction sequence.
[0033] This embodiment realizes the acquisition and integration of high-quality multi-modal data, and establishes a solid foundation for subsequent preference parsing. The multi-channel listening mechanism is used for synchronous acquisition, and signal separation is performed through a modal recognition algorithm, effectively retaining the temporal correspondence relationship between modalities; combined with statistical outlier detection and fluctuation threshold filtering, the data quality can be improved and the misjudgment rate can be reduced; the conversation segmentation method based on time interval and semantic coherence analysis solves the problem that continuous interaction data is difficult to divide in traditional methods; the feature extraction and standardization processing for different modalities ensure the comparability and consistency of multi-modal data, laying a data foundation for subsequent preference recognition.
[0034] According to one aspect of the present application, step S2 is further as follows:
[0035] S21. Analyze Interaction Frequency and Duration: Read the cleaned multi-modal interaction sequence, calculate the occurrence frequency, duration, and repetition pattern of each preference expression, establish the temporal distribution characteristics of preferences, and output a preference temporal distribution mapping.
[0036] S22. Construct a Time Decay Function: According to the preference temporal distribution mapping, construct a non-linear time decay function, which can dynamically adjust the decay rate according to the historical distribution characteristics of preference expressions, so that different types of preferences have different "memory retention periods", and output an adaptive time decay model.
[0037] S23. Analyze Multi-scale Time Windows: Read the cleaned multi-modal interaction sequence, adopt multi-scale sliding window technology, and analyze the stability and change trend of preferences on short time windows (hour level), medium time windows (day level), and long time windows (month level) respectively, and output a multi-scale preference stability matrix.
[0038] S24. Hierarchical Clustering and Preference Classification: Read the multi-scale preference stability matrix and the data processed by the adaptive time decay model, and use the hierarchical clustering algorithm to classify preference expressions into short-term temporary preferences (high volatility, low stability), medium-term task preferences (medium stability), and long-term stable preferences (low volatility, high stability), and output a multi-level temporal preference classification result.
[0039] S25, Temporal Preference Conversion Boundary Detection: Read the multi-level temporal preference classification results, analyze the conditions and triggering factors for the preference to convert from one level to another, establish conversion rules and boundary conditions, and output the temporal preference conversion rule set for subsequent preference conflict coordination.
[0040] This embodiment solves the key problem that traditional preference models ignore the differences in the time dimension. Through the preference time distribution mapping technology, it can capture the time-varying characteristics of user preferences; the constructed non-linear time decay function dynamically adjusts the decay rate according to the historical distribution of preferences, enabling different types of preferences to have corresponding "memory retention periods"; the multi-scale sliding window technology analyzes the preference stability from three dimensions: short-term, medium-term, and long-term, providing unprecedented time granularity; the hierarchical clustering algorithm accurately classifies preferences into short-term temporary preferences, medium-term task preferences, and long-term stable preferences, improving the system's ability to distinguish different temporal preferences; the temporal preference conversion boundary detection technology effectively identifies the conditions and triggering factors for preference level conversion, providing an important basis for subsequent conflict coordination.
[0041] Since in a memory enhancement system, users' preferences usually have variations on multiple time scales (short-term temporary preferences, medium-term task preferences, and long-term stable preferences), and existing solutions do not address the conflict problem between preferences on different time scales. For example, a user's long-term preference when processing landscape photos may be "natural saturation", but in a specific scenario, they may temporarily need "high saturation", and the system cannot automatically determine which preference should take precedence, nor can it determine when the temporary preference should expire. There is a lack of a mechanism that can automatically identify, classify, and coordinate multi-level temporal preference conflicts, and cannot reasonably handle short-term preference changes while maintaining the stability of long-term preferences. Existing memory enhancement models usually treat all preferences as the same level without considering the hierarchical structure of the time dimension. Therefore, according to one aspect of this application, step S22 is further as follows:
[0042] S221, Extract preference expression time features; Read the preference time distribution mapping, and extract four key time features for each preference expression: frequency periodicity feature (the periodic pattern of preference appearance), duration stability feature (the duration of a single preference expression), repetition interval feature (the time interval between multiple occurrences of the same preference), and recent activity feature (the difference between the timestamp of the most recent preference expression and the current time), and combine these four features to form a preference time feature vector.
[0043] S222. Preference category time pattern clustering: Read the preference time feature vectors of all preferences, and use the K-means++ clustering algorithm to divide the preferences into multiple categories according to the time behavior pattern. Evaluate the optimal number of clusters through the silhouette coefficient and the Davies-Bouldin index to obtain the preference time behavior categories, where each category represents a set of preferences with similar time performance characteristics.
[0044] S223. Calculate the category-specific decay parameters: Read the preference time behavior categories, analyze the time feature distribution for each category, and use the maximum likelihood estimation method to determine the decay function form (exponential, power-law, or hyperbolic) and parameter range that best suits the category. Output the category decay parameter set, which includes the function type and parameter value range for each category.
[0045] S224. Construct an adaptive decay function: Read the category decay parameter set, set the weight coefficients and decay rate parameters, and construct a hybrid decay function template with adjustable parameters: D(t, c, p) = w1(c)*exp(-α(p)*t) + w2(c)*(1 / (1 + β(p)*t)) + w3(c)*(t (-γ(p)) ) where t represents the time variable, c represents the preference category, p represents the personalized parameter, w1, w2, and w3 are the weight coefficients (satisfying w1 + w2 + w3 = 1), and α, β, and γ are the decay rate parameters. This function combines the exponential decay, hyperbolic decay, and power-law decay modes, can adapt to the memory characteristics of different types of preferences, and outputs the decay function template for subsequent personalized parameter learning and tuning.
[0046] S225. Personalized parameter learning and tuning: Read the preference time feature vectors, preference time behavior categories, and decay function template, and use the Bayesian optimization algorithm to learn the personalized decay parameters of each user on the historical data, minimizing the error between the predicted preference intensity and the real user feedback to obtain the user-specific personalized decay parameters.
[0047] S226. Implement a dynamic adjustment mechanism: Construct a responsive adjustment mechanism to enable the decay function to dynamically update the parameters according to the user's latest interaction data. Read the personalized decay parameters and real-time user feedback data, and use the online learning method to continuously fine-tune the decay parameters so that the function can adapt to the changes in the user preference stability, and output an adaptive time decay model with self-adjustment ability.
[0048] S227. Decay model verification and calibration: Read the adaptive time decay model, verify the model performance on the retained test set of historical data, calculate the prediction accuracy, recall rate, and F1 score, and conduct a comparative analysis with the fixed decay rate model. Finally, calibrate the model parameters according to the verification results, and output the optimized adaptive time decay model that has been verified and calibrated. This model will be passed to step S23 for multi-scale time window analysis.
[0049] As Figure 2 shown, according to one aspect of the present application, the temporal preference hierarchy identification and classification includes the step of constructing an adaptive time decay model, and the step of constructing an adaptive time decay model includes:
[0050] Based on the cleaned multi-modal interaction sequences, construct a preference time distribution mapping, extract the time features of preference expressions, and generate preference time feature vectors; conduct a clustering analysis on the preference time feature vectors to obtain preference time behavior categories;
[0051] Calculate category-specific decay parameters for the preference time behavior categories to generate a set of category decay parameters; construct a mixed decay function based on the set of category decay parameters to generate a decay function template;
[0052] Conduct personalized parameter learning and tuning based on the preference time feature vectors, preference time behavior categories, and decay function templates to generate personalized decay parameters; implement a dynamic adjustment mechanism based on the personalized decay parameters to generate an adaptive time decay model for multi-scale time window analysis.
[0053] According to one aspect of the present application, the steps of extracting the time features of preference expressions and clustering analysis include:
[0054] Extract frequency periodicity features, duration stability features, repetition interval features, and recent activity features from the preference time distribution mapping; combine the frequency periodicity features, duration stability features, repetition interval features, and recent activity features to form preference time feature vectors; apply the K-means++ clustering algorithm to the preference time feature vectors to generate an initial clustering result; evaluate the optimal number of clusters through the silhouette coefficient and Davies-Bouldin index to generate preference time behavior categories.
[0055] In this embodiment, by constructing an adaptive time decay function, the problem that the traditional fixed decay rate model cannot adapt to the memory characteristics of different preference types is solved. It can accurately distinguish the time attributes of different preferences, assign the most suitable "memory retention period" to each preference, and improve the accuracy of temporal preference classification.
[0056] According to one aspect of the present application, step S3 is further:
[0057] S31. Construct a semantic-visual alignment space; read the standardized feature vector set, and construct a shared semantic space for the text modality and the image modality through contrastive learning methods, enabling preference expressions in different modalities to be compared in the same space, and output a cross-modal alignment mapping function.
[0058] S32. Conduct behavior-semantics correlation modeling; read the standardized feature vector set and the context correlation graph, analyze the correspondence between the user behavior modality and semantic expressions, establish a mapping model from the behavior sequence to the semantic intention, and output a behavior-semantics mapping matrix.
[0059] S33. Calculate the within-modal consistency; respectively read the standardized feature vector sets of each modality, calculate the consistency degree of preference expressions at different time points within the same modality, identify the preference conflicts and changes within the modality, and output the internal consistency scores of each modality.
[0060] S34. Measure the between-modal consistency; read the cross-modal alignment mapping function, the behavior-semantics mapping matrix, and the standardized feature vector sets of each modality, calculate the semantic consistency of preference expressions between different modalities, identify the expression differences between modalities, and output a between-modal consistency matrix.
[0061] S35. Locate and quantify the consistency conflicts; read the internal consistency scores and the between-modal consistency matrix, accurately locate the conflict points of preference expressions between different modalities, quantify the conflict degree and influence range, and output a cross-modal preference consistency report, including the conflict intensity, location, and possible explanations.
[0062] Through cross-modal preference consistency analysis in this embodiment, the limitations of traditional single-modal preference analysis are broken through, and the deep fusion of multi-modal preference signals is achieved. By constructing a semantic-visual alignment space, the preference expressions of the text and image modalities can be compared in the same semantic space; the behavior-semantics correlation modeling effectively solves the problem of understanding user behavior intentions; the within-modal consistency calculation can accurately identify the preference conflicts and changes within the same modality; the between-modal consistency measurement realizes the calculation of the semantic consistency of preference expressions between different modalities; the consistency conflict location and quantification technology can accurately locate and quantify the conflicts of preference expressions between different modalities, improve the system's ability to understand complex multi-modal signals, and reduce the preference misunderstanding rate caused by modal inconsistencies.
[0063] As Figure 3 shown, according to one aspect of the present application, step S31 is further as follows:
[0064] S311. Extract text semantic features and image semantic features from the standardized feature vector set;
[0065] S312. Based on the cleaned multi-modal interaction sequence, screen the time- and semantic-related text-image pairs to construct a selected parallel corpus;
[0066] S313. Construct a neural network model with a two-tower structure using text semantic features, image semantic features, and a selected parallel corpus for shared latent space learning to generate a cross-modal alignment mapping network;
[0067] S314. Perform semantic consistency verification and fine-tuning on the cross-modal alignment mapping network to generate an improved cross-modal alignment mapping network;
[0068] S315. Construct a dedicated adaptation layer for preference expression based on the improved cross-modal alignment mapping network to generate a preference-enhanced cross-modal mapping network, i.e., a semantic-visual alignment space;
[0069] S316. Evaluate and optimize the preference-enhanced cross-modal mapping network to generate a cross-modal alignment mapping function for inter-modal consistency measurement.
[0070] In an embodiment of the present application, modal feature extraction and initial representation are performed; a standardized feature vector set is read, and feature extraction is respectively performed on text modal and image modal data. For text data, a variant model of the pre-trained Bidirectional Encoder Representations from Transformers (BERT) is used to extract context-sensitive word vectors to obtain text semantic features; for image data, a pre-trained Vision Transformer (Transformer) model is used to extract visual features to obtain image semantic features. Both features are mapped to a 512-dimensional initial feature space, retaining the original semantic information.
[0071] Parallel data pair screening and construction are performed. The cleaned multi-modal interaction sequence is read, text-image pairs that are closely related in time are identified, and an initial parallel corpus is constructed. Then, a dual filtering mechanism based on temporal correlation and semantic correlation is designed to eliminate noise samples and retain high-quality text-image pairs, and the selected parallel corpus is output as the training data for cross-modal alignment.
[0072] Shared latent space learning. The text semantic features, image semantic features, and the selected parallel corpus are read, a neural network model with a two-tower structure is constructed, and each of the text tower and the image tower contains multiple layers of transformation and is finally mapped to the same-dimensional latent space. The model is trained using a cross-modal contrast loss function, and the training process is optimized through a random batch processing and hard negative sample mining strategy, and an initial cross-modal alignment mapping network is output.
[0073] Semantic consistency verification and fine-tuning are performed. The cross-modal alignment mapping network and the validation set data are read, a cross-modal retrieval task (bidirectional retrieval of text-image and image-text) is performed to evaluate the mapping quality, and metrics such as R@1, R@5, and average ranking are calculated. According to the verification results, a gradient quadratic sampling strategy is used to focus on learning difficult-to-align samples, and the mapping network parameters are fine-tuned to output an improved cross-modal alignment mapping network.
[0074] Construct a dedicated adaptation layer for preference expression. Read the improved cross-modal alignment mapping network. For the special properties of preference expression, construct a dedicated adaptation layer that can strengthen the semantic features related to preference and weaken the irrelevant features. Fine-tune with a small amount of labeled preference expression data to achieve the migration from the general semantic space to the preference-specific semantic space, and output a preference-enhanced cross-modal mapping network.
[0075] Conduct alignment space evaluation and optimization. Read the preference-enhanced cross-modal mapping network, construct three evaluation metrics: inter-modal alignment degree, semantic fidelity, and preference sensitivity, and conduct a comprehensive evaluation on real user data. According to the evaluation results, use the Bayesian hyperparameter optimization method to finally tune the network structure and training strategy, and output the final cross-modal alignment mapping function, which will be used for subsequent inter-modal consistency measurement.
[0076] According to one aspect of the present application, shared latent space learning includes the step of constructing a cross-modal contrast loss function, specifically:
[0077] Construct a global alignment loss component to ensure that the overall representation distributions of different modalities are similar; construct a local semantic consistency loss component to ensure that the distances of semantically similar content in the cross-modal space are within a preset range; construct a structure-preserving loss component to ensure that the semantic structure within the modality is retained after mapping;
[0078] Combine the global alignment loss component, the local semantic consistency loss component, and the structure-preserving loss component to form a cross-modal contrast loss function; use the cross-modal contrast loss function, text semantic features, image semantic features, and a selected parallel corpus to train a neural network model with a two-tower structure to generate a cross-modal alignment mapping network.
[0079] According to one aspect of the present application, the steps of constructing a dedicated adaptation layer for preference expression include:
[0080] Based on the improved cross-modal alignment mapping network, construct a dedicated adaptation layer for the special properties of preference expression to form a preference-specific adaptation layer; use the preference-specific adaptation layer to strengthen the semantic features related to preference and weaken the irrelevant features to generate enhanced preference expression features;
[0081] Fine-tune the preference-specific adaptation layer with a small amount of labeled preference expression data to obtain fine-tuning parameters; based on the improved cross-modal alignment mapping network, the preference-specific adaptation layer, and the fine-tuning parameters, achieve the migration from the general semantic space to the preference-specific semantic space, and generate a preference-enhanced cross-modal mapping network.
[0082] This embodiment solves the technical problem that the preference expressions of text and image modalities cannot be directly compared. It can compare and correlate the preference expressions of different modalities in a unified semantic space, improve the accuracy of cross-modal preference understanding, solve the problem that it is difficult to integrate preference information of different modalities in traditional methods, and lay a solid foundation for multi-modal consistency analysis.
[0083] According to one aspect of the present application, step S4 is further as follows:
[0084] S41. Identify micro-behavior patterns; read the behavior modality data stream in the cleaned multi-modal interaction sequence, and use a fine-grained behavior analysis algorithm to identify micro-behavior features such as the user's slight hesitation, repeated operations, and changes in residence time, which imply intentions, and output a micro-behavior feature set.
[0085] S42. Mine context condition associations; read the micro-behavior feature set and the context association graph, analyze the change rules of the user's behavior patterns under specific context conditions, identify the associations between environmental factors and preference expressions, and output a context-preference conditional probability table.
[0086] S43. Generate implicit preference hypotheses; read the micro-behavior feature set and the context-preference conditional probability table, generate candidate hypotheses about the user's implicit preferences based on the Bayesian inference framework, and assign an initial confidence level to each hypothesis, and output an implicit preference hypothesis set.
[0087] S44. Construct a low-interference verification strategy; read the implicit preference hypothesis set, and design a verification strategy that minimizes user interference for each implicit preference hypothesis to be verified, including indirect inquiries that naturally integrate into the conversation, alternative plan comparisons, and observing reactions to fine-tune parameters, etc., and output a low-interference verification plan.
[0088] S45. Verify execution and update hypotheses; execute the low-interference verification plan, collect verification feedback data, update the confidence level of the implicit preference hypothesis according to the user's reaction, confirm the high-confidence hypothesis as a valid preference, and output an implicit preference mapping table, which includes the semantic description, trigger conditions, and confidence score of the implicit preference.
[0089] Through implicit preference extraction and verification, this embodiment solves the major problem that traditional systems can only capture explicit preferences while ignoring implicit preferences. The micro-behavior pattern recognition technology can capture the subtle behavior characteristics of users, such as signals of implicit intentions like hesitation and repeated operations; the context-conditioned association mining analyzes the relationship between environmental factors and preference expressions, improving the context sensitivity of preference understanding; the implicit preference hypothesis generation method based on Bayesian inference enables the system to actively speculate on the potential preferences of users; the low-interference verification strategy design realizes the verification of implicit preference hypotheses without disturbing users; the verification execution and hypothesis update mechanism form a closed-loop implicit preference learning system, enabling the model to continuously optimize the understanding of users' implicit preferences and enhancing the system's ability to capture preference signals hidden in interaction behaviors.
[0090] During the interaction process, users usually only explicitly express part of their preferences, while a large number of preferences are implicit and not explicitly expressed. The preference capture module in existing solutions mainly relies on explicit feedback from users and cannot effectively capture implicit preferences, such as users' slight hesitation, repeated adjustment of the same parameters, and preference patterns in specific scenarios. There is currently a lack of a mechanism that can extract implicit preference signals from user interactions and confirm these implicit preferences through low-interference verification. Existing technologies either ignore implicit preferences or cause a decline in the user experience by frequently asking for verification. This is a characteristic problem because it involves how to accurately capture the subtle preference signals hidden in interaction behaviors without interrupting the user experience. Therefore, as Figure 4 shown, according to one aspect of the present application, step S44 is further as follows:
[0091] S441. Based on the cleaned multi-modal interaction sequence, generate an implicit preference hypothesis set, conduct verification requirement analysis and classification, and generate classified preference hypotheses;
[0092] S442. Based on the classified preference hypotheses, respectively construct a natural dialogue fusion verification template, a parameter fine-tuning reaction observation plan, and an alternative comparison strategy, and generate a dialogue fusion verification template library, a parameter fine-tuning verification plan set, and an alternative comparison verification strategy;
[0093] S443. Based on the dialogue fusion verification template library, the parameter fine-tuning verification plan set, and the alternative comparison verification strategy, construct a multi-round progressive verification plan and generate a time-sequential verification plan;
[0094] S444. Based on the time-sequential verification plan, construct a verification process adaptability control mechanism and generate verification control parameters;
[0095] S445. Integrate the dialogue fusion verification template library, the parameter fine-tuning verification plan set, the alternative comparison verification strategy, the time-sequential verification plan, and the verification control parameters to form a low-interference verification strategy for verification execution and hypothesis update.
[0096] In one embodiment of the present application, verification requirement analysis and classification are performed. The implicit preference hypothesis set is read, and the difficulty of each hypothesis is evaluated. Considering three dimensions of the interaction complexity required for verification, the user attention cost, and the certainty of the verification result, the hypotheses are classified into three categories: high-certainty hypotheses, medium-certainty hypotheses, and low-certainty hypotheses, and differentiated verification strategies are designed for different categories.
[0097] Construct a natural dialogue fusion verification template. For high-certainty hypotheses, a series of natural dialogue templates are designed so that the verification questions can be seamlessly integrated into the normal dialogue process. The dialogue history data of the user is read, the common language patterns and communication habits are extracted, and personalized dialogue templates are customized according to these features to ensure that the verification questions do not interrupt the dialogue rhythm or arouse the user's awareness, and a dialogue fusion verification template library is output.
[0098] Construct a parameter fine-tuning reaction observation scheme. For medium-certainty hypotheses, an indirect verification method based on minor parameter adjustments is designed. The relevant hypotheses in the implicit preference hypothesis set are read, and 3-5 minor parameter change schemes are designed for each hypothesis, with the change range controlled near the user's perception threshold. The hypotheses are verified by observing the user's reaction to these fine-tunings, and a parameter fine-tuning verification scheme set is output.
[0099] Construct an alternative comparison strategy. For low-certainty hypotheses, a verification strategy based on a limited number of alternatives is designed. The hypotheses to be verified in the implicit preference hypothesis set are read, and 2-3 opposing alternatives are constructed for each hypothesis, and a natural presentation method is designed to let the user make a choice without being aware of being tested, and an alternative comparison verification strategy is output.
[0100] Generate a multi-round progressive verification plan. The dialogue fusion verification template library, the parameter fine-tuning verification scheme set, and the alternative comparison verification strategy are read, and a multi-round progressive verification plan is designed to ensure that the verification process is dispersed in multiple interactions to avoid user fatigue caused by concentrated verification. According to the user's interaction frequency and habits, the best verification timing is assigned to each hypothesis, and a timing verification plan is output.
[0101] Control the adaptability of the verification process. A dynamic adaptation mechanism for the verification process is constructed, the real-time user feedback signals (such as response time, engagement, etc.) are read, and the verification intensity and frequency are automatically adjusted according to the user state. When user fatigue or inattention is detected, the verification frequency is automatically reduced or the verification is paused, and verification control parameters with adaptive capabilities are output.
[0102] Integrated low-interference verification solution. Integrate the dialogue fusion verification template library, parameter fine-tuning verification solution set, alternative solution comparison verification strategy, timing verification plan, and verification control parameters to form a complete verification solution. Conduct conflict checks and optimizations on the overall solution to ensure the coordination among various verification strategies, and output the final low-interference verification solution, which will be used for verification execution in step S45.
[0103] According to one aspect of the present application, the steps for constructing a natural dialogue fusion verification template include:
[0104] Screen high-certainty hypotheses from the classified preference hypotheses to form high-certainty hypotheses to be verified; analyze the user's dialogue history data, extract the user's common language patterns and communication habits, and form the user's personalized language features;
[0105] Based on the high-certainty hypotheses to be verified and the user's personalized language features, design a series of natural dialogue templates so that the verification questions can seamlessly integrate into the normal dialogue process, forming an initial dialogue template set; customize and adjust the initial dialogue template set according to the user's personalized language features to ensure that the verification questions do not interrupt the dialogue rhythm or arouse the user's awareness, and generate a dialogue fusion verification template library.
[0106] According to one aspect of the present application, the steps for constructing a parameter fine-tuning reaction observation solution include:
[0107] Screen medium-certainty hypotheses from the classified preference hypotheses to form medium-certainty hypotheses to be verified; design 3-5 minor parameter change solutions for the medium-certainty hypotheses to be verified to form a set of parameter change solutions;
[0108] Optimize the set of parameter change solutions to control the change range near the user's perception threshold to form an optimized set of parameter change solutions; construct indicators for observing the user's reaction to parameter changes to form a set of reaction observation indicators;
[0109] Combine the optimized set of parameter change solutions and the set of reaction observation indicators to form a parameter fine-tuning verification solution set for verifying hypotheses by observing the user's reaction to fine-tuning.
[0110] This embodiment solves the problem of interference with the user experience when verifying implicit preference hypotheses. It realizes the verification of implicit preference hypotheses with almost no interruption to the user's normal experience, reduces the verification cost and user disturbance, and at the same time maintains the reliability of the verification results, solving the problem of the decline in user experience caused by the traditional explicit inquiry verification method.
[0111] According to one aspect of the present application, step S45 is further:
[0112] S451. Perform verification task scheduling and execution based on a low-interference verification scheme and an implicit preference hypothesis set, and collect original verification feedback data;
[0113] S452. Extract and clean feedback signals from the original verification feedback data to generate structured feedback data;
[0114] S453. Establish a Bayesian evidence accumulation model based on the structured feedback data and the implicit preference hypothesis set to generate a hypothesis posterior probability table;
[0115] S454. Conduct multi-source evidence consistency evaluation on the hypothesis posterior probability table and the structured feedback data to generate an evidence consistency score;
[0116] S455. Make hypothesis confirmation and elimination decisions based on the hypothesis posterior probability table and the evidence consistency score to generate a hypothesis status update result;
[0117] S456. Generate hypotheses for incremental learning based on the hypothesis status update result and user interaction data to generate new implicit preference hypotheses;
[0118] S457. Construct and update an implicit preference mapping table based on the hypothesis status update result and the new implicit preference hypotheses to generate an implicit preference mapping table.
[0119] In an embodiment of the present application, verification task scheduling and execution are performed. Read the low-interference verification scheme and the implicit preference hypothesis set, and arrange the execution order and timing of verification tasks according to the timing verification plan. Trigger corresponding verification strategies at appropriate interaction nodes, collect user response data, including selection results, response times, and interaction patterns, etc., and output the original verification feedback data.
[0120] Extract and clean feedback signals. Read the original verification feedback data, apply signal processing techniques to extract effective feedback signals, and filter out noise and outliers. For text feedback, extract sentiment polarity and semantic orientation; for behavioral feedback, extract operation patterns and time characteristics; for selection feedback, extract preference intensity and certainty, and output the cleaned structured feedback data.
[0121] Construct a Bayesian evidence accumulation model. Construct a Bayesian evidence accumulation model, read the prior probabilities in the structured feedback data and the implicit preference hypothesis set, and calculate the posterior probabilities of each feedback data corresponding to the hypotheses. Adopt an incremental update mechanism to combine new feedback with historical evidence accumulation to form a continuously updated evidence chain, and output the hypothesis posterior probability table.
[0122] Perform multi-source evidence consistency evaluation. Read the hypothesis posterior probability table and structured feedback data to evaluate the evidence consistency from different verification strategies. When there are conflicts in the evidence from different sources, weighted processing is performed according to the reliability and timeliness of the evidence to solve the evidence conflict problem, and the evidence consistency score is output.
[0123] Make hypothesis confirmation and elimination decisions. Read the hypothesis posterior probability table and the evidence consistency score, and set dynamic thresholds to determine the status of the hypothesis. When the posterior probability exceeds the high threshold and the consistency score is good, the hypothesis is confirmed as a valid preference; when the posterior probability is below the low threshold or the consistency score is extremely poor, the hypothesis is eliminated; the remaining hypotheses remain in the to-be-verified state, and the hypothesis status update result is output.
[0124] Generate hypotheses for incremental learning. Based on the hypothesis status update result and the latest user interaction data, apply incremental learning techniques to generate new hypotheses to fill the gaps of the eliminated hypotheses. The new hypothesis generation process utilizes the feature patterns of the confirmed hypotheses to improve the quality and relevance of the hypotheses, and the newly added implicit preference hypotheses are output.
[0125] Construct and update the implicit preference mapping table. Integrate the confirmed hypotheses and the newly added implicit preference hypotheses in the hypothesis status update result to construct a complete implicit preference mapping table. This mapping table contains the semantic description, triggering conditions, confidence scores, and relevant contexts of each confirmed preference, sorted by confidence for use by subsequent modules. At the same time, the hypotheses that need further verification are looped back to step S44 to enter the next round of verification process.
[0126] According to one aspect of the present application, the steps of establishing a Bayesian evidence accumulation model and hypothesis confirmation and elimination decisions include:
[0127] Extract the prior probability of each hypothesis from the implicit preference hypothesis set to form the hypothesis prior probability; calculate the posterior probability of each feedback data corresponding to the hypothesis according to the structured feedback data and the hypothesis prior probability to generate the single-feedback posterior probability;
[0128] Adopt an incremental update mechanism to combine the single-feedback posterior probability with the historical evidence accumulation to form a continuously updated evidence chain, and generate a hypothesis posterior probability table; set dynamic thresholds based on the hypothesis posterior probability table and the evidence consistency score to generate a confirmation threshold and an elimination threshold;
[0129] Apply the confirmation threshold and the elimination threshold to judge the hypothesis. When the posterior probability exceeds the confirmation threshold and the consistency score is good, the hypothesis is confirmed as a valid preference. When the posterior probability is below the elimination threshold or the consistency score is extremely poor, the hypothesis is eliminated. The remaining hypotheses remain in the to-be-verified state, and the hypothesis status update result is generated.
[0130] This embodiment forms a closed-loop system for implicit preference learning. It can continuously learn and verify implicit preferences from users' daily interactions, forming a self-improving closed-loop learning system, improving the system's ability to capture and confirm implicit preferences, and solving the problems of insufficient understanding of implicit preferences or overly invasive verification methods in traditional systems.
[0131] Some of the preferences established by users in one domain (such as landscape image processing) may be applicable to other domains (such as portrait processing), while some are not. Existing solutions cannot automatically identify which preferences can be transferred across domains and which preferences need to be isolated within the domain, resulting in the system either over-generalizing preferences (wrongly applying specific domain preferences to irrelevant domains) or being overly conservative (unable to utilize transferable preference knowledge). There is currently a lack of a mechanism that can automatically identify the boundaries of preference transfer and accurately judge the applicable scope and transfer conditions of specific preferences. This involves how to extract implicit domain boundary information from historical interaction data and how to verify transfer hypotheses in new scenarios. This involves the representation and reasoning of domain knowledge, rather than simple memory storage and retrieval. Therefore, according to one aspect of this application, step S5 is further as follows:
[0132] S51. Extract application domain features: Read the cleaned multimodal interaction sequence and the multi-level temporal preference classification results, extract the domain features of different application scenarios, establish a domain feature vector, and output the application domain feature space.
[0133] S52. Calculate the similarity between domains: Read the domain feature vectors in the application domain feature space, calculate the similarity matrix between different application domains, identify the association strength between domains, and output the domain similarity matrix.
[0134] S53. Analyze the preference-domain relevance: Read the multi-level temporal preference classification results and the application domain feature space, analyze the applicability and performance differences of specific preferences in different domains, establish an association mapping between preferences and domain attributes, and output the preference-domain association table.
[0135] S54. Automatically identify the transfer boundary: Read the domain similarity matrix and the preference-domain association table, and identify under what conditions preferences can be transferred across domains and under what conditions they need to be isolated within the domain through a decision tree algorithm, and output the definition of the preference transfer boundary.
[0136] S55. Generate and verify transfer rules: Based on the definition of the preference transfer boundary, generate specific cross-domain preference transfer rules, and verify the effectiveness of these rules through historical data simulation, adjust the rule parameters, and output the final cross-domain preference transfer rule set.
[0137] This embodiment solves the problem of migrating preference knowledge between different application fields. It improves the system's generalization ability for user preferences, avoids the dilemma of over-generalization and over-conservatism, and enables users to enjoy the convenience of established preferences in new scenarios.
[0138] In one embodiment of the present application, automatically identifying the migration boundary specifically includes: analyzing the differences in domain features; reading the domain similarity matrix and the application domain feature space, calculating the feature distances between different domains, using the Mahalanobis distance to measure the differences in the multi-dimensional feature space, identifying the key distinguishing features, and outputting the domain distinguishing feature set and the feature difference weight table.
[0139] Evaluating the preference sensitivity. Reading the preference-domain association table and the multi-level time-series preference classification results, calculating the performance change sensitivity of each preference in different domains. Constructing a sensitivity scoring function: S(p, d1, d2) = |P(p|d1) - P(p|d2)| / max(P(p|d1), P(p|d2)); where p represents the preference, d1 and d2 represent two different domains, and P(p|d) represents the applicable probability of the preference p in the domain d. Calculating the sensitivity of all preferences between all domain pairs and outputting the preference sensitivity matrix.
[0140] Selecting the migration decision tree features. Reading the domain distinguishing feature set, the feature difference weight table, and the preference sensitivity matrix, constructing a feature selection algorithm based on information gain, and identifying the most critical feature subset for the preference migration decision. Calculating the information gain of each feature: IG(f) = H(T) - Σ(v∈V(f)) |T_v| / |T| * H(T_v); where f is the feature, T is the training set, V(f) is the value set of the feature f, H represents the information entropy, and T_v is the data subset corresponding to the value v of the feature f in the training data. Selecting the feature subset with the highest information gain and outputting the decision key feature set.
[0141] Learning the migration condition rules. Reading the decision key feature set, the preference sensitivity matrix, and the historical successful cases of preference migration, constructing a decision tree model to learn the migration condition rules. Constructing an initial decision tree, where each internal node represents a test on a certain feature, and each leaf node represents a migration decision (transferable / non-transferable), and outputting the initial migration condition decision tree.
[0142] Construct a migration risk assessment model. Read the migration condition decision tree and historical user feedback data to construct a migration risk assessment model and quantify the potential negative impact of incorrect migration. Design a risk scoring function: R(p, d1, d2) = I(p) * D(d1, d2) * (1 - C(p, d1, d2)); where I(p) is the preference importance, D(d1, d2) is the domain distance, and C(p, d1, d2) is the confidence of migration. Calculate the risk score for each potential migration decision and output a migration risk scoring table.
[0143] Perform decision boundary optimization and pruning. Read the migration condition decision tree and the migration risk scoring table, and apply pruning and boundary optimization techniques to improve the generalization ability of the decision tree. Use cost-complexity pruning to reduce overfitting, and at the same time adjust the decision threshold according to the risk score, setting more conservative migration conditions for high-risk nodes, and output an optimized migration decision rule tree.
[0144] Visualize and verify the migration boundary. Read the migration decision rule tree and the domain similarity matrix, and visualize the migration boundary in the feature space for easy understanding and verification. Evaluate the boundary accuracy through randomly generated test cases, calculate the precision, recall, and F1 score, and make final adjustments according to the verification results, and output a verified definition of the preference migration boundary, which will be used for generating migration rules in step S55.
[0145] This embodiment solves the key problem of the precise migration of preference knowledge between different domains through automatic identification of the migration boundary. The breakthrough enables the system to accurately judge under what conditions preferences can be migrated across domains and under what conditions domain isolation is required, avoiding the dilemma of overgeneralization and overconservatism, and improving the system's ability to reuse existing preference knowledge in new scenarios.
[0146] In multimodal interaction, users express preferences through different channels such as text, images, and interaction behaviors, and the signals of these channels sometimes appear inconsistent or even contradictory. For example, a user expresses the hope that an image is "more natural" in text, but always selects a high-saturation effect in interaction. Existing solutions cannot effectively handle this problem of multimodal signal inconsistency, resulting in deviations in the system's understanding of the user's true preferences. There is currently a lack of a mechanism that can detect, quantify, and coordinate the inconsistency of multimodal preference signals, and it is impossible to extract the user's true intention from contradictory signals. This involves how to design a cross-modal preference consistency metric standard and how to establish a reliable mapping relationship between different modalities. This involves the in-depth semantic understanding and fusion of multimodal information, rather than preference processing within a single modality. Therefore, according to one aspect of this application, step S6 is further as follows:
[0147] S61. Detect preference conflicts; read the multi-level temporal preference classification results, cross-modal preference consistency reports, and implicit preference mapping tables to comprehensively detect the conflict points between preferences at different levels, different modalities, and different types (explicit / implicit), and output a preference conflict graph.
[0148] S62. Construct a preference priority matrix; construct a multi-dimensional preference priority scoring system according to the temporal preference conversion rule set, the confidence levels in the cross-modal preference consistency report, and the implicit preference mapping table, assign dynamic weights to each preference, and output a preference priority matrix.
[0149] S63. Construct context-aware conflict resolution strategies; read the preference conflict graph, the preference priority matrix, and the current application context data, design a set of context-based conflict resolution strategies, adopt different coordination methods for different types of conflicts, and output a set of conflict resolution strategies.
[0150] S64. Perform conflict coordination execution and verification; apply the set of conflict resolution strategies to handle the conflicts in the preference conflict graph, generate a coordinated preference set, and simulate and verify the consistency and user satisfaction of the coordination results through historical data, and output a conflict coordination report.
[0151] S65. Generate a unified preference decision; read the conflict coordination report and all non-conflicting preferences, integrate them to form a final unified preference decision, ensure the consistency and coherence of the decision results in all dimensions, and output the unified preference decision result.
[0152] This embodiment solves the problem of multiple preference conflicts in complex scenarios. It can make decisions that conform to the true intentions of users in complex multi-preference conflict scenarios, improve the user experience and satisfaction, and solve the problems of chaotic decision-making or simply taking the latest preference when traditional systems face preference conflicts.
[0153] In an embodiment of the present application, constructing context-aware conflict resolution strategies specifically includes: classifying conflict types and extracting features; reading the preference conflict graph, analyzing the structural features of the conflicts, and classifying the conflicts into four basic types: temporal hierarchy conflicts, modal expression conflicts, explicit / implicit conflicts, and domain applicability conflicts. Extract specific feature vectors for each conflict type, including conflict intensity, the number of preferences involved, historical resolution patterns, etc., and output a conflict type feature table.
[0154] Construct a historical conflict resolution case library. Read historical user interaction data and preference update records, identify past successfully resolved conflict cases, extract conflict features and resolution methods, and construct a historical conflict resolution case library. Represent the cases in a structured manner, including fields such as conflict descriptions, resolution strategies, and user satisfaction scores, to form the basis for case retrieval.
[0155] Perform situational feature extraction and representation. Read the current application context data, extract key situational features, including task type, time pressure, user status, and environmental conditions, etc. Design a situational representation model to map multi-dimensional situational information into vector representation for subsequent policy matching, and output the current situational feature vector.
[0156] Perform case similarity calculation and retrieval. Read the conflict type feature table, historical conflict resolution case library, and the current situational feature vector, and construct a multi-factor similarity calculation formula: Sim(c1, c2) = w1 * SimType(c1, c2) + w2 * SimContext(c1, c2) + w3 * SimUser(c1, c2); where SimType represents the conflict type similarity, SimContext represents the situational similarity, SimUser represents the user feature similarity, w1, w2, and w3 are weight coefficients (satisfying w1 + w2 + w3 = 1), and c1, c2 are two different conflict cases. Retrieve the most matching historical cases according to the similarity, and output the set of similar cases.
[0157] Generate situation-adaptive strategies. Read the set of similar cases and the current situational feature vector, and generate solution strategies suitable for the current situation through case adaptation technology. Make situation-adaptive adjustments to the retrieved historical strategies, considering factors such as the current user status, task importance, and time limit, etc., and output the set of candidate solution strategies.
[0158] Perform policy utility evaluation and optimization. Construct a multi-dimensional policy utility evaluation framework, and calculate the expected utility for each policy in the set of candidate solution strategies: U(s) = Σ(i = 1 to n) w_i * u_i(s); where s is the policy, u_i is the scoring function of the i-th utility dimension, and w_i is the weight. The evaluation dimensions include user satisfaction, decision consistency, execution efficiency, and cognitive load. Sort and screen the policies according to the utility scores, and output the optimized policy ranking table.
[0159] Perform adaptive policy combination and integration. Read the optimized policy ranking table and the preference conflict graph, analyze the dependency relationships between conflicts, and design an overall coordination plan. For interrelated conflicts, ensure that consistent solution strategies are adopted; for independent conflicts, select the optimal solution strategies for each. Integrate all strategies into a unified solution, and output the complete conflict resolution strategy set, which will be used for conflict coordination execution in step S64.
[0160] In this embodiment, through the context-aware conflict resolution strategy, the complex problem of different coordination methods being required for preference conflicts in different scenarios is solved. It can dynamically select the most suitable conflict resolution strategy according to the specific context instead of adopting fixed conflict resolution rules, improving the flexibility and accuracy of conflict coordination and solving the problem of rigid conflict handling in traditional systems in complex contexts.
[0161] In one embodiment of the present application, the process of generating a unified preference decision is specifically as follows: Integrate the conflicted preference set; Read the conflict coordination report and all non-conflicted preferences, merge the coordinated conflicted preferences with the original non-conflicted preferences to form a preliminary unified preference set. Check the logical consistency within the set to ensure there are no remaining contradictions, and output the preliminary unified preference set.
[0162] Recalculate the preference priorities. Read the preliminary unified preference set and the preference priority matrix, and recalculate the priorities of the preferences according to the adjustments during the conflict coordination process. Consider the changes in weights during the conflict resolution process and the new context conditions, update the priority scores of each preference, and output the updated priority preference set.
[0163] Identify the decision constraint conditions. Read the updated priority preference set and the current application context data, and identify the constraint conditions that the decision must satisfy, including system capacity limitations, resource constraints, and explicit bottom-line requirements of the user, etc. Formalize these constraints as boundary conditions in the decision space, and output the decision constraint condition set.
[0164] Generate a preference implementation plan. Read the updated priority preference set and the decision constraint condition set, and design a preference implementation plan that meets the constraints. For multiple preferences that cannot be satisfied simultaneously, determine the execution order according to the priorities; for preferences that can be partially satisfied, determine the degree of satisfaction and the parameter values, and output the preference implementation plan.
[0165] Conduct cross-dimensional consistency checks. Read the preference implementation plan, construct a multi-dimensional consistency check mechanism to ensure the coherence of the decision in multiple dimensions. The checked dimensions include temporal consistency (continuity of decisions at different time points), semantic consistency (compatibility of different expression methods), and functional consistency (coordination of different functional domains), identify and repair potential inconsistencies, and output the consistency check report.
[0166] Conduct parameter fine-tuning and optimization. Read the preference implementation plan and the consistency check report, and perform fine-tuning at the parameter level for the detected inconsistencies. Use the simulated annealing algorithm to search for the optimal solution in the parameter space, and fine-tune the specific parameter values while maintaining the preference intention to make the overall decision more harmonious and consistent, and output the optimized parameter solution.
[0167] Generate and encapsulate the final decision. Integrate and optimize the parameter scheme and the high-priority preferences in the updated priority preference set to generate the final unified decision. Encapsulate the decision into a standard format, including the core decision content, implementation parameters, priority information, and explanatory text, which is convenient for system execution and user understanding, and output the unified preference decision result, which will be used for the update of the memory-enhanced preference model in step S7.
[0168] In this embodiment, by generating a unified preference decision, the output of the final decision that is consistent in multiple dimensions and meets the constraints is realized. It can generate consistent, coherent, and constraint-compliant decisions in a complex multi-preference environment, avoiding problems such as decision fragmentation, contradictions, or violations of constraints in traditional systems, and providing an overall solution that conforms to the user's true intention.
[0169] According to one aspect of the present application, step S7 is further as follows:
[0170] S71. Extract the differential increment; Read the unified preference decision result and the existing user preference model, calculate the difference between the old and new models, extract the meaningful changed parts, and output the preference increment data.
[0171] S72. Construct a memory optimization storage strategy; Read the preference increment data and the multi-level time-series preference classification results, design a differentiated storage strategy according to the time-level attributes of the preferences, allocate different storage priorities and update frequencies for different types of preferences, and output the memory optimization storage scheme.
[0172] S73. Update the preference model increment; Integrate the preference increment data into the existing preference model according to the memory optimization storage scheme, adopt an incremental update mechanism to avoid catastrophic forgetting, and maintain the stability and coherence of the model, and output the updated user preference model.
[0173] S74. Construct a personalized parameter mapping; Read the updated user preference model, establish a mapping relationship between the user's semantic expression and the specific operation parameters, form a personalized parameter converter, and output the personalized parameter mapping table.
[0174] S75. Generate and apply the recommendation strategy; Based on the updated user preference model and the personalized parameter mapping table, generate a personalized recommendation strategy that meets the user's current needs and long-term preferences, apply it to the interaction interface and content generation, output the personalized recommendation result, and collect user feedback for the update of the next-round preference model.
[0175] This embodiment realizes the long-term accumulation and optimized utilization of preference knowledge. It can continuously learn and improve the understanding of user preferences like humans, maintain the long-term stability of memory, and at the same time allow appropriate updates and adjustments, solving the problems of untimely update of the preference model or loss of important information during the update process in traditional systems.
[0176] In an embodiment of the present application, the process of updating the preference model increment is specifically as follows: Construct an incremental data structure; Read the preference incremental data and construct a dedicated incremental data structure, which can efficiently represent preference changes and includes fields such as old value, new value, change type (newly added, modified, deleted), timestamp, and confidence. Organize the incremental data into a tree structure to reflect the hierarchical relationship between preferences, and output a structured preference incremental tree.
[0177] Analyze the scope of update impact. Read the preference incremental tree and the existing user preference model, and analyze the scope of impact of each incremental change on the model. Identify directly and indirectly affected model components through dependency graph analysis, predict possible chain reactions, and output an update impact map.
[0178] Perform catastrophic forgetting prediction and prevention. Read the update impact map and apply a feature retention detection algorithm to predict areas where catastrophic forgetting may occur. Design an important feature protection mechanism: L_protect = λ * ||M_new(x_i) - M_old(x_i)|| 2 ; where M_new and M_old are the models before and after the update respectively, x_i is an important sample, and λ is a protection strength parameter. Integrate the protection mechanism into the update process and output a forgetting prevention strategy.
[0179] Perform progressive update of knowledge distillation. Read the existing user preference model, preference incremental tree, and forgetting prevention strategy, and design a progressive update method based on knowledge distillation. Use the original model as the teacher model and the new model as the student model, and transfer stable knowledge through the distillation loss function: L_distill = α*L_new + (1-α)*KL(P_teacher||P_student); where L_new is the learning loss of new data, KL is the KL divergence, α is a balance parameter, P_teacher is the probability distribution of the teacher model, and P_student is the probability distribution of the student model. Gradually integrate new knowledge through this mechanism and output a distillation update plan.
[0180] Perform balance adjustment of stability and plasticity. Construct an adaptive learning rate adjustment mechanism to balance the stability and plasticity of the model. Read the distillation update plan and the historical interaction stability analysis of the user, and dynamically adjust the learning rate according to the historical stability of the user's preferences: η_t = η_base * exp(-β*S(t)); where η_t is the learning rate at time t, S(t) is the preference stability index, β is a decay parameter, and η_base is the initial learning rate. Use a lower learning rate in areas with high stability to ensure slow updates; use a higher learning rate in areas with high plasticity to allow for quick adaptation, and output adaptive learning parameters.
[0181] Perform incremental update execution and monitoring. Execute model update operations according to the distillation update scheme and adaptive learning parameters. Adopt a phased update strategy, first update the model components with less impact, and then gradually update the core components. Monitor the changes in performance metrics during the entire update process. Immediately roll back once an anomaly is detected, and output the update execution log and the preliminary updated model.
[0182] Perform post-update consistency verification and repair. Read the updated model, design a multi-dimensional consistency verification scheme, and test the model performance on the retained dataset. Calculate the decision consistency index before and after the update: C = |{x | M_new(x) == M_old(x), x ∈ X_stable}| / |X_stable|; where X_stable is the sample set that should remain stable. For regions with low consistency, apply repair strategies for fine-tuning to ensure the continuity of key functions. Finally, output the verified and repaired updated user preference model, which will be used for personalized parameter mapping construction in step S74.
[0183] This embodiment solves the catastrophic forgetting problem that may occur during the model update process by updating the preference model incrementally. It can maintain the stability of existing preference knowledge while continuously absorbing new preference knowledge, solving the dilemma of traditional systems that either cannot fully absorb new knowledge or lose important old knowledge during the update process, providing technical support for long-term and stable preference learning.
[0184] In another embodiment of the present application, a user preference parsing system based on a memory-augmented multi-modal large model includes three main modules: a memory module, a content generation module, and a preference capture module. The preference capture module captures the user's conversation habits and image enhancement preferences and saves them in the memory module. The content generation module performs conversation generation and image enhancement based on the user's conversation and the memory in the memory module.
[0185] The memory module includes two parts: a preference memory bank and a memory retriever. The preference memory bank stores the mapping between the current user conversation and its corresponding image enhancement process, which is automatically saved in the memory bank through the preference capture module during the conversation with the user and is continuously updated as the interaction with the user progresses. The memory retriever includes a text classifier (TextClassifier) and a dense retrieval encoder (Dense Retrieval Encoder). The classifier is used to determine whether there is an intention of image enhancement in each conversation input of the user, and the dense retriever is used to match the saved memories to extract the user preferences in the memories.
[0186] The content generation module includes a multi-modal large language model backbone (MLLM) and an image enhancement toolbox. The multi-modal large language model backbone is used to implement continuous natural multi-modal conversations with users. The image enhancement toolbox contains a variety of mainstream image enhancement tools and can be called sequentially according to instructions to achieve more complex and flexible image enhancement effects.
[0187] The preference capture module is connected to the memory module and reuses the multi-modal large language model in the content generation module. The preference capture module uses the text classifier in the memory module to determine whether the user input dialogue contains an image enhancement intention. Then, through the multi-modal large language model combined with the image enhancement tools in the toolbox, a candidate image enhancement process is generated, and the user is interacted with to confirm the user's real needs. Finally, the memory library is updated to complete the capture of the user's preferences. At the same time, after each enhancement is completed, the user feedback is obtained through the dialogue to adjust the enhancement process.
[0188] During the entire running process of the algorithm, the preference capture module captures the user's personal preferences in the continuous dialogue and saves them in the memory module. The memory module is responsible for extracting the correct memory in the dialogue and mapping the user input to its personalized enhancement method. While the content generation module is responsible for the dialogue, it receives the enhancement method extracted by the memory module and performs image enhancement processing.
[0189] According to one aspect of the present application, a method for parsing user preferences based on a memory-enhanced multi-modal large model includes the following steps:
[0190] Step 1: First, determine whether the user's input contains an image enhancement intention. The classifier of the retriever in the memory library is a text classification model, which is based on the RoBERTa (Robustly Optimized BERT Pretraining Approach) model represented by a bidirectional encoder. The classification (CLS) vector of its encoded text is fed into a feed-forward neural network (FFN) model to perform a binary classification task to determine whether the dialogue contains an image enhancement intention.
[0191] Step 2: If not, use the multi-modal language model backbone in the content generation module to conduct a regular conversation; if it contains, check whether the corresponding memory is empty, which is completed through the dense retrieval model in the retriever. This is a text encoder that encodes the text into a vector in the retrieval space. Similar texts have higher similarity. In the present invention, when the dot product value of two vectors is greater than 0.5, the two are matched. If multiple texts in the memory library meet the condition, the maximum value is taken; otherwise, they do not match and are determined to be empty memory.
[0192] Step 3: If the memory is empty, the system will use the multi-modal language model backbone in the content generation module and combine it with the tools in the image enhancement toolbox to infer the user's possible enhancement intentions, generate several candidate enhancement processes, and ask the user for confirmation to obtain the accurate user intention and establish a mapping relationship to update the memory bank.
[0193] Step 4: If the memory is not empty, directly match the most similar process memory and perform enhancement processing.
[0194] Step 5: Finally, the system will ask for the user's satisfaction. If the user is satisfied, it will end; otherwise, the language model will judge the adjustment direction based on the user's input, adjust the image enhancement process, and update the memory bank.
[0195] Step 6: Perform image enhancement again according to the adjusted process and execute Step 5.
[0196] In an embodiment of image enhancement in this application, the user: <Picture> I think this picture is too dark; judge whether there is an intention: yes; obtain the toolbox process; generate an inquiry text. System: Do you want to perform the following operations? 1. Increase brightness; 2. Perform low-light enhancement; 3. Have no modification intention. User: I choose 1. Analyze the input; obtain the process [{Brightness: +10}]; store the memory:
I think this picture is too dark: [{Brightness: +10}]
I think this picture is too dark: [{Brightness: +25}, {Noise reduction: 30}]
I think this picture is too dark: [{Brightness: +25}, {Noise reduction: 30}]
[0197] In this embodiment, when the memory bank is empty, 1. User input: I think this picture is too dark ( For the picture placeholder token). 2. The text classifier in the "I think this picture is too dark" retriever determines that the text may contain an image processing intention. 3. Determine whether the memory bank is empty. It is. 4. Call the preference capture module. The module obtains the saved fixed prompt words and has a conversation with the backbone of the large language model to obtain the query text: Example of prompt words: According to the user's description of the picture problem, combined with the existing list of image enhancement tools, generate a concise text asking about the user's intention. Execute according to the following steps: Semantic analysis: Identify the keywords in the user's description; Tool matching: Screen all relevant tools from the tool library (allow 1-3 relevant items); Option generation: List the matching items with digital numbers, keeping the original names of the tools; Exclusion rule: Ignore tools irrelevant to the user's question; Output format: Use a friendly question template: "Do you want to perform the following operations?" + List items on new lines, and finally add an option of no intention; The tool list is [Brightness adjustment, Low-light enhancement, Denoising, Super-resolution]; The user input is I think this picture is too dark. 5. System output: Do you want to perform the following operations? (1). Increase brightness; (2). Perform low-light enhancement; (3). Have no modification intention. 6. User input: I choose 1. 7. Obtain the user's response text, call the large language model again for judgment, and obtain the accurate answer. Example of prompt words: Determine the user's needs based on the user input, and strictly output the name of a tool; The options include (1). Increase brightness; (2). Perform low-light enhancement; (3). Have no modification intention; The user input is I choose 1. 8. The large language model judges and obtains: (1). Increase brightness. 9. According to the output tool name, match the enhancement process in the image enhancement toolbox (preset some processes at the first startup), increase brightness, corresponding to [{Brightness: +10}]. 10. Save the mapping of the user input and the enhancement process in the memory bank: "I think this picture is too dark" [{Brightness: +10}], and process the picture. If the output is no intention, it means the user has no enhancement intention, and store it in the memory bank I think this picture is too dark - None. 11. Encode the user input in the memory bank using a dense retrieval encoder. 12. After processing, output the picture and output a text querying the effect. The system outputs: Are you satisfied? <The processed picture>. 13. The user inputs: "Overall it is too dark, especially the face of the person, and there are also a small amount of noise points", obtain the user feedback text, and call the large model for judgment. The example is as follows: Prompt: According to the user's feedback on the image enhancement effect, perform the following operations: (1). Satisfaction judgment: If the feedback contains a clear affirmation (such as "satisfied" / "good" / "ok") or no modification opinion, directly output: "satisfied". (2). Parameter adjustment rules (only when not satisfied): a) Adjust based on the given current process: Current process: [{Brightness: +10}]; b) Adjustment rules: 1) Modify the parameters in the process according to the user's feedback; 2) Add the following process according to the user's feedback: {{The processes to be selected in the toolbox}}; (3). Output format: When not satisfied, output a list of dictionaries: [{"step": value},...], and obtain the adjusted process: [{Brightness: +25}, {Noise reduction: 30}]. 14. Update the memory bank with the new process, replacing [{Brightness: +10}] with [{Brightness: +25}, {Noise reduction: 30}]. 15. Ask again and loop until "satisfied" is output.
[0198] When the memory bank is not empty, 1. When there is memory in the memory bank, the dense retriever encodes the input and performs a dot product with the data in the memory bank. If the dot product value < 0.5, the retrieval fails; if > 0.5, it is successful. If there are multiple successful results, take the maximum dot product value. 2. If the retrieval fails, repeat step 1; if the retrieval is successful, extract the mapped process for enhancement. If the mapping is unintentional, no processing is performed.
[0199] Among them, the text classifier is trained for classification using a large number of regular conversations and conversations with intentions. The dense retriever can use the retriever in the existing RAG, as long as it can match the text similarity, or it can also be retrained through contrastive learning. The large language model can use the COT model of R1, o1 to enhance the effect. The large language model can be fine-tuned by instructions to obtain better process optimization capabilities, and basically no training is required for the rest of the judgment tasks. The image enhancement toolbox can be a multi-layer structure, such as requirement - tool - process, which is gradually more refined and accurate. The process format of this embodiment is schematic and can be more complex. In addition to processing images, this application can also process audio, video, text, etc.
[0200] This embodiment utilizes the backbone of a multimodal large language model and a memory module to achieve seamless connection between colloquial conversations and professional image processing processes. The memory bank in the memory module stores a large amount of memory data mapping natural language and image enhancement processes. These data are collected and updated with the help of the preference capture module during conversations, enabling precise matching of the user's vague and popular input in each round to the corresponding complex image enhancement scheme. This allows users to obtain high-quality image processing effects without having to master professional terms, simply through daily conversations. This embodiment has the ability to understand the image enhancement intentions behind different expressions of different users. The built-in large language model of the system can understand various language styles and expressions. The preference capture module captures and analyzes the user's personalized preferences through continuous interaction with the user. When parsing diverse and complex conversations, it accurately records the user's expression characteristics and changing needs. The retriever in the memory module can quickly extract the historical record most matching the current input from the memory bank during the interaction, achieving precise recognition and response to the user's habits. The memory bank of the system can store the user's preference information for a long time. During each interaction process, the memory bank automatically records the user's language habits, expression characteristics, and their corresponding specific needs. During continuous interactions, the memory bank continuously receives information from the preference capture module to update the memory bank, gradually building a complete user profile. As the usage time increases, the system's understanding of the user deepens, thereby providing more personalized and customized image enhancement services in subsequent interactions, not only improving the accuracy of responses but also optimizing the user experience.
[0201] The present invention effectively solves the coordination problem of multi-level temporal preference conflicts through a temporal preference hierarchy identification and classification module and a multi-level preference conflict coordination and decision-making module. Specifically: The system uses interaction frequency and duration analysis to identify the temporal distribution characteristics of preferences and establish a preference temporal distribution mapping. Then, an adaptive time decay function is used to dynamically adjust the decay rate for different types of preferences, achieving a differentiated "memory retention period". Multi-scale time window analysis analyzes preference stability at three scales: short-term window, medium-term window, and long-term window, while hierarchical clustering precisely classifies preferences into short-term temporary preferences, medium-term task preferences, and long-term stable preferences. Temporal preference transition boundary detection further analyzes the conditions and triggering factors for preference transitions between different levels. When facing preference conflicts, the preference conflict detection in the module can comprehensively identify conflicts between different levels of preferences; the preference priority matrix assigns dynamic weights to each preference based on temporal characteristics; the context-aware conflict resolution strategy dynamically selects the best solution according to the specific scenario; and finally, unified preference decision generation ensures that the decision results are consistent in all dimensions. This enables the system to intelligently judge the priority according to the specific situation when there is a conflict between "high-saturation" temporary needs and "natural-saturation" long-term preferences, and can accurately determine when the temporary preferences should expire.
[0202] The problem of capturing implicit preferences and low-interference verification is solved by the implicit preference extraction and verification module: Micro-behavior pattern recognition can capture the behavioral characteristics of implicit intentions such as the user's slight hesitation and repeated operations; Contextual condition correlation mining analyzes the variation rules of the user's behavior patterns in specific contexts; Implicit preference hypothesis generation generates and preliminarily evaluates implicit preference hypotheses based on the Bayesian inference framework. The key lies in the design of the low-interference verification strategy and the verification execution and hypothesis update. After classifying the verification requirements, three low-interference verification methods are designed: Natural dialogue fusion verification seamlessly integrates verification questions into normal conversations; Parameter fine-tuning response observation indirectly verifies through small parameter changes; Alternative comparison strategy allows users to make choices without being aware of being tested. The multi-round progressive verification plan and the adaptive control of the verification process ensure that the verification process is decentralized and dynamically adjusted according to the user's state. The credibility of the hypothesis is continuously updated through the Bayesian evidence accumulation model and the multi-source evidence consistency evaluation, forming a closed-loop learning system, improving the system's ability to capture and verify implicit preferences, and avoiding interfering with the user experience at the same time.
[0203] The problem of accurately identifying the cross-domain transfer boundary of preferences is solved by the cross-domain preference boundary identification and transfer rule generation module: Application domain feature extraction and inter-domain similarity calculation provide a basis for identifying domain relationships; Preference-domain relevance analysis deeply explores the relationship between preferences and domain attributes. The core lies in the automatic identification of the transfer boundary, which calculates the feature distance between different domains through domain feature difference analysis; Preference sensitivity assessment quantifies the performance changes of preferences in different domains; Transfer decision tree feature selection and transfer condition rule learning establish an accurate transfer judgment mechanism; Transfer risk assessment model quantifies the potential negative impact of incorrect transfers; Decision boundary optimization and pruning improve the generalization ability of the decision tree; Transfer boundary visualization and verification ensure the accuracy of the boundary definition. Finally, transfer rule generation and verification generate specific rules based on the transfer boundary definition and verify their effectiveness. This enables the system to accurately judge when the preference for processing landscape photos can be transferred to portrait processing, avoiding the dilemma of over-generalization and over-conservatism.
[0204] The problem of inconsistent multimodal preference signals is solved by the cross-modal preference consistency analysis module: The semantic-visual alignment space construction realizes the expression of text and image modalities in a unified semantic space through steps such as modal feature extraction, parallel data pair screening, contrast loss function design, and shared latent space learning. The behavior-semantic association modeling establishes the mapping between user behavior and semantic intention. Through intra-modal consistency calculation and inter-modal consistency measurement, the system can accurately identify the differences in preference expressions within the same modality and between different modalities. The consistency conflict location and quantification further precisely locate the conflict points and quantify the conflict degree. When the user expresses in the text the hope that the image is "more natural" but the behavior prefers high saturation, the system can deeply understand the contradiction between different modal signals through these technologies, and combined with the conflict coordination mechanism, comprehensively analyze the user's true intention. This multi-modal consistency analysis improves the system's ability to understand complex preference signals and reduces the preference misunderstanding rate caused by modal inconsistency.
[0205] The present invention constructs a complete memory-enhanced multi-modal preference parsing closed-loop system. From multi-modal data acquisition and processing to temporal preference classification, from cross-modal consistency analysis to implicit preference mining, from cross-domain boundary recognition to conflict coordination decision-making, and then to memory-enhanced model update, a complete technical chain of data → model → application → optimization is formed. It can capture explicit and implicit preferences, handle cross-modal and cross-domain preferences, coordinate multi-level preference conflicts, and has long-term memory and progressive learning capabilities. The present invention breaks through the limitations of traditional preference parsing, realizes a preference understanding mechanism closer to human cognition, improves the personalized service quality of intelligent systems, solves the core challenges faced by current preference parsing systems, and provides a technical foundation for constructing a truly user-centered artificial intelligence system.
[0206] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A user preference parsing method based on a memory-enhanced multimodal large model, characterized in that: The following steps are involved: Collect and preprocess multimodal interaction data to generate standardized feature vector sets and cleaned multimodal interaction sequences; The multimodal interaction data are text and pictures; Based on the cleaned multimodal interaction sequence, the temporal preference hierarchy is identified and classified to generate a multi-level temporal preference classification result and a temporal preference conversion rule set; including: reading the cleaned multimodal interaction sequence, calculating the frequency, duration and repetition pattern of each preference expression, establishing the temporal distribution characteristics of the preference, and outputting the preference time distribution map; constructing a nonlinear time decay function based on the preference time distribution map, and outputting an adaptive time decay model; reading the cleaned multimodal interaction sequence, using a multi-scale sliding window technology, analyzing the stability and change trend of preferences at the hourly, daily and monthly levels, and outputting a multi-scale preference stability matrix; reading the multi-scale preference stability matrix and the data processed by the adaptive time decay model, using a hierarchical clustering algorithm to divide the preference expressions into short-term temporary, medium-term tasks and long-term stable preferences, and outputting a multi-level temporal preference classification result; reading the multi-level temporal preference classification result, analyzing the conditions and triggering factors of preference conversion, establishing conversion rules and boundary conditions, and outputting a temporal preference conversion rule set; Based on the multi-level temporal preference classification results and temporal preference conversion rule set, multi-level preference conflict coordination and decision-making are carried out to generate a unified preference decision result; Based on the unified preference decision results, the memory-enhanced preference model is updated and applied to generate an updated user preference model, which is then applied to the interactive interface and content generation to obtain personalized recommendation results. After generating the multi-level temporal preference classification results and the temporal preference conversion rule set, it also includes: Based on the standardized feature vector set, cross-modal preference consistency analysis is performed to generate a cross-modal preference consistency report; including: reading the standardized feature vector set, constructing a shared semantic space of text modality and image modality through contrastive learning methods, and outputting a cross-modal alignment mapping function; reading the standardized feature vector set and the context association graph, analyzing the correspondence between user behavior modality and semantic expression, establishing a mapping model from behavior sequence to semantic intention, and outputting a behavior-semantic mapping matrix; reading the standardized feature vector set of each modality separately, calculating the consistency degree of preference expression at different time points in the same modality, identifying preference conflicts and changes within the modality, and outputting the internal consistency score of each modality; reading the cross-modal alignment mapping function, the behavior-semantic mapping matrix and the standardized feature vector set of each modality, calculating the semantic consistency of preference expression between different modalities, identifying expression differences between modalities, and outputting an inter-modal consistency matrix; reading the internal consistency score and the inter-modal consistency matrix, accurately locating the conflict points of preference expression between different modalities, quantifying the degree of conflict and the scope of influence, and outputting a cross-modal preference consistency report, including the conflict intensity, location and possible explanations; Based on the cleaned multimodal interaction sequence, implicit preference extraction and verification are performed to generate an implicit preference mapping table containing semantic descriptions, trigger conditions, and confidence scores. Based on the results of multi-level temporal preference classification, cross-domain preference boundary identification and migration rule generation are performed to obtain a cross-domain preference migration rule set; including: reading the cleaned multimodal interaction sequence and multi-level temporal preference classification results, extracting domain features of different application scenarios, establishing domain feature vectors, and outputting application domain feature space; reading domain feature vectors in the application domain feature space, calculating the similarity matrix between different application domains, identifying the correlation strength between domains, and outputting the domain similarity matrix; reading the multi-level temporal preference classification results and the application domain feature space, analyzing the applicability and performance differences of specific preferences in different domains, establishing an association mapping between preferences and domain attributes, and outputting a preference-domain association table; reading the domain similarity matrix and the preference-domain association table, identifying the conditions for cross-domain migration or intra-domain isolation of preferences through a decision tree algorithm, and outputting a preference migration boundary definition; based on the preference migration boundary definition, generating specific cross-domain preference migration rules, and verifying the effectiveness through historical data simulation, adjusting the rule parameters, and outputting the final cross-domain preference migration rule set.
2. The method according to claim 1, characterized in that The temporal preference hierarchy identification and classification includes the steps of constructing an adaptive time decay model. The steps of constructing an adaptive time decay model include: Based on the cleaned multimodal interaction sequence, a preference time distribution map is constructed, the preference expression time features are extracted, and the preference time feature vector is generated; Perform cluster analysis on the preferred time feature vectors to obtain the preferred time behavior categories; calculating category-specific decay parameters for the preferred time behavior categories and generating a category decay parameter set; Constructing a mixed attenuation function based on the category attenuation parameter set and generating an attenuation function template; Based on the preferred time feature vector, preferred time behavior category and attenuation function template, personalized parameter learning and tuning are performed to generate personalized attenuation parameters; A dynamic adjustment mechanism is implemented based on personalized decay parameters to generate an adaptive time decay model.
3. The method according to claim 2, characterized in that The steps of extracting the time characteristics of preference expression and performing cluster analysis to obtain the preference time behavior categories include: Extract frequency periodicity features, continuous stability features, repetition interval features and recent activity features from the preference time distribution map, and combine them to form a preference time feature vector; Apply K-means++ clustering algorithm to the preference time feature vector to generate initial clustering results; Based on the initial clustering results, the optimal number of clusters was evaluated by the silhouette coefficient and Davies-Bouldin index to generate the preferred time behavior categories.
4. The method according to claim 2, characterized in that: The steps of constructing a mixed attenuation function and generating an attenuation function template include: The structure of the hybrid function is determined based on the category decay parameter set, including three modes: exponential decay, hyperbolic decay and power law decay; Set the weight coefficient and attenuation rate parameters to construct a parameterized mixed attenuation function: D(t, c, p) = w1(c)*exp(-α(p)*t)+w2(c)*(1 / (1+β(p)*t))+w3(c)*(t (-γ(p)) ), where t represents the time variable, c represents the preference category, p represents the personalization parameter, w1, w2, w3 are weight coefficients and satisfy w1+w2+w3=1, α, β, γ are decay rate parameters; Construct a decay function template based on a parameterized mixed decay function.
5. The method according to claim 2, characterized in that: The steps of performing personalized parameter learning and tuning and implementing a dynamic adjustment mechanism to generate an adaptive time decay model include: Based on the preference time feature vector, preference time behavior category and decay function template, a Bayesian optimization algorithm is used to learn personalized decay parameters on historical data and minimize the error between the predicted preference strength and the real user feedback; A responsive adjustment mechanism is built based on personalized decay parameters. The personalized decay parameters are continuously fine-tuned through online learning methods according to real-time user feedback data to form an adaptive time decay model with self-adjustment capabilities. The performance of the adaptive time decay model is verified on the reserved test set of historical data, and the prediction accuracy, recall rate and F1 score are calculated to generate the verification results. The parameters of the adaptive time decay model are finally calibrated according to the verification results to generate the optimized adaptive time decay model.
6. The method according to claim 1, characterized in that The cross-modal preference consistency analysis includes the steps of constructing a semantic-visual alignment space. The steps of constructing a semantic-visual alignment space include: Extract text semantic features and image semantic features from the standardized feature vector set; Based on the cleaned multimodal interaction sequences, temporally and semantically related text-image pairs are selected to construct a curated parallel corpus. A dual-tower neural network model is constructed using text semantic features, image semantic features, and a selected parallel corpus to learn shared latent spaces and generate a cross-modal alignment mapping network. Verify and fine-tune the semantic consistency of the cross-modal alignment mapping network to generate an improved cross-modal alignment mapping network; Based on the improved cross-modal alignment mapping network, a preference expression-specific adaptation layer is constructed to generate a preference-enhanced cross-modal mapping network, i.e., a semantic-visual alignment space. The preference-enhanced cross-modal mapping network is evaluated and optimized to generate a cross-modal alignment mapping function.
7. The method according to claim 6, characterized in that Shared latent space learning includes the steps of constructing a cross-modal contrastive loss function. The steps of constructing a cross-modal contrastive loss function include: Construct a global alignment loss component to ensure that the overall representation distributions of different modalities are similar; Construct a local semantic consistency loss component to ensure that the distance between semantically similar content in the cross-modal space is within a preset range; Constructing a structure-preserving loss component to ensure that the semantic structure within the modality is preserved after mapping; The global alignment loss component, the local semantic consistency loss component and the structure preservation loss component are combined to form a cross-modal contrast loss function; A dual-tower neural network model is trained using a cross-modal contrast loss function, text semantic features, image semantic features, and a selected parallel corpus to generate a cross-modal alignment mapping network.
8. The method according to claim 6, characterized in that The steps of constructing a preference expression-specific adaptation layer and generating a preference-enhanced cross-modal mapping network include: Based on the improved cross-modal alignment mapping network, a dedicated adaptation layer is constructed for the special properties of preference expression to form a preference-specific adaptation layer; Use preference-specific adaptation layers to strengthen preference-related semantic features and weaken irrelevant features to generate enhanced preference expression features; Based on the enhanced preference expression features, the preference expression data is extracted and the preference specific adaptation layer is fine-tuned to obtain the fine-tuning parameters; Based on the improved cross-modal alignment mapping network, preference-specific adaptation layer and fine-tuning parameters, the migration from the general semantic space to the preference-specific semantic space is achieved to generate a preference-enhanced cross-modal mapping network.
9. The method according to claim 1, characterized in that: Implicit preference extraction and verification includes the steps of constructing a low-intrusion verification strategy. The steps of constructing a low-intrusion verification strategy include: Based on the cleaned multimodal interaction sequence, generate implicit preference hypothesis set and conduct verification demand analysis and classification to generate classified preference hypothesis; Based on the classified preference hypothesis, natural dialogue fusion verification templates, parameter fine-tuning reaction observation schemes, and alternative scheme comparison strategies are constructed to generate dialogue fusion verification template libraries, parameter fine-tuning verification scheme sets, and alternative scheme comparison verification strategies. Based on the dialogue fusion verification template library, parameter fine-tuning verification solution set and alternative solution comparison verification strategy, a multi-round progressive verification plan is constructed to generate a timing verification plan; Construct an adaptive control mechanism for the verification process based on the timing verification plan and generate verification control parameters; Integrate the dialogue fusion verification template library, parameter fine-tuning verification solution set, alternative solution comparison verification strategy, timing verification plan and verification control parameters to form a low-interference verification strategy.
Citation Information
Patent Citations
Multi-modal personalized content generation method
CN118260483A
Memory retrieval method for enhancing multi-modal long-context dialogue ability of large language model
CN119293139A