Personalized regimen generation method, apparatus, device, and medium
Patent Information
- Application Number
- CN202511170116.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-08-20
AI Technical Summary
[0005]本发明的主要目的在于提供一种个性化方案生成方法、装置、设备及存储介质,旨在解决现有技术无法对多模态数据进行融合处理并生成包含需求、偏好和情绪状态的动态画像,从而限制了个性化方案生成的精准性与实时性的技术问题
[0020] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating personalized solutions, comprising: collecting multimodal data and performing preprocessing operations to generate a preprocessed data set and data relationship indices; in a fusion processing module, performing cross-modal interaction processing and attention interaction processing on the preprocessed data set based on the data relationship indices to generate a comprehensive feature vector; inputting the comprehensive feature vector into a recognition model for recognition operations to generate a dynamic profile of the object; generating a set of personalized optimization paths based on the dynamic profile of the object; extracting the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combining them with the set of personalized optimization paths to generate a personalized solution. This invention generates a dynamic profile of an object that simultaneously includes demand, preference, and emotional state by fusing and processing multimodal data and extracting deep features. By combining this profile with an optimization path generation mechanism, it can output accurate and differentiated personalized solutions for objects with different characteristics, achieving a comprehensive characterization of object needs and improving the real-time and targeted nature of solution generation.
Smart Images

Figure CN121032544B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for generating personalized solutions. Background Technology
[0002] In the fintech sector, existing marketing support systems still suffer from significant shortcomings in terms of intelligence and effectiveness. Most systems rely heavily on the analysis of structured data, such as customer account information and transaction records, exhibiting weak capabilities in processing unstructured data like social media posts, images, and voice recordings, and lacking the technological foundation for multimodal data fusion. This limitation prevents the systems from fully capturing customer behavioral characteristics and potential needs, resulting in often one-sided data analysis that fails to capture deeper information such as customer purchase motives and changing interests, thus impacting the reliability and effectiveness of targeted marketing decisions.
[0003] In the healthcare sector, intelligent support systems also have shortcomings in patient needs analysis and service customization. Existing systems largely focus on analyzing structured data such as medical records and examination reports, lacking comprehensive utilization of unstructured data such as doctors' consultation voice recordings, patients' facial expressions, and social interactions. This leads to insufficient identification of patients' emotions, health preferences, and potential health risks. This not only affects the personalized development of health management plans but also limits the ability to dynamically adjust communication strategies in doctor-patient interactions, easily resulting in inaccurate information transmission or untimely responses.
[0004] Furthermore, in cross-industry intelligent interaction applications, existing technologies generally rely on template-based logic for demand analysis and personalized solution generation, making it difficult to output customized results based on individual differences. For different customer or patient groups, the product recommendations and service strategies output are often highly similar, lacking specificity. Simultaneously, in the data processing chain, the coordination between various functional modules is low; data collection, feature analysis, and solution generation are often fragmented, resulting in delayed information transmission and an inability to form a real-time, dynamic optimization support system. These shortcomings directly affect the application effectiveness of interactive systems in precise and personalized service scenarios. Summary of the Invention
[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for generating personalized solutions, aiming to solve the technical problem that existing technologies cannot fuse multimodal data and generate dynamic profiles containing needs, preferences, and emotional states, thus limiting the accuracy and real-time performance of personalized solution generation.
[0006] To achieve the above objectives, the present invention provides a method for generating personalized schemes, comprising:
[0007] Collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes;
[0008] In the fusion processing module, based on the data relationship index, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector.
[0009] The comprehensive feature vector is input into the recognition model for recognition operation to generate a dynamic portrait of the object;
[0010] Based on the dynamic profile of the object, a set of personalized optimization paths is generated;
[0011] Extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the set of personalized optimization paths to generate a personalized solution.
[0012] Furthermore, to achieve the above objectives, the present invention provides a personalized solution generation apparatus, comprising:
[0013] The preprocessing module is used to collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes;
[0014] A cross-modal fusion processing module is used to perform cross-modal interaction processing and attention interaction processing on the preprocessed data set based on the data relationship index in the fusion processing module, and generate a comprehensive feature vector;
[0015] The multi-dimensional feature recognition module is used to input the comprehensive feature vector into the recognition model for recognition operations and generate a dynamic portrait of the object.
[0016] The optimized path generation module is used to generate a set of personalized optimized paths based on the dynamic profile of the object.
[0017] The personalized solution generation module is used to extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the personalized optimization path set to generate a personalized solution.
[0018] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a personalization scheme generation program stored in the memory and executable on the processor, wherein when the personalization scheme generation program is executed by the processor, it implements the steps of the personalization scheme generation method as described above.
[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a personalization scheme generation program, which, when executed by a processor, implements the steps of the personalization scheme generation method as described above.
[0020] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating personalized solutions, comprising: collecting multimodal data and performing preprocessing operations to generate a preprocessed data set and data relationship indices; in a fusion processing module, performing cross-modal interaction processing and attention interaction processing on the preprocessed data set based on the data relationship indices to generate a comprehensive feature vector; inputting the comprehensive feature vector into a recognition model for recognition operations to generate a dynamic profile of the object; generating a set of personalized optimization paths based on the dynamic profile of the object; extracting the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combining them with the set of personalized optimization paths to generate a personalized solution. This invention generates a dynamic profile of an object that simultaneously includes demand, preference, and emotional state by fusing and processing multimodal data and extracting deep features. By combining this profile with an optimization path generation mechanism, it can output accurate and differentiated personalized solutions for objects with different characteristics, achieving a comprehensive characterization of object needs and improving the real-time and targeted nature of solution generation. Attached Figure Description
[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0022] Figure 1 This is a schematic diagram of an application environment for a personalized scheme generation method according to an embodiment of the present invention;
[0023] Figure 2 This is a flowchart illustrating an embodiment of the personalized solution generation method of the present invention;
[0024] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the personalized solution generation device of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0026] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0027] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0028] The personalized scheme generation method provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can collect multimodal data from the user terminal and perform preprocessing operations to generate a preprocessed data set and data relationship indexes. In the fusion processing module, based on the data relationship indexes, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector. The comprehensive feature vector is input into the recognition model for recognition operations to generate a dynamic profile of the object. A personalized optimization path set is generated based on the dynamic profile of the object. The demand intensity item, preference weight item, and emotional state item are extracted from the dynamic profile of the object, and combined with the personalized optimization path set to generate a personalized solution. This invention generates a dynamic profile of the object that simultaneously contains demand, preference, and emotional state by fusing multimodal data and extracting deep features, and combines it with an optimization path generation mechanism. This enables the output of accurate and differentiated personalized solutions for objects with different characteristics, achieving a comprehensive characterization of object needs and improving the real-time and targeted nature of solution generation. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0029] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the personalized solution generation method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0030] like Figure 2 As shown, the personalized solution generation method proposed in this invention includes the following steps:
[0031] S10: Collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes;
[0032] In this embodiment, the process of collecting multimodal data and performing preprocessing operations first involves acquiring multi-source data. This process is accomplished by constructing a data acquisition interface that can interface with different business systems, devices, and third-party channels. The interface supports the parallel access of structured and unstructured data. Structured data may include transaction logs, account information, health check values, etc., output from business databases; unstructured data may include voice recordings from customer service or doctors, text content from online communications, images obtained by shooting or scanning, etc. The acquisition interface can be implemented using a combination of pull and push methods. In scenarios with high real-time requirements, a subscription mechanism is used to ensure that new data is processed as soon as it arrives. To ensure the traceability of data in subsequent processing, each piece of data is appended with a source identifier and a collection time stamp during acquisition and is stored sequentially in the access queue to ensure that data of the same object is arranged in chronological order.
[0033] The first step in preprocessing is data cleaning, which aims to remove records with severe missing values or obvious outliers, correct correctable errors, and standardize data formats. For example, units in numerical data need to be converted to consistent measurement standards, text data needs to be deduplicated and encoded uniformly, image data can undergo size and resolution alignment, and audio data can undergo noise suppression and silent segment removal. The cleaning process can also remove or replace sensitive information in the desensitization module to ensure compliant data use in subsequent stages.
[0034] After cleaning, data from different modalities are converted into formats that are easy to compute and store. Image data is matrix-processed and stored in a uniform height, width, and channel order, with color space information added for subsequent restoration or alignment. Speech data undergoes feature extraction to form a spectrogram representation, which can be a linear spectrum or a Mel spectrum, and includes time step and sampling rate information to synchronize with the temporal information of other modalities. Numerical data is normalized to map different dimensions to the same numerical range or distribution, reducing convergence difficulties caused by differences in numerical ranges in subsequent algorithms. Text data is vectorized into dense or sparse vector representations. Vector generation methods can be based on statistical models after word segmentation or semantic vectors generated by deep pre-trained models, while recording the vocabulary or model version used.
[0035] Object unique identifiers are crucial in this step, used to associate multimodal data from different data sources with the same entity. The source of a unique identifier can be a single account number, patient number, or multiple identifiers from different sources can be mapped to a unified identifier through a multi-identifier matching strategy. During the identifier establishment process, duplicate detection and conflict resolution are required to ensure the uniqueness and stability of each identifier.
[0036] The data relationship index uses the object's unique identifier as the index key, recording the storage location, time range, source information, and version information of matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data. This allows subsequent processing modules to directly locate the corresponding data slice when they need to access multimodal data of the same object.
[0037] The preprocessed dataset is generated by merging modal data that has been processed in a unified format according to object unique identifiers and time windows. The merging process maintains the time alignment and modal completeness of the data. The preprocessed dataset provides a highly consistent and low-latency input foundation for subsequent fusion computation, avoiding the need to repeatedly perform data preparation operations in downstream modules.
[0038] In scenarios with large volumes of real-time data, distributed stream processing systems can be used for data acquisition and preprocessing. Tasks such as speech feature extraction, image matrixing, and text vectorization are deployed as independent processing nodes. Data segments are passed through message queues, and partitioned routing is performed using unique object identifiers. This ensures that data streams of the same object are sent to the same processing node, guaranteeing time sequence consistency. For offline scenarios, cleaning and format conversion jobs can be run in batches in data lakes or columnar storage systems, which is suitable for processing large amounts of historical data.
[0039] The specific strategies for data cleaning can be adjusted according to the characteristics of the industry. For example, in medical data, the unit conversion of numerical fields needs to be adapted to medical standards, and noise suppression of voice data needs to take into account the background noise characteristics of wards and clinics. In financial data, outlier detection of transaction records needs to be combined with transaction type and customer historical behavior model, and the processing of image data may involve resolution unification and watermarking of contract images or identity authentication images.
[0040] In multi-source merging scenarios, unique object identifiers can be generated by establishing a primary key mapping table and introducing a priority strategy to select retained identifiers, while recording the mapping relationship for subsequent tracking. When there are conflicts in the identifier fields of the source data, a matching algorithm can be used to calculate similarity, and high-risk matches can be manually reviewed to ensure accurate identification.
[0041] Data relationships can be stored in a key-value database for fast querying, or a graph database can be used to represent many-to-many relationships between different modal data fragments and objects. This approach is more suitable for applications that require complex cross-modal relationship tracking. When generating the preprocessed dataset, a strict merging pattern that is fully modalized or a lenient merging pattern that allows for partial omissions can be selected based on business needs to balance data integrity and availability.
[0042] Example Description: In the healthcare business field, patient consultation voice, imaging examination results, laboratory values, and electronic medical record text can be accessed through multiple source interfaces, and then subjected to spectrogram generation, matrix processing, normalization processing, and vectorization processing, respectively. The consultation number and medical record number are mapped as unique identifiers, and all multimodal examination results of the patient are indexed and recorded. The preprocessed data set can be directly used by the disease identification and treatment plan recommendation module.
[0043] In the fintech business, customer transaction records, account information, customer service calls, online consultation texts, and identity verification images can be uniformly collected, cleaned, and transformed to generate matrix images, spectrogram voice data, normalized transaction features, and vectorized text features with consistent structure. Account numbers and device fingerprints are merged into a unique identifier, and the index records cross-channel interaction data of the same customer. The pre-processed data set can be directly called by risk assessment and personalized product recommendation modules.
[0044] This embodiment unifies the collection, cleaning, transformation, normalization, and vectorization of multimodal data, and establishes data associations with unique object identifiers as the core. Data from different sources and modalities can achieve accurate matching and time alignment, reducing data loss and mismatch during downstream processing, and improving the stability and accuracy of fusion processing and recognition calculation.
[0045] S20, In the fusion processing module, based on the data relationship index, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector;
[0046] In this embodiment, the fusion processing module utilizes data association indices to establish access paths between different modalities. This index quickly locates text, image, speech, and numerical data belonging to the same object, ensuring data consistency in time series and source identification. Cross-modal interaction processing refers to feature mapping and information exchange between two or more different modalities to extract joint features that cannot be expressed by a single modality. In practical implementation, text and images can be associated through cross-modal feature alignment mechanisms. For example, using a shared semantic space model, the vector representation of text and the visual features of images are projected into a unified embedding space, where similar semantic content has a high similarity score. Speech and numerical data can be processed through a feature interaction network, fusing the temporal energy distribution features in the spectrogram with the temporal trend features in the numerical sequence to form an interactive feature representation with both temporal correlation and quantification information.
[0047] Attention-based interactive processing introduces attention weighting mechanisms into acquired intra-modal or cross-modal features. By calculating the relevance scores of each feature under a specific task, the contribution weights of the features are dynamically adjusted. For sentiment features, attention mechanisms can be used to highlight semantic segments or visual regions related to emotion judgment, such as emphasizing areas containing facial expressions in images and strengthening the vector weights of emotional words in text. For decision-style features, higher attention weights can be applied to segments of speech with intonation changes and key fluctuation points in numerical data, thereby improving the accuracy and stability of subsequent analysis.
[0048] The comprehensive feature vector is generated by fusing important features from multiple modalities according to a unified vector structure after cross-modal interaction and attention weighting. This fusion can be achieved through weighted concatenation, element-wise weighted summation, tensor fusion, etc., and the source identifiers and timestamps of each modal feature are retained during generation so that the source of specific features can be traced back in subsequent use.
[0049] When processing large-scale business data, the fusion processing module can be deployed as a parallel multi-channel structure, with different channels specializing in handling specific modal combinations. For example, one channel might handle text-image alignment, while another handles voice-numerical interaction. The outputs from each channel are then integrated through a fusion layer. In low-latency scenarios, an online feature fusion approach can be used. After the data is linked and located, a lightweight cross-modal network and attention module are used to directly generate a comprehensive feature vector, reducing intermediate storage and loading overhead.
[0050] Cross-modal feature alignment can utilize pre-trained multimodal coding models, selecting training corpora based on industry characteristics. For example, in the healthcare field, paired data containing medical images and diagnostic descriptions can be used for pre-training, while in the fintech field, transaction document images and transaction description texts can be used. Voice and numerical feature interaction can be achieved by extracting time-frequency patterns from speech using multi-layer convolutional networks, aligning them with numerical time-series patterns extracted using long short-term memory networks, and then fusing them along the time dimension.
[0051] The implementation of attention interaction processing can be diversified. It can employ a self-attention mechanism to allocate weights within the same modality, or a cross-attention mechanism to calculate a correlation matrix across different modalities and then weight the fused features based on the matrix distribution. To adapt to different scenarios, the attention mechanism can be parameterized, allowing it to automatically optimize the weight allocation strategy during the training phase based on specific task objectives.
[0052] Example: In the healthcare business, patients' imaging examination results and doctors' diagnostic descriptions are fused through a text-image alignment mechanism. Textual information in medical records is matched with lesion areas in images in semantic space. Voice consultation records and time series of vital signs are associated through a feature interaction network to generate a comprehensive feature vector containing disease tendencies and risk levels, providing input for disease prediction and personalized treatment suggestions.
[0053] In the fintech business, customer identity verification images are aligned with the text of account information entered, and voice features from customer service calls are integrated with customer transaction behavior data to obtain a comprehensive feature vector that includes risk preferences, trading habits, and emotional state, providing basic data support for personalized product recommendations and risk control models.
[0054] This embodiment utilizes data associations in the fusion processing module to perform cross-modal and attention interactions. Features from different modalities can be fused under a unified spatiotemporal coordinate system, improving the semantic integrity and task relevance of feature representations. Cross-modal interaction expands the information dimension of a single modality, while the attention mechanism ensures the weight enhancement of key features. The combined feature vector generated by these two mechanisms provides a richer and more accurate input foundation for subsequent recognition.
[0055] S30, input the comprehensive feature vector into the recognition model for recognition operation to generate a dynamic portrait of the object;
[0056] In this embodiment, the comprehensive feature vector undergoes morphological calibration and version marking before entering the recognition model to ensure consistency in input dimensions, time windows, and embedding space. Morphological calibration includes vector fragment alignment, missing segment imputation, and scale normalization. Version marking records the encoder and timestamp used when generating the vector, facilitating subsequent backtracking and comparison. The recognition model is divided into a demand recognition model, a preference recognition model, and an emotion recognition model, all sharing the same comprehensive feature vector. Each model completes vector channel selection and noise reduction through its respective front-end adaptation layer. The demand recognition model receives the joint representation of fused text, image, speech, and numerical signals, outputs a multi-dimensional score oriented towards consumption intention and service demands, and then generates a demand intensity item by the mapping unit. The preference recognition model focuses on long-term stable tendencies and interaction preferences, using cross-modal attention gating to filter sub-vectors related to communication channels, presentation styles, and delivery forms, outputting a preference distribution, and then generating preference weight items by the weight shaping unit. The emotion recognition model emphasizes short-term fluctuations and emotional cues. It synchronously aligns the speech spectrum and text emotion word embeddings along the timeline, and combines them with facial region embedding from images to output emotion classification and confidence scores. Then, the state induction unit generates emotion state items. The three sets of results enter the profile construction unit, which aggregates them according to the object's unique identifier, time window, and scene label to form a dynamic profile of the object. The profile structure includes a demand intensity item, a preference weight item, an emotion state item, and a source index and time stamp, ensuring that subsequent calls can accurately retrieve information by object, time period, and scene.
[0057] To reduce the impact of noise on judgment, the recognition process introduces consistency checks and cross constraints. Consistency checks compare the stability of the demand intensity term and preference weight term across multiple observations, eliminating extreme outliers. Cross constraints use the emotion state term to adjust for instantaneous fluctuations in the demand intensity term, preventing short-term emotions from causing abnormally high or low values. To balance real-time performance and robustness, the profile building unit provides two paths: sliding window aggregation and incremental updates. The sliding window generates a snapshot of the profile for the current time period, while incremental updates write new recognition results to the profile storage and record the version, allowing rollback to any historical version if necessary.
[0058] In implementation, the input manager receives the comprehensive feature vector and distributes it to the three types of models; the model execution engine runs in parallel, supporting both batch processing and streaming scheduling; the result shaper converts the model's native output into a unified profile field and adds a source index; the profile storage uses the object's unique identifier as the key to store structured profiles, supporting time-range queries and scene filtering. The entire chain uses data relationships to locate the object's data source, comprehensive feature vectors to carry cross-modal information, and the dedicated outputs of the three types of models to constitute the three core elements of the profile, ensuring a closed and traceable mapping from input to profile.
[0059] In low-latency interaction scenarios, a streaming execution chain can be used. Upon receiving the comprehensive feature vector, the input manager immediately triggers parallel inference for the three types of models. The model execution engine enables resident memory and graph optimization for the lightweight subnet. The result shaper generates demand intensity, preference weight, and emotion state items in real time. The profile building unit aggregates using a short-cycle sliding window to ensure that profile updates closely follow the interaction process. In high-throughput offline evaluation scenarios, a batch processing chain can be used. The comprehensive feature vector is divided into batches by object and time. The model execution engine runs in a distributed manner. The result shaper performs denoising and consistency checks on the batch output. The profile building unit generates steady-state profiles using a longer time window for policy library training and evaluation.
[0060] In healthcare scenarios, the input adaptation layer of the demand recognition model can increase the weight of vital signs and image region embedding, the preference recognition model emphasizes remote follow-up preferences and information reception methods, and the emotion recognition model focuses on increasing the weight of hesitant speech segments and worried text expressions. In fintech scenarios, the demand recognition model strengthens clues about transaction motives and fund usage, the preference recognition model highlights channel selection and risk tolerance, and the emotion recognition model focuses on changes in speaking speed and areas of facial tension.
[0061] To adapt to different production environments, the profile building unit provides a configurable set of fields. For finer-grained control, fields can be expanded without changing the three item names; for example, adding source contribution ratio to the demand intensity item, adding a scenario dimension to the preference weight item, and adding duration to the emotional state item. For resource-constrained devices, the three models can be compressed into a shared backbone and three lightweight headers, maintaining consistency in the three outputs. For highly reliable clusters, model integration and voting can be enabled, with the result shaper selecting outputs to write to the profile based on a consistency threshold.
[0062] To optimize different parameters, the input manager can adjust the vector channel selection strategy and increase the channel weights related to the application scenario; the model execution engine can adjust the batch size and parallelism to balance throughput and latency; the portrait building unit can adjust the sliding window length and incremental write frequency to control the ratio of new and old information; and the consistency check can adjust the outlier removal intensity to adapt to the noise level.
[0063] Example Explanation: In the healthcare business field, the comprehensive feature vector generated during a patient's visit enters the recognition process. The demand recognition model outputs the strength of the intention to return for a follow-up visit, forming a demand intensity item. The preference recognition model outputs the communication style and follow-up frequency tendency, forming a preference weight item. The emotion recognition model outputs states such as tension or relaxation based on the consultation voice and facial expressions, forming an emotion state item. The profile building unit aggregates the visit time window into a dynamic profile of the object, which is used for subsequent follow-up paths and customized educational materials.
[0064] In the fintech business, customer inquiries and historical behaviors are integrated into a comprehensive feature vector that enters the identification process. The demand identification model outputs financial planning and product interest to form demand intensity items. The preference identification model outputs channel and service method preferences to form preference weight items. The emotion identification model outputs positive or hesitant states based on call and video footage to form emotion state items. The profile building unit aggregates the data into dynamic profiles of objects according to the conversation time window, which are then used for subsequent resource matching and interaction strategy generation.
[0065] This embodiment inputs comprehensive feature vectors in parallel into the demand recognition model, preference recognition model, and emotion recognition model, and constructs a dynamic profile of the object after shaping the results. By jointly utilizing cross-modal and temporal information, it outputs a profile that is structurally stable, traceable, and versionable. Based on consistency verification and cross-constraints, the impact of short-term noise on the profile is reduced. The profile update can respond to real-time changes while maintaining a stable tendency, providing reliable input for subsequent path matching and personalized generation.
[0066] S40, Based on the dynamic profile of the object, generate a set of personalized optimization paths;
[0067] In this embodiment, the dynamic profile of the object serves as input, including demand intensity, preference weight, emotional state, time stamp, source index, and other auxiliary information. The goal of generating a personalized optimization path set is to map the three quantitative results in the profile into executable and traceable path entries, organized in a set structure for direct use in subsequent decision-making processes. First, the profile parsing unit reads the demand intensity, preference weight, and emotional state, and verifies whether the time stamp is within a valid window; if multiple versions of the profile exist, the latest stable version or the version matching the current scenario is selected according to the version strategy. Subsequently, the resource mapping unit loads the index view of the resource entry library, which is derived from a structured list of products and services, including class tags, adaptation conditions, dependency constraints, and availability markers; simultaneously, the strategy mapping unit loads the index view of the scenario strategy entry set, which records execution entries such as communication order, presentation format, follow-up method, and threshold control. The resource mapping unit establishes a resource candidate set based on the demand intensity item. It employs a combined approach of conditional filtering and similarity retrieval, first eliminating entries incompatible with the profile based on fit criteria, then sorting the remaining entries by their matching score with the demand intensity item. The matching score is calculated by weighted similarity between the demand intensity item and the entry's fit vector, and the source index is retained for traceability. The strategy mapping unit establishes a strategy candidate set based on the emotion state item and the preference weight item. It controls the activation order of different sub-rules within the strategy entries through gating weights. For example, when the preference weight item displays offline communication preferences, strategy entries containing face-to-face communication sub-rules are prioritized; when the emotion state item displays hesitation, the weight of explanation and confirmation rules is increased.
[0068] The path construction unit converts the resource candidate set and strategy candidate set into product matching optimization paths and interaction strategy optimization paths, respectively. The path structure uses a phased representation, with each phase including target items, execution actions, preconditions, successor pointers, evaluation points, and rollback conditions. The product matching optimization path uses demand intensity as the main control variable, defining the presentation order and switching conditions in stages; the interaction strategy optimization path uses emotional state as the main control variable, defining discourse organization, media switching, and rhythm control in stages. To avoid conflicts between the two paths, the path coordination unit performs constraint consistency checks: when the product matching optimization path requires rapid progress while the interaction strategy optimization path requires delayed explanation, priority is determined based on profile time tags and preference weights, and intermediate actions of waiting or supplementary explanation are injected at path nodes. The path scoring unit calculates a comprehensive score for the two paths, composed of matching degree, usability, and historical effectiveness; historical effectiveness is obtained by referencing data relationships to trace the execution records of similar objects in similar scenarios. Paths whose comprehensive scores reach the threshold are encapsulated as set entries and written into the personalized optimization path set. Each entry carries a unique object identifier, version number, time window, source index, and conflict resolution summary to ensure subsequent use and traceability. The final output is a personalized optimization path set, which contains at least one product matching optimization path and one interaction strategy optimization path, and the two paths have been verified for consistency and executability.
[0069] In a low-latency interactive environment, an online generation process can be adopted. The portrait parsing unit resides in memory and immediately triggers the resource mapping unit and strategy mapping unit upon receiving the latest dynamic portrait of the object. The two units search their respective candidate sets in parallel, and the path construction unit incrementally generates path nodes in a streaming manner, outputting set entries first when the minimum executable segment is met, with the remaining nodes completed in the background. To reduce computational overhead, the resource entry library and the scene strategy entry set construct a dual-channel retrieval structure of inverted index and vector index. Conditional filtering uses the inverted index, while matching and sorting use the vector index. Thresholds, weights, and window lengths can be distributed through the configuration center to avoid frequent redeployment.
[0070] In a high-throughput batch generation environment, object dynamic profiles can be aggregated into batches by time slices. The resource mapping unit and the strategy mapping unit generate candidate sets in a distributed manner. The path construction unit compresses staged nodes into the shortest feasible sequence through graph optimization and eliminates redundant jumps and invalid loops using dynamic programming. The path coordination unit abstracts the conflict patterns shared across objects into constraint templates and applies them to the entire batch of paths, significantly reducing the per-object solution cost.
[0071] In the healthcare business scenario, the resource item library can be expanded into two sub-libraries: health service items and educational material items. The demand intensity item drives the selection of service items, the preference weight item drives the selection of educational material carriers, and the emotional state item controls the communication pace. In the fintech business scenario, the resource item library is mainly composed of product categories and ancillary services. The demand intensity item controls the display depth and trial calculation nodes, the preference weight item controls the channels and presentation style, and the emotional state item controls the explanation and confirmation ratio.
[0072] To adapt to different production environments, the personalized optimized path set supports multiple versions coexisting and canary releases. In the online environment, a snapshot version can be enabled, retaining only the limited number of items with the highest scores in the current window; in offline evaluation, a full version can be enabled, retaining all executable paths and scoring details for training and playback. The three weights of the path scoring unit can be adjusted according to business objectives: increase the weight of historical performance when emphasizing sales, and increase the weight of matching degree and usefulness when emphasizing user experience.
[0073] To improve stability and interpretability, collection entries undergo two types of checks before being written. Consistency checks ensure that the preconditions for path nodes can be met under the current profile and available resources; readability checks map node actions to human-readable descriptions, including source indexes and threshold information, making it easier for business personnel to understand and adjust. If necessary, graph optimization and consistency checks can be moved to the database side, using stored procedures or computing engines to perform them closer to the data, reducing network round trips.
[0074] Example Description: In the healthcare business domain, a patient's dynamic profile shows high demand intensity, online reception preference, and mild anxiety. Resource mapping generates a service candidate set including remote follow-up consultations and medication management, while strategy mapping generates a communication candidate set primarily focused on reassurance and confirmation. Path construction forms a product matching optimization path: Phase 1 pushes the remote follow-up consultation arrangement, and Phase 2 guides the medication management plan; the interaction strategy optimization path inserts explanation and confirmation nodes in Phase 1 and reiteration and reminder nodes in Phase 2. After coordination, the two paths are written into a personalized optimization path set for subsequent execution to call in sequence.
[0075] In the fintech business, a customer's dynamic profile shows moderate demand intensity, mobile preference, and hesitation. Resource mapping generates a candidate set of installment payment and liquidity management products, while strategy mapping generates a candidate set of communication products characterized by mobile display and delayed explanation. Path construction yields a product matching optimization path, first displaying liquidity management, then switching to installment payment after preconditions are met; the interaction strategy optimization path inserts objection clarification and trial calculation demonstrations at key nodes. After scoring and consistency verification, the two paths form a personalized optimized path set, which is distributed segment by segment throughout the conversation and versions are retained for backtracking.
[0076] This embodiment generates a personalized set of optimized paths based on dynamic object profiling, transforming three quantitative pieces of information into executable, traceable, and evaluable path entries. This approach preserves the stability of the profiling while reflecting the immediate constraints of the current scenario. Resource mapping and strategy mapping are executed in parallel, path construction and coordination ensure that product presentation and interaction rhythms do not conflict, and path scoring and version management make the output comparable and iterable, providing a clearly structured input for subsequent execution and feedback loops.
[0077] S50, extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the set of personalized optimization paths to generate a personalized solution.
[0078] In this embodiment, the dynamic profile of the object serves as input, including metadata such as demand intensity, preference weight, emotional state, time stamp, version number, and source index. The demand intensity, derived from the quantitative output of the demand identification model, expresses the object's immediate desire and urgency for a certain type of resource or service, and can be mapped to display depth, pace of progress, and confirmation frequency. The preference weight, derived from the quantitative output of the preference identification model, expresses long-term tendencies such as channel preference, presentation style preference, and interaction media preference, and can be mapped to content organization, media selection, and presentation order. The emotional state, derived from the quantitative output of the emotion identification model, expresses states such as stability, hesitation, and resistance, and can be mapped to explanation ratio, clarification intensity, and pace control. The personalized optimization path set serves as another input, consisting of product matching optimization paths and interaction strategy optimization paths. The former provides the phased presentation and switching conditions for resource items, while the latter provides the phased organization and control signals for communication actions.
[0079] The process of generating a personalized solution begins with self-portrait analysis. The analysis unit verifies whether the time tag falls within the valid window; if multiple versions of the profile exist, the latest version that matches the current scenario is selected based on the version number and source index. Then, three extractions are performed to obtain the current moment's demand intensity, preference weight, and emotional state. These three items are then normalized and their confidence levels checked to form a ternary vector for decision-making. The mapping unit uses this ternary vector as control variables to read the personalized optimization path set. In the product matching optimization path, it locates the stage interval corresponding to the demand intensity threshold and extracts the resource items, preconditions, and successor pointers within that interval. In the interaction strategy optimization path, it locates the stage interval corresponding to the emotional state threshold and extracts the communication actions, rhythm parameters, and explanation ratios within that interval. In the common interval of the two paths, the preference weight is used to apply biases to the display media, layout, and sorting factors.
[0080] Conflicts and couplings require explicit handling. The coordination unit establishes a constraint matrix, with rows representing the progress constraints of resource items and columns representing the rhythm constraints of communication actions. Matrix elements record control signals such as allow, delay, or replacement. When a resource item requires rapid progress but the emotional state item suggests delayed explanation, the control signal is set to delay and a clarification action is inserted; when the preference weight item favors offline communication but the interaction strategy optimization path provides online actions, the control signal is set to replacement and a media switch is specified. After consistency verification, the solution construction unit generates four types of outputs into the same solution body. The product portfolio solution comes from the phased items of the product matching optimization path and the progress threshold of the demand intensity item; the key content of the product introduction comes from the weighted projection of the preference weight item onto the item attributes and the explanation ratio of the interaction strategy optimization path; the communication skills come from the action sequence of the interaction strategy optimization path and the gating weight of the emotional state item; the follow-up time node solution comes from the scaling of the demand intensity item on the cycle length and the time window of the path successor pointer.
[0081] The solution body adopts a structured representation, including a solution identifier, a unique object identifier, a version number, a time window, and four types of sub-solutions. Each sub-solution contains a source index and a traceable summary, recording the correspondence between three quantified values and path nodes, facilitating subsequent execution and replay. To ensure stability, two types of checks are performed before output. The consistency check verifies that the preconditions can be met under the current resource availability and verifies that there are no mutually exclusive actions within the same time window. The readability check converts key nodes into human- and machine-readable expressions, retaining thresholds, weights, and replacement reasons, facilitating understanding and review by external systems. If the check fails, it reverts to the previous stable interval, lowers the advancement threshold or adjusts the interpretation ratio, regenerates the solution body, and repeats the check until a version that meets the constraints is obtained.
[0082] In an online real-time environment, streaming generation can be used. After receiving the dynamic profile of the object, the parsing unit immediately outputs a ternary vector; the mapping unit searches the current interval of two optimization paths in parallel; the solution construction unit generates the first segment of the solution body at the smallest executable segment level and sends it out, while the remaining segments are completed in the background. Threshold and ratio parameters are stored in the configuration center, supporting the issuance of differentiated parameter sets according to business scenarios. When the emotional state fluctuates during the conversation, the interval switching of the interaction strategy optimization path is triggered instantly by a gating signal, and the corresponding paragraph of the solution body is replaced with a delayed explanation or reassurance and clarification action sequence.
[0083] Batch offline environments can be generated using batch processing. Object dynamic profiles are aggregated into batches by time slices. The mapping unit constructs candidate intervals using both inverted and vector indices. The solution construction unit transforms staged nodes into a directed acyclic graph, using dynamic programming to obtain the shortest feasible sequence that satisfies the constraints, reducing redundant jumps. Conflict templates and replacement templates are shared at the batch level, reducing per-object solution costs.
[0084] Compact representations can be used in resource-constrained environments. Only key nodes and jump conditions of each path are retained, and intermediate actions are inferred through interpolation strategies; the sentiment state item retains only three thresholds, and the preference weight item only affects the ranking without changing the medium, thereby reducing computational and bandwidth overhead.
[0085] In a high-consistency environment, a strong consistency strategy can be enabled. All range switches must pass consistency and readability checks, and any replacement action must be accompanied by the source index and the reason for the replacement to ensure transparency in replay and auditing.
[0086] Parameter adaptation can be achieved through feedback. Feedback data from the previous cycle's execution process is used to estimate the effective range of each threshold. The threshold for advancing the demand intensity item is adjusted upwards or downwards, and the threshold for the explanatory proportion is fine-tuned according to the historical distribution of the emotional state item. When a certain type of preference weight item has been dominant for a long time, the ranking factor is adjusted accordingly to better reflect actual acceptance habits.
[0087] Example Explanation: In the healthcare business sector, the profile shows high demand intensity, a preference for remote communication, and mild anxiety. The personalized optimization path set provides product matching optimization paths for remote follow-up consultations and medication management, as well as an interaction strategy optimization path focused on reassurance and confirmation. When generating the solution, the product portfolio solution arranges remote follow-up consultations in the first paragraph, followed by medication management; the product introduction highlights the follow-up consultation process and remote tool guidance; communication skills include reassurance and confirmation phrases in the first paragraph; and the follow-up timeline solution schedules follow-up consultations within a short timeframe and includes confirmation callbacks.
[0088] In the fintech business sector, the profile reveals moderate demand intensity, mobile preference, and hesitation. The personalized optimization path set provides product matching optimization paths for liquidity management and installment products, as well as interactive strategy optimization paths for objection clarification and trial calculation demonstrations. When generating the solution, the product portfolio first displays liquidity management, then switches to installment products after the preconditions are met; key product introductions on mobile devices emphasize repayment flexibility and fee structure; communication techniques insert objection clarification and trial calculation demonstration phrases at crucial points; and the follow-up timeline uses a longer observation window and sets up secondary reminders.
[0089] This embodiment uses a dynamic object profile and a personalized optimization path set as dual inputs, coupling three quantitative results with two phased paths within the same scheme to achieve the joint generation of resource presentation, content organization, communication actions, and time scheduling. Through threshold, gating, and constraint matrix mechanisms, conflicts between the pace of advancement and explanation are avoided; through source indexing and readability summaries, the output is ensured to be traceable, reviewable, and replayable; and through both online streaming and offline batch processing, real-time interaction is satisfied while balancing the stability and cost of batch generation.
[0090] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating personalized solutions, comprising: collecting multimodal data and performing preprocessing operations to generate a preprocessed data set and data relationship indices; in a fusion processing module, performing cross-modal interaction processing and attention interaction processing on the preprocessed data set based on the data relationship indices to generate a comprehensive feature vector; inputting the comprehensive feature vector into a recognition model for recognition operations to generate a dynamic profile of the object; generating a set of personalized optimization paths based on the dynamic profile of the object; extracting the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combining them with the set of personalized optimization paths to generate a personalized solution. This invention generates a dynamic profile of an object that simultaneously includes demand, preference, and emotional state by fusing and processing multimodal data and extracting deep features, and combines this with an optimization path generation mechanism. This enables the output of accurate and differentiated personalized solutions for objects with different characteristics, achieving a comprehensive characterization of object needs and improving the real-time and targeted nature of solution generation.
[0091] In one embodiment, step S10 includes:
[0092] S101, Construct a multi-source data acquisition interface to collect multimodal data;
[0093] S102, Perform a cleaning operation on the multimodal data to generate a cleaned data set;
[0094] S103, Perform a format conversion operation on the image data in the cleaned dataset to generate matrix format image data;
[0095] S104, Perform a format conversion operation on the speech data in the cleaned dataset to generate spectrogram speech data;
[0096] S105, normalization processing is performed on the numerical data in the cleaned data set to generate normalized numerical data;
[0097] S106, Perform vectorization processing on the text data in the cleaned data set to generate vectorized text data;
[0098] S107, Extract the unique identifier of the object from the cleaned data set;
[0099] S108, bind the matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data with the object's unique identifier to generate a data relationship index;
[0100] S109, merge the matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data to generate a preprocessed data set.
[0101] In this embodiment, the multi-source data acquisition interface establishes a unified access channel for both structured and unstructured inputs, covering sources such as transaction records, behavior logs, registration information, social text, images, and voice. The interface uniformly organizes access frequency, data format, time series labels, and source markers to form a traceable acquisition pipeline. During the acquisition phase, source markers and time tags serve as anchors for subsequent processing; data that fails source verification or has missing time signatures does not enter the downstream process. After multimodal data enters the cleaning phase, missing data detection, anomaly detection, and consistency verification are performed on structured fields; noise reduction and segmentation are performed on text; and integrity verification and decoding validity verification are performed on images and voice. The cleaning operation generates a cleaned data set, which retains source markers, time tags, object-related clue fields, and processing records for easy traceability and review.
[0102] After cleaning, the image data in the dataset undergoes a format conversion process. This process unifies the color space and scale, constrains the aspect ratio, fills in boundary areas, and converts the data into matrix-formatted image data. This conversion ensures that each image frame is mapped to a fixed-size numerical grid, maintaining a one-to-one correspondence between grid cells and pixel positions. The original resolution and cropping position are recorded for playback when needed. The speech data enters a time-frequency transformation process, performing endpoint segmentation and energy gating to convert the time series into a two-dimensional time-frequency representation, outputting a spectrogram of the speech data. The discretization method of the time-frequency axis and the window overlap strategy are defined by configuration options to adapt to different sampling conditions. The conversion result includes time alignment information to ensure a consistent reference for subsequent alignment with images or text.
[0103] Structured numerical data enters the scaling transformation process, where unit alignment, range compression, and distribution correction are performed based on field type, outputting normalized numerical data. This process maintains the relative relationships between fields and generates a transformation summary to ensure the feasibility of the inverse transformation. Text data enters the word segmentation and semantic mapping process, handling expressions, spoken language, and multilingual content, producing fixed-length representations, and outputting vectorized text data. The mapping process considers both lexical sequences and contextual semantics, recording vocabulary snapshots and mapping configurations to ensure consistency and verifiability across batches.
[0104] The object unique identifier is extracted from the cleaned dataset, prioritizing registration fields and stable identifiers. When missing, candidate identifiers are aggregated based on source markers and timestamps, and a single identifier is determined through consistency verification. The object unique identifier serves as a multi-modal aggregation anchor, establishing stable mapping relationships between different levels such as individuals, sessions, or devices. Matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data are bound to the object unique identifier to form data relationship indices. Each data product is assigned a link entry, recording the object unique identifier, source marker, timestamp, and verification summary. These link entries support bidirectional cross-modal navigation and time-series playback, ensuring the location and association of different modalities within the same object.
[0105] Finally, matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data are merged to generate a preprocessed dataset. The merging process is not a simple splicing; rather, it involves alignment and deduplication based on data relationships, ensuring that only one valid record is retained for the same time period, the same object, and the same semantic unit. The preprocessed dataset is organized using partitioning and hierarchical structures. Partitions are based on time windows and source markers, while hierarchical structures manage the original products, derived products, and metadata separately. Each record within the dataset carries a unique object identifier and a link entry, allowing direct retrieval by object and cross-modal access in subsequent fusion processing modules. To avoid inconsistencies across batches, consistency and integrity checks are performed on the preprocessed dataset after generation. Records that fail the checks are returned to the cleaning and transformation stages for reprocessing. Through these connections, acquisition, cleaning, transformation, identifier extraction, binding, and merging form a closed loop, outputting a stable input base that can be used for cross-modal and attention-based interactions.
[0106] This embodiment constructs a traceable data pipeline in sequence, consisting of a multi-source data acquisition interface, a cleaned dataset, matrix-formatted image data, spectrogram speech data, normalized numerical data, vectorized text data, object unique identifiers, data association links, and a preprocessed dataset. Object unique identifiers and data association links align disparate modalities within the same object and time window, reducing cross-modal mismatches and redundant calculations. Matrixing and time-frequency conversion unify heterogeneous media into a unified numerical domain, lowering the alignment cost of subsequent fusion. Normalization and vectorization improve numerical stability and semantic density, enhancing the distinguishability of downstream recognition and fusion. This structurally alleviates the information loss caused by a single modality, improves input quality through consistency and integrity checks, provides stable input for cross-modal interaction processing and attention-based interaction processing, and ultimately supports the accurate generation of dynamic object profiles and personalized optimization path sets.
[0107] In one embodiment, step S20 above includes:
[0108] S201, based on the data relationship index, extract text data, image data, voice data and numerical data of the same object from the preprocessed data set;
[0109] S202, Perform cross-modal feature alignment operation on the text data and image data to generate text-image alignment features;
[0110] S203, Perform cross-modal feature interaction operation on the speech data and numerical data to generate speech-numerical interaction features;
[0111] S204, Extract sentiment features from the text image alignment features;
[0112] S205, Extract decision style features from the voice numerical interaction features;
[0113] S206, Perform an attention weighting operation on the sentiment tendency features to generate weighted sentiment tendency features;
[0114] S207, Perform an attention weighting operation on the decision style features to generate weighted decision style features;
[0115] S208, the weighted sentiment tendency features and weighted decision style features are integrated to generate a comprehensive feature vector.
[0116] In this embodiment, text data, image data, voice data, and numerical data of the same object are retrieved from the preprocessed dataset based on data association. The retrieval process uses the object's unique identifier and time tag as the search key, prioritizing sample segments within overlapping time intervals. When multiple candidates exist within the same time window, the one with the most complete source tag is retained, and the reason for discarding it is recorded to ensure the consistency and traceability of downstream inputs. After the retrieval is completed, a cross-modal alignment batch is constructed. Each record in the batch carries the object's unique identifier, source tag, and time tag, and subsequent steps maintain the alignment relationship through these tags.
[0117] Text and image data undergo cross-modal feature alignment. On the text side, vectorized text data is first mapped to a context-sensitive sequence representation, preserving word and paragraph boundaries. On the image side, matrix-formatted image data is divided into region units, and region representations are extracted, preserving spatial coordinates and scale information. Both representations are projected onto a unified representation domain. Within this domain, mapping relationships are established based on the indicative relationships between text fragments and image regions, generating text-image alignment features. These mapping relationships are stored in a one-to-one or many-to-many format, accompanied by confidence and spatial location information, avoiding the separation of semantics and visual cues.
[0118] Speech and numerical data undergo cross-modal feature interaction. The spectrogram speech data is first aligned to a fixed time base, preserving frame boundaries and energy peak information. Numerical data is resampled and slid-aggregated along time labels to generate temporal segments synchronized with the speech frames. These two temporal representations converge to the same temporal grid, establishing frame-level coupling and performing information interaction on this coupling, outputting speech-numerical interaction features. These features simultaneously contain covariation cues of prosodic rhythm and key numerical trajectories, and retain timestamp indices for subsequent localization.
[0119] Sentiment features are extracted from text-image alignment features. During extraction, a semantic-visual channel matching gate is set, prioritizing the retention of common cues between semantic segments and high-confidence visual regions, while suppressing background regions and meaningless word segments unrelated to the object. A stable set of sentiment vector components is output, and the contribution source is recorded in the metadata to ensure interpretability. Decision-making style features are extracted from speech-numerical interaction features. During extraction, a rhythm window and a persistence window are set in the temporal dimension, granting greater priority to segments coupled with high-persistent rhythm and key numerical changes, forming vector components that reflect style characteristics such as robustness, hesitation, or impulsivity.
[0120] Sentiment characteristics are then processed through an attention-weighted operation. The weighting mechanism assigns weights to each component based on contextual relevance and alignment confidence, and then normalizes and converges these weights to obtain weighted sentiment characteristics. Weights on low-confidence or conflicting components are suppressed to avoid noise amplification. Decision style characteristics are also processed through an attention-weighted operation. The weighting mechanism assigns weights to each component based on temporal persistence and cross-modal consistency, and then converges these weights to obtain weighted decision style characteristics. Short-term isolated components are weakened when there is no consistency support, ensuring stable style characterization.
[0121] Weighted sentiment features and weighted decision style features are then fused. The fusion process begins with dimensional and scale alignment, followed by gated splicing and learnable projection to form a fixed-length representation, outputting a comprehensive feature vector. This vector retains source markers, unique object identifiers, and time-stamped reference indices, ensuring that the original text fragment, image region, audio frame, and numerical segment can be traced back at any stage. The construction of the comprehensive feature vector adheres to constraints of single-object, synchronous time window, and cross-modal consistency, avoiding information leakage and cross-object crosstalk, providing a stable, compact, and interpretable input for subsequent recognition stages.
[0122] This embodiment utilizes data relationships to achieve multimodal retrieval and alignment across both object and time dimensions, enabling text, images, speech, and numerical values to converge within the same object and synchronization window, reducing semantic drift caused by mismatches. Text-image alignment features and speech-numerical interaction features respectively carry semantic-visual cues and prosodic-behavioral cues, covering complementary dimensions of emotion and decision-making style. Attention weighting assigns weights to source confidence, temporal persistence, and cross-modal consistency, suppressing noise and isolated fragments and improving representation purity. Finally, after dimensional and scale alignment, a comprehensive feature vector is output, with fixed length, traceable source, and cross-modal consistency, providing a stable input foundation for subsequent dynamic object profiling. This also improves recognition accuracy and convergence efficiency with equal amounts of data, while reducing sensitivity to quality fluctuations in a single modality.
[0123] In one embodiment, step S30 above includes:
[0124] S301, Input the comprehensive feature vector into the demand identification model to perform demand identification operation and generate demand identification result;
[0125] S302, Extract the demand intensity item from the demand identification result;
[0126] S303, input the comprehensive feature vector into the preference recognition model to perform preference recognition operation and generate preference recognition result;
[0127] S304, Extract preference weight terms from the preference recognition results;
[0128] S305, input the comprehensive feature vector into the emotion recognition model to perform emotion recognition operation and generate emotion recognition result;
[0129] S306, Extract the emotion state item from the emotion recognition result;
[0130] S307, Perform a profile integration operation based on the demand intensity item, preference weight item, and emotional state item to generate a dynamic profile of the object.
[0131] In this embodiment, the integrated feature vector undergoes input normalization and channel segmentation before entering the three recognition channels. Input normalization ensures scale and encoding consistency without altering semantic information, guaranteeing comparable representations of the same object across different time slices. Channel segmentation constructs a routing table based on feature subdomains related to sentiment, preference, and needs, marking scenario-independent components as low priority to avoid interfering with downstream recognition. The segmented integrated feature vector is simultaneously delivered to the need recognition model, preference recognition model, and sentiment recognition model, executing in parallel to maintain consistency between time labels and object identifiers.
[0132] The demand identification model receives demand subdomain representations from comprehensive feature vectors and outputs demand identification results through hierarchical representation aggregation and saliency filtering. Hierarchical aggregation focuses on triggering cues consistent across modalities, while saliency filtering suppresses isolated cues with a single source and no cross-support, outputting structured results containing candidate demand categories, confidence information, and trigger fragment references. Demand intensity extraction uses the demand identification results as the sole source, reads the corresponding channel for the target category, and sets a threshold based on cross-time window stability and cross-modal support. Components that do not meet stability or support requirements are downweighted and removed, while continuous components that pass through the threshold are retained, mapped to single-object demand intensity terms, and their contribution sources are recorded to ensure subsequent traceability.
[0133] The preference recognition model receives preference subdomain representations from a comprehensive feature vector and generates preference recognition results through multi-head attention routing and mutual exclusion constraints. Multi-head attention routing establishes independent perspectives for different preference dimensions, while mutual exclusion constraints resolve contradictory candidate preferences within the same dimension. The output is a labeled preference vector set, carrying dimension identifiers and confidence markers. Preference weight extraction uses this vector set as input. First, conflict resolution is performed within the same dimension, suppressing components with low confidence and no cross-modal support. Then, cross-dimensional normalization and aggregation are performed to obtain preference weights with stable ranking relationships and clear labels. Simultaneously, the alignment index with the time label is retained to ensure direct reference in subsequent time-related generation stages.
[0134] The emotion recognition model receives the emotion subdomain representation from the comprehensive feature vector and outputs the emotion recognition result through the collaborative processing of short-term and long-term windows. The short-term window captures instantaneous cues, while the long-term window assesses persistent trends. The two are fused with confidence in the consistency judgment unit to avoid misjudgments caused by a single instantaneous perturbation. The extraction of the emotion state item is based on the fused result, establishing a ternary representation of state label, duration, and fluctuation features. When multiple states coexist, the state with stronger persistence within the long-term window is retained first, and the trigger fragment reference of the state transition is recorded, and the emotion state item is output.
[0135] The profile integration operation uses demand intensity, preference weight, and emotional state as the only inputs to construct a unified field space and complete alignment and fusion. In the alignment phase, the time tags of the three items are first aligned, and then scale consistency is achieved to eliminate dimensional differences. When missing items exist, a gap-filling strategy is triggered, adhering to a conservative principle of only filling gaps when cross-modal support is available to avoid misleading subsequent generation. In the fusion phase, a field dependency graph is established, using demand intensity as the driving field, preference weight as the configuration field, and emotional state as the moderating field, and writing them into a structured profile container according to their dependencies. The container also stores the reference index and source tag of each field, forming a dynamic object profile. The dynamic object profile provides a machine-readable, traceable, and updatable structured characterization of the same object within a unified time slice, directly providing key-value input for subsequent decision generation and path combination.
[0136] This embodiment decouples and quantifies the three types of cues—demand, preference, and emotion—in the comprehensive feature vector through unified input, three-way parallel recognition, and field integration, reducing the interference of cross-modal noise on single-path judgment. The demand intensity item, preference weight item, and emotion state item are consistent in both time and scale dimensions, reducing the risk of mismatch caused by different sources and different time windows. The profile integration operation completes the structured disking by using the dependency relationship between driving fields, configuration fields, and adjustment fields, enabling the dynamic profile of the object to have stable fields, traceable references, and incremental update capabilities.
[0137] In one embodiment, step S40 above includes:
[0138] S401, Load the scene strategy entry set and resource entry library;
[0139] S402, extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object;
[0140] S403, based on the demand intensity item and the emotional state item, match the target strategy item from the set of scenario strategy items;
[0141] S404, Based on the demand intensity item and preference weight item, match the target resource item from the resource item library;
[0142] S405, Generate a product matching optimization path based on the target resource item;
[0143] S406, Generate an interaction strategy optimization path based on the target strategy item and the emotion state item;
[0144] S407, combine the product matching optimization path with the interaction strategy optimization path to generate a personalized optimization path set.
[0145] In this embodiment, the loading of the scenario strategy entry set and resource entry library provides a searchable entry space and resource space for subsequent matching processes. The scenario strategy entry set is organized by entry as the basic unit. Each entry includes a scenario tag, a reach channel tag, a rhythm tag, a profile adaptation tag, a priority tag, and an effectiveness boundary tag. The indexing method uses an object-independent multi-key index, facilitating tag-based aggregation and rapid location. The resource entry library is organized by resource entry as the basic unit. Each resource entry includes a resource category tag, an attribute tag set, an availability tag, a compliance tag, and an update tag. An object-independent inverted index and lightweight caching are established to ensure high-concurrency retrieval. During loading, integrity checks and tag standardization are performed, and abnormal entries are moved to an isolation area to prevent contamination of subsequent path generation.
[0146] When extracting the demand intensity, preference weight, and emotional state items from the dynamic profile of an object, the original field names and timestamps are kept unchanged. First, the timestamps of the three items are aligned, using window alignment and gap placeholders to ensure comparability within the same time frame. The demand intensity item serves as the driving field, determining the resource selection scope and strategy activation gate; the preference weight item serves as the configuration field, determining resource attribute bias and presentation order; and the emotional state item serves as the adjustment field, controlling the reach channels, communication style, and pace. Before entering the matching stage, the three items undergo scale consistency to eliminate dimensional differences and are accompanied by source tags to support traceability.
[0147] When matching target strategy entries from the set of scenario strategy entries based on demand intensity and emotional state, the demand intensity item is first used to trigger scenario tag aggregation, limiting the scenario scope of candidate entries. Then, the emotional state item is used to trigger rhythm and channel filtering, eliminating entries incompatible with the current state. Subsequently, gating filtering is performed based on priority tags and effectiveness boundary tags to generate the set of target strategy entries. To avoid conflicts, a mutual exclusion graph is introduced among candidate entries. If a mutual exclusion relationship exists, the entry with higher priority and stronger consistency with the object's dynamic profile is retained; if a dependency relationship exists, entries are grouped according to the dependency relationship, and the dependency order is recorded. Throughout the process, the original identifiers and tags of the entries remain unchanged, facilitating direct referencing in subsequent path generation.
[0148] When matching target resource entries from the resource entry library based on demand intensity and preference weight, the resource category and threshold label are first set using demand intensity to form an initial screening set. Then, the attribute label set is weighted and filtered using preference weight to retain resource entries that are more consistent with the preference. Next, usability and compliance labels are checked to remove unusable or non-compliant entries. Finally, newer entries are selected based on updated labels to reduce interference from outdated content. If there is attribute or functional overlap among target resource entries, deduplication and clustering are performed to obtain a set of target resource entries, while retaining the mapping relationship between entries and preference weight, supporting the determination of the information unfolding order in the path.
[0149] When generating a product matching optimization path based on target resource items, the target resource items are sorted into a directed sequence according to mapping relationships and preference weights. The path consists of nodes and connections. Nodes correspond to the display or interaction units of resource items, and connections carry entry and exit conditions. Entry conditions reference the demand intensity range and availability tags, while exit conditions reference lightweight event markers such as clicks, favorites, and inquiries. Branch nodes are introduced within the path to express the switching logic of replacing or supplementing resources, and converging nodes are introduced to return to the main presentation sequence after multiple branches. After the path is generated, the node sequence, node attributes, connection conditions, and reference indexes are output, forming the product matching optimization path.
[0150] When generating an interaction strategy optimization path based on target strategy items and emotional state items, the set of target strategy items serves as the framework, and the execution order of reach channels, speech style, and rhythm is written according to dependencies. Emotional state items are used to set adjustment parameters for tone strength, information density, and dwell time at each node; when the emotional state item indicates resistance or low receptivity, a slowing-down branch and a frequency-reducing branch are triggered; when the emotional state item indicates positive or high receptivity, an accelerating branch and an information-deepening branch are triggered. To ensure consistency, the binding to the original item labels is retained at the path nodes to avoid deviations in tone and channel from the strategy intent. The path output includes a node sequence, channel labels, speech style labels, rhythm parameters, and reference indexes, forming the interaction strategy optimization path.
[0151] When generating a personalized optimization path set by combining product matching optimization paths and interaction strategy optimization paths, the alignment coordinates of the two types of paths are first constructed. A dual-coordinate approach using time and event coordinates ensures that the display sequence and the reach sequence can be executed in parallel or sequentially. Dependency checks are then performed to establish subordinate or parallel relationships between interaction nodes and product nodes. Subordinate relationships are used to insert specific dialogue or confirmation actions before or after a specified display node, while parallel relationships are used to send light reach prompts synchronously during the display. Conflict detection is performed again. If two path nodes produce incompatible channel and display behaviors at the same coordinate, a comprehensive judgment is made based on the priority of the target strategy item and the availability of resource items, retaining the side with lower execution risk and generating a delayed node for the other side. Finally, a personalized optimization path set is output, containing several paths that can be executed independently or collaboratively. Each path maintains the integrity of nodes, connections, parameters, and reference indexes, facilitating direct reference in subsequent personalized solution generation.
[0152] This embodiment achieves direct alignment between strategy retrieval and resource retrieval at the tag level through domain-specific loading and standardization of the scenario strategy item set and resource item library, reducing cross-domain mapping overhead; the time and scale of demand intensity items, preference weight items, and emotional state items are consistent before entering the matching stage, reducing mismatch and drift; the dual-channel matching of target strategy items and target resource items ensures that the reach path and content path are independent yet constrainedly integrated, avoiding information imbalance caused by a single path; the product matching optimization path and the interaction strategy optimization path are combined into a personalized optimization path set under dual coordinates, and the display and reach are coordinated and executed according to dependency and parallel relationships, improving the coupling between reach timing and content presentation.
[0153] In one embodiment, step S50 above includes:
[0154] S501, Separate the product matching optimization path and the interaction strategy optimization path from the set of personalized optimization paths;
[0155] S502, Based on the demand intensity item in the dynamic profile of the object and the product matching optimization path, determine the product combination scheme;
[0156] S503, Based on the aforementioned demand intensity item, determine the follow-up timeline plan;
[0157] S504, Based on the preference weight items in the object dynamic profile and the product matching optimization path, determine the key content of the product introduction;
[0158] S505, Based on the emotional state item in the dynamic profile of the object and the interaction strategy optimization path, determine the communication skills;
[0159] S506, combine the product portfolio plan, key product introduction content, communication skills and follow-up timeline plan to generate a personalized plan.
[0160] In this embodiment, when separating the product matching optimization path and the interaction strategy optimization path from the personalized optimization path set, it is necessary to classify them according to the source tags of the nodes in the path set. The source tags are already bound to the node attributes during the path set generation stage and can be directly used to identify whether a node belongs to a product path or a strategy path. The separation process maintains the node order and connection conditions of each path unchanged, while recording the mapping relationship between the path and dynamic profile elements to ensure that the matching basis can be traced back when combining schemes later. During separation, it is also necessary to handle possible cross nodes, such as composite nodes that are associated with both resource display and channel strategy. These nodes need to be split into two parallel child nodes, which are retained in their respective paths, and the dependency relationship tags are maintained.
[0161] When determining product portfolio solutions based on the demand intensity item and product matching optimization path in the dynamic profile, the nodes in the product matching optimization path are first filtered using the numerical range of the demand intensity item, eliminating resource nodes whose demand intensity is insufficient to trigger. The remaining nodes are combined into a candidate product sequence according to the sorting priority in the preference weight item and the original execution order of the path. In this sequence, the portfolio structure is determined through attribute clustering or functional complementarity analysis. For example, products of the same category are grouped according to differentiated attributes, and products of different categories are combined according to functional links, forming multiple candidate portfolio solutions. When finally selecting a portfolio solution, availability and compliance checks can be introduced to eliminate combinations that are unavailable or do not meet specific regulatory requirements.
[0162] When determining the follow-up timeline based on demand intensity, the demand intensity is combined with the path execution time model to generate the time interval configuration between nodes. The time interval can be dynamically adjusted according to the demand intensity; for example, high demand intensity corresponds to a shorter time between the first and second reach, while low demand intensity extends the interval to avoid excessive disruption. For multiple parallel paths, time nodes need to be coordinated to ensure that different reach does not conflict or overlap. Simultaneously, event triggering conditions are introduced; when demand intensity changes during the execution cycle, the time nodes can be recalculated and adjusted in real time.
[0163] When determining the key content of a product introduction based on preference weights in dynamic profiles and product matching optimization paths, preference weights are mapped to attribute tags on product nodes, selecting attributes with higher weights as the focus of the introduction. For example, if a resource item has multiple attribute tags (price, function, brand, service, etc.), the weight of the corresponding tag in the preference weights determines the depth and order of content display. In actual content generation, information output can be arranged in descending order of weight; attributes with similar weights can be displayed side-by-side or presented in stages to avoid information overload. The key content of the introduction is linked to the product portfolio plan, ensuring that the introduction content of each product in the portfolio aligns with customer preferences.
[0164] When determining communication techniques based on emotional state items in dynamic profiles and interaction strategy optimization paths, the emotional state items are first mapped to node labels in the strategy path to identify the necessary adjustments to the outreach methods. For example, when the emotional state is positive, a more direct, facilitating approach can be chosen; when the emotional state is neutral or hesitant, an explanatory or reassuring approach can be selected; and when the emotional state is resistant, a reduced frequency or shift to indirect guidance is triggered. Communication techniques include not only adjustments to the content itself but also the configuration of non-verbal parameters such as tone, rhythm, and interaction frequency. When generating communication techniques, the channel labels of the strategy path nodes need to be referenced to ensure that the communication matches the characteristics of the channel; for example, telephone channels emphasize tone of voice, while text channels emphasize conciseness and keywords.
[0165] When generating a personalized solution by combining product portfolio strategies, key product introductions, communication techniques, and follow-up timelines, it's crucial to establish a dependency order and synchronization mechanism among these elements. During the combination process, the product portfolio strategy serves as the core carrier, with key product introductions bound to it. Communication techniques are embedded into corresponding product or strategy nodes according to the order of outreach. The follow-up timeline serves as the overall pacing control framework, coordinating the sequence and intervals of different outreaches. The final personalized solution is structurally represented as a multi-layered relationship diagram. Each product and strategy node includes introductory content, communication techniques, and timeline configurations, ensuring that execution unfolds according to a profile-driven logical sequence.
[0166] This embodiment refines the personalized optimization path set into product paths and strategy paths, enabling independent optimization of path content for different user profile elements. It utilizes demand intensity-driven product selection and timing adjustments to ensure a match between reach timing and resource delivery. Preference weighting guides the generation of key information, ensuring high consistency between displayed content and user preferences. Emotional state-controlled communication techniques enhance the contextual adaptability of the communication process. The resulting personalized solution achieves synergy between displayed content, reach methods, and execution pace, significantly improving the solution's targeting and conversion potential.
[0167] In one embodiment, after step S50 above, the method further includes:
[0168] S601, the personalized plan is pushed through the terminal interface;
[0169] S602, Monitor emotional state change data during the execution of the personalized plan;
[0170] S603, when the emotional state change data reaches the resistance threshold, a strategy adjustment prompt is generated;
[0171] S604, record the execution response time of the strategy adjustment prompt;
[0172] S605 records terminal interaction behavior data;
[0173] S606, merge the emotional state change data, execution response time, and terminal interaction behavior data to generate execution process feedback data;
[0174] S607, using the execution process feedback data, perform incremental training operations on the demand recognition model, preference recognition model, and emotion recognition model, and update the demand recognition model parameters, preference recognition model parameters, and emotion recognition model parameters in the model parameter set;
[0175] S608, Prioritize the strategy entries in the set of scenario strategy entries based on the execution response time optimization;
[0176] S609, update the resource item weights in the resource item library based on the terminal interaction behavior data.
[0177] In this embodiment, when a personalized solution is pushed to the terminal interface, the personalized solution is structured into two parts: renderable data and executable instructions. The renderable data includes fields such as product combinations, key product introduction points, communication scripts, and follow-up time nodes. The executable instructions are used to trigger the display order of interface components and the interaction entry point. Before pushing, identity verification and session binding are completed, and the object's unique identifier, session identifier, and personalized solution version identifier are written into the context cache to ensure that subsequent collection and updates can be accurately traced back to the same execution process. The push action is completed through a message channel or interface call, and the return value carries the delivery time and presentation status, which serves as the base timestamp for subsequent calculation of execution response time and alignment of interaction logs.
[0178] When monitoring changes in emotional state during the implementation of personalized solutions, emotional signals from sources such as text, speech, and images are collected uniformly. Text sources are mapped to the emotion dimension through word segmentation and semantic vectorization; speech sources are mapped to the intensity and fluctuation trend of emotions through time-frequency domain features; and image sources are mapped to emotion labels through facial expression key points and expression categories. Emotional signals from each source are aligned according to the unique identifier of the object and the session identifier, and resampling and missing data completion are performed on a unified timeline to output a continuous emotional state change data stream. Source labels and confidence scores are added simultaneously to support threshold judgment and sample selection for subsequent model training.
[0179] The resistance threshold is used to trigger a strategy adjustment prompt when emotional state change data reaches a preset boundary. The threshold is derived from historical conversation statistics and offline verification, and stored as a threshold table or threshold function that can be loaded online. During online judgment, emotional intensity and rate of change are calculated using a sliding window. The aggregated results within the window are compared with the threshold; once the trigger condition is met, a strategy adjustment prompt is generated. The strategy adjustment prompt includes fields such as prompt type, associated channel, suggested script tag, and frequency reduction strategy flag, and carries the trigger time and trigger reason, facilitating subsequent calculation of execution response time and evaluation of strategy effectiveness.
[0180] Execution response time measures the delay between the issuance of a policy adjustment prompt and its completion, whether manually or automatically. For accurate measurement, a trigger timestamp is written when the policy adjustment prompt is generated, and an execution completion timestamp is sent back when the operation is completed at the execution end. The difference is calculated and recorded by the data collection layer. If the execution process involves multiple steps, representative timestamps are extracted from predefined key nodes, such as receiving, viewing, adopting, and implementing, and the node sequence is retained in the response time record, supporting fine-grained bottleneck localization.
[0181] Terminal interaction behavior data covers events such as clicks, swipes, pauses, jumps, dialing, callbacks, content expansion, and form submissions. Each event, along with a unique object identifier, session identifier, UI component identifier, timestamp, event type, and event context, is written into the event stream. To avoid noise interference, the event stream undergoes deduplication and anomaly filtering before being stored in the database. Short-cycle, highly repetitive events are merged into a single valid behavior, and behavior counts and pause distributions are summarized by session phase for subsequent weight updates.
[0182] The execution process feedback data is obtained by merging emotional state change data, execution response time, and terminal interaction behavior data. During merging, the session identifier and timeline are used as the primary keys for alignment, constructing segmented sample units. Each sample unit includes the stage start and end times, a summary of the emotional curve within the stage, statistics on response time within the stage, statistics on interaction behavior within the stage, and the identifier of the personalized solution fragment corresponding to the stage. For records that cannot be aligned, a nearest neighbor completion or discard strategy is used, and alignment quality markers are retained in the sample units for weighting during training.
[0183] The model parameter set is updated incrementally, covering the demand identification model, preference identification model, and emotion identification model. Feedback data from the execution process is mapped to usable training samples and supervision signals for the model: the demand identification model uses event density and conversion indicators related to product interaction as target signals; the preference identification model uses interaction distribution and dwell distribution of different content tags as target signals; and the emotion identification model uses multi-source emotion tags and intensity trajectories as target signals. The training process uses small-batch online updates, first performing sample cleaning and class equalization, then completing parameter updates and offline verification. After successful verification, the current effective version of the model parameter set is written, and the version number and effective time are recorded to support rollback.
[0184] Priority optimization of the scenario strategy item set depends on the execution response time and the trigger records of strategy adjustment prompts. For each strategy item, the number of triggers, adoption rate, and response latency under different scenarios are statistically analyzed, and a ranking index is constructed based on latency thresholds and adoption results. Smoothing is performed during ranking updates to avoid over-adjustment caused by short-term fluctuations. After the update is complete, the priority field of the item header is written back and synchronized to the online matching component, making subsequent matching more biased towards strategy items with fast response and high adoption rates.
[0185] The weighting of resource items relies on terminal interaction behavior data and product exposure and interaction records in personalized solutions. For each resource item, the conversion ratio between exposure, effective interaction, and target behavior trigger is calculated, and this is combined with tiered statistics based on target user profiles, with different tiers using their own weighting coefficients. Weighting updates employ a decay-cumulative approach, balancing historical stability with the latest session performance. After the update, during the resource retrieval and product portfolio generation stages, recall and reordering are performed according to the new weights, giving high-value resource items a higher probability of presentation.
[0186] Example Description: In a large comprehensive hospital's intelligent follow-up and personalized health management system, a multi-source data acquisition interface is first deployed to connect with the hospital's electronic medical record system, image archiving and communication system, wearable health monitoring devices, patient mobile applications, and online consultation platforms to acquire multimodal data covering both structured and unstructured data. This data includes patient medical records, test and examination results, CT and MRI images, electrocardiograms, voice consultation recordings, online questionnaires, and heart rate and blood pressure monitoring data from wearable devices. After data access, a unified cleaning process is performed to remove invalid, duplicate, and abnormally formatted data. Image data is converted to matrix format, voice data is converted to spectrograms, numerical data is normalized, and text data is vectorized to ensure that different modalities can be processed uniformly. Subsequently, a unique identifier for each patient (such as a patient card number + encryption rules) is extracted, and the aforementioned multimodal data is bound to the unique identifier to generate data associations. Simultaneously, the different modal data are merged into a preprocessed dataset.
[0187] In the fusion processing module, the system matches text data (medical record summaries, consultation records), image data (X-rays, CT scans), voice data (consultation recordings, rehabilitation guidance feedback), and numerical data (monitoring values such as blood glucose, blood pressure, and heart rate) of the same patient based on data relationship indexing. For text and images, cross-modal feature alignment is performed, for example, mapping lesion areas in radiological images to pathological features in medical record descriptions to generate text-image alignment features. For voice and numerical data, cross-modal interaction analysis is performed, combining changes in the patient's voice intonation with fluctuations in physiological indicators to generate voice-numerical interaction features. Based on this, affective information (such as emotional changes when describing pain locations) is extracted from the text-image alignment features, and decision-making style information (such as the patient's hesitation or determination when choosing a treatment plan) is extracted from the voice-numerical interaction features. Affective and decision-making style features are each attention-weighted to highlight the signals most valuable for personalized health management, and finally fused into a comprehensive feature vector for the patient.
[0188] The comprehensive feature vector is input into a multi-model recognition system, including a needs recognition model, a preference recognition model, and an emotion recognition model. The needs recognition model analyzes recent examination results and disease progression to identify the patient's most pressing health needs, such as whether further imaging examinations, rehabilitation training, or medication adjustments are required, and extracts the intensity of these needs. The preference recognition model analyzes the patient's past medical behavior, acceptance of different treatment methods, and lifestyle choices, extracting preference weights, such as whether the patient prefers conservative treatment or surgical intervention. The emotion recognition model uses voice, facial expressions, and textual feedback to determine the patient's emotional state, such as anxiety, calmness, or positivity. The results of the three recognitions are integrated to generate a dynamic patient profile for subsequent intervention design.
[0189] When generating personalized optimization paths, the system loads a set of medical scenario strategy items (such as follow-up frequency, intervention methods, and communication channels) and a medical resource item library (including medicines, rehabilitation equipment, examination items, and health education materials). It extracts demand intensity, preference weight, and emotional state items from the patient profile. Based on the demand intensity and emotional state items, it matches the most suitable health management strategy from the strategy item set, such as frequent remote health monitoring or face-to-face consultation. Based on the demand intensity and preference weight, it matches corresponding medical resources from the resource item library. For example, if a patient prefers traditional Chinese medicine and has a high demand intensity, it prioritizes recommending traditional Chinese medicine departments and related drug resources. The matching results are used to generate product matching optimization paths (medical service and resource combinations) and interaction strategy optimization paths (communication frequency, methods, and key points of communication), which are ultimately combined into a personalized optimization path set.
[0190] When generating personalized plans, the personalized optimization path set is separated into product matching optimization path and interaction strategy optimization path. Based on the intensity of needs in the patient profile and the product matching optimization path, a combination of treatment and rehabilitation resources is determined, such as rehabilitation training + nutritional supplementation + regular imaging follow-ups. Based on the intensity of needs, a follow-up and intervention timeline is determined, such as weekly follow-ups for high-risk patients. Based on the preference weighting and the product matching optimization path, key health content to be emphasized in communication and intervention is determined. Based on the emotional state and the interaction strategy optimization path, communication scripts and care methods are determined, such as using more reassuring and explanatory communication in anxious states. Finally, the combination of treatment resources, key health content, communication strategies, and timelines are combined into a complete personalized health management plan.
[0191] In the solution delivery and closed-loop feedback phase, the system delivers personalized solutions through the doctor's workstation interface or the patient's mobile device. During execution, it monitors changes in the patient's emotional state in real time, such as through voice feedback and online interactive content. Once the emotional state reaches a resistance threshold (e.g., showing significant rejection of the treatment plan), the system generates a strategy adjustment prompt, alerting the doctor or health manager to adjust communication methods or intervention content. The system records the execution response time of the strategy adjustment prompt (time from prompt to completion) and the patient's interactive behavior data on the terminal (e.g., whether they viewed the solution, clicked on health education videos, or filled out follow-up questionnaires). The emotional state change data, response time, and interactive behavior data are combined into execution process feedback data, which is used for incremental training of the needs identification model, preference identification model, and emotion identification model, allowing the models to more closely reflect the patient's actual state through continuous iteration. Simultaneously, the system optimizes the priority of medical scenario strategy items based on response time and updates the weight of medical resource items based on interactive behavior data, making future personalized solutions more accurate in resource and strategy selection.
[0192] In the intelligent wealth management and personalized investment advisory platform, the first step is to deploy multi-source data acquisition interfaces to establish data connections with the bank's core business system, customer relationship management system, third-party credit reporting platform, transaction matching system, and financial information aggregation platform, acquiring multimodal data covering both structured and unstructured data. This data includes customers' historical transaction records, account asset and liability information, credit scores, risk tolerance assessment results, online customer service dialogue records, investment questionnaire feedback, social media financial opinion interaction records, video or voice investment consultation recordings, and real-time market data. After integration, a unified cleaning process is performed to remove expired, missing, or abnormal data; image data (such as ID card images and scanned copies of signed documents) is converted into matrix format; voice data is converted into spectrograms; transaction and market numerical data are normalized; and text data (news sentiment, dialogue records) is vectorized. Subsequently, a unique identifier (such as customer ID + encryption rules) is extracted for each customer, and all modal data are bound to this unique identifier to generate data association indices. Simultaneously, different modal data are merged into a preprocessed data set.
[0193] In the fusion processing module, the system extracts text data (customer service communication content, risk questionnaire answers, social media opinions), image data (signed document images, KYC information), voice data (telephone financial consultation recordings), and numerical data (transaction records, net asset value curves, market volatility indicators) from the same customer based on data relationship indexes. Text and image data undergo cross-modal feature alignment, correlating key information in the document images with the financial product preferences mentioned in the customer's text statements to generate text-image alignment features. Voice and numerical data undergo cross-modal feature interaction, combining emotional fluctuations in voice communication with market trends or trading behavior data to generate voice-numerical interaction features. Investment sentiment tendencies (e.g., cautious, aggressive) are extracted from the text-image alignment features, and decision-making style features (e.g., preference for short-term trading or long-term holding) are extracted from the voice-numerical interaction features. These are then processed with attention weighting to highlight the signals most influential on investment strategy formulation in the investment advisory scenario, and fused to form a comprehensive feature vector.
[0194] The comprehensive feature vector is input into a cluster of recognition models, including a demand recognition model, a preference recognition model, and an emotion recognition model. The demand recognition model analyzes the strength of a customer's investment needs in the current market cycle, such as whether they have a strong need for capital appreciation, liquidity, or hedging. The preference recognition model extracts preference weights by combining behavioral data such as historical portfolios, trading frequency, and product category selection, for example, a preference for equity assets, fixed-income assets, or alternative investments. The emotion recognition model extracts emotional state items through voice tone, text expression, and interactive content to determine whether the customer is optimistic, cautious, or panicked. These recognition results are integrated to form a dynamic customer profile, used for generating personalized investment advisory strategies.
[0195] When generating a personalized optimization path, the system loads a set of financial scenario strategy items (such as asset allocation strategies, transaction execution strategies, and risk management rules) and a library of financial product resources (including funds, stocks, bonds, structured financial products, and insurance). It extracts demand intensity, preference weight, and emotional state items from the customer profile. Based on demand intensity and emotional state, it matches target strategy items from the strategy item set that match the customer's risk tolerance and market sentiment. Based on demand intensity and preference weight, it matches target product items from the resource item library that match the customer's investment preferences. It then generates a product matching optimization path (corresponding to specific product portfolio recommendations) and an interaction strategy optimization path (corresponding to investment advisor communication methods, push notification frequency, and information depth), and combines them into a personalized optimization path set.
[0196] When generating personalized investment plans, the optimization path set is separated into product matching optimization paths and interaction strategy optimization paths. Based on the demand intensity factor and the product matching optimization path, a portfolio plan is determined, for example, allocating a higher proportion of high-yield products under high demand intensity. Follow-up timing is determined based on demand intensity, for example, increasing follow-up frequency during periods of high short-term market volatility. Based on the preference weight factor and the product matching optimization path, the product characteristics that need to be emphasized during communication are identified, for example, highlighting the coupon yield and risk level for clients with a fixed-income preference. Based on the emotional state factor and the interaction strategy optimization path, communication techniques are developed, for example, providing more information on market risk mitigation and medium- to long-term investment logic for clients experiencing panic. Finally, the portfolio plan, communication focus, communication techniques, and follow-up timing are combined into a complete personalized investment advice plan.
[0197] During the solution delivery and closed-loop feedback phase, the system pushes personalized investment solutions through the financial manager's work interface or the customer's mobile device. During execution, it monitors changes in the customer's emotional state in real time, capturing emotional fluctuations through online interactions, video conferences, and voice communications. When the emotional state reaches a resistance threshold (such as rejecting product recommendations or questioning investment logic), a strategy adjustment prompt is generated, reminding the financial manager to adjust communication strategies or re-optimize product recommendations. The system records the execution response time of the strategy adjustment prompt and the customer's interactive behavior data on the terminal (such as whether they viewed solution details, placed an order, or added products to their favorites). The emotional state change data, execution response time, and interactive behavior data are combined to generate execution process feedback data, which is used for incremental training of the demand identification model, preference identification model, and emotion identification model to continuously improve recognition accuracy. Based on the execution response time, the system optimizes the priority of strategy items in the financial scenario strategy item set, and updates the weights of financial product resource items based on interactive behavior data, enabling subsequent solution generation to more accurately meet the customer's personalized investment needs.
[0198] This embodiment establishes a single session link and a unified timeline between push notifications, monitoring, judgment, recording, merging, and updating, creating a closed loop between personalized solution execution and feedback collection. It uses emotional state change data to trigger strategy adjustment prompts and quantifies execution response time, improving the reliability of strategy item priority ranking. Terminal interaction behavior data drives resource item weight updates, and execution process feedback data is used for incremental training of the three models, enabling synchronous evolution of parameters upon which profile updates and path generation depend. As a result, subsequent personalized solutions are more closely aligned with the target audience's state in terms of timing, content, and presentation order, and push efficiency and conversion potential steadily improve with iteration.
[0199] In one embodiment, a personalized scheme generation apparatus is provided, which corresponds one-to-one with the personalized scheme generation method described in the above embodiments. (Refer to...) Figure 3, Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the personalized solution generation device of the present invention. The modules include a preprocessing module 10, a cross-modal fusion processing module 20, a multi-dimensional feature recognition module 30, an optimized path generation module 40, and a personalized solution generation module 50. Detailed descriptions of each functional module are as follows:
[0200] Preprocessing module 10 is used to collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes;
[0201] The cross-modal fusion processing module 20 is used to perform cross-modal interaction processing and attention interaction processing on the preprocessed data set based on the data relationship index in the fusion processing module to generate a comprehensive feature vector;
[0202] The multi-dimensional feature recognition module 30 is used to input the comprehensive feature vector into the recognition model for recognition operation and generate a dynamic portrait of the object;
[0203] The optimized path generation module 40 is used to generate a set of personalized optimized paths based on the dynamic profile of the object.
[0204] The personalized solution generation module 50 is used to extract the demand intensity item, preference weight item and emotional state item from the dynamic profile of the object, and combine them with the personalized optimization path set to generate a personalized solution.
[0205] In one embodiment, the preprocessing module 10 is specifically used for:
[0206] Construct a multi-source data acquisition interface to collect multimodal data;
[0207] Perform a cleaning operation on the multimodal data to generate a cleaned data set;
[0208] Perform a format conversion operation on the image data in the cleaned dataset to generate matrix format image data;
[0209] Perform a format conversion operation on the speech data in the cleaned dataset to generate spectrogram speech data;
[0210] Normalization is performed on the numerical data in the cleaned dataset to generate normalized numerical data;
[0211] The text data in the cleaned dataset is vectorized to generate vectorized text data;
[0212] Extract unique identifiers for objects from the cleaned dataset;
[0213] The matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data are bound to the unique identifier of the object to generate a data relationship index;
[0214] The matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data are merged to generate a preprocessed data set.
[0215] In one embodiment, the cross-modal fusion processing module 20 is specifically used for:
[0216] Based on the data relationship index, text data, image data, voice data, and numerical data of the same object are extracted from the preprocessed data set;
[0217] Perform cross-modal feature alignment on the text and image data to generate text-image alignment features;
[0218] Perform cross-modal feature interaction operations on the speech data and numerical data to generate speech-numerical interaction features;
[0219] Extract sentiment characteristics from the text image alignment features;
[0220] Extract decision-making style features from the voice-numerical interaction features;
[0221] An attention-weighted operation is performed on the sentiment tendency features to generate weighted sentiment tendency features;
[0222] An attention-weighted operation is performed on the decision style features to generate weighted decision style features;
[0223] By integrating the weighted sentiment tendency features and the weighted decision style features, a comprehensive feature vector is generated.
[0224] In one embodiment, the multidimensional feature recognition module 30 is specifically used for:
[0225] The comprehensive feature vector is input into the demand identification model to perform a demand identification operation and generate a demand identification result.
[0226] Extract the demand intensity term from the demand identification results;
[0227] The comprehensive feature vector is input into the preference recognition model to perform preference recognition operations and generate preference recognition results.
[0228] Extract preference weight terms from the preference identification results;
[0229] The comprehensive feature vector is input into the emotion recognition model to perform emotion recognition operations and generate emotion recognition results.
[0230] Extract the emotion state item from the emotion recognition results;
[0231] Based on the demand intensity item, preference weight item, and emotional state item, a profile integration operation is performed to generate a dynamic profile of the object.
[0232] In one embodiment, the optimized path generation module 40 is specifically used for:
[0233] Load the scene strategy entry set and resource entry library;
[0234] Extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object;
[0235] Based on the demand intensity item and the emotional state item, the target strategy item is matched from the set of scenario strategy items;
[0236] Based on the demand intensity item and preference weight item, target resource items are matched from the resource item library;
[0237] Generate a product matching optimization path based on the target resource entries;
[0238] Generate an interaction strategy optimization path based on the target strategy items and emotion state items;
[0239] The product matching optimization path and the interaction strategy optimization path are combined to generate a set of personalized optimization paths.
[0240] In one embodiment, the personalized scheme generation module 50 is specifically used for:
[0241] Separate the product matching optimization path and the interaction strategy optimization path from the set of personalized optimization paths;
[0242] Based on the demand intensity item in the dynamic profile of the object and the product matching optimization path, a product combination scheme is determined;
[0243] Based on the aforementioned demand intensity, a follow-up timeline plan is determined;
[0244] Based on the preference weight items in the dynamic profile of the object and the product matching optimization path, the key content of the product introduction is determined.
[0245] Based on the emotional state item in the dynamic profile of the object and the interaction strategy optimization path, communication skills are determined.
[0246] Combine the aforementioned product portfolio, key product introduction content, communication skills, and follow-up timeline to generate a personalized solution.
[0247] In one embodiment, the personalized scheme generation module 50 is specifically used for:
[0248] The personalized plan is pushed through the terminal interface;
[0249] Monitor emotional state changes during the implementation of the personalized plan;
[0250] When the emotional state change data reaches the resistance threshold, a strategy adjustment prompt is generated.
[0251] Record the execution response time of the strategy adjustment prompt;
[0252] Record terminal interaction behavior data;
[0253] The emotional state change data, execution response time, and terminal interaction behavior data are combined to generate execution process feedback data;
[0254] The incremental training operations of the demand recognition model, preference recognition model, and emotion recognition model are performed using the feedback data of the execution process, and the demand recognition model parameters, preference recognition model parameters, and emotion recognition model parameters in the model parameter set are updated.
[0255] Prioritize the strategy entries in the set of strategy entries for the scenario optimization based on the execution response time;
[0256] The weights of resource entries in the resource entry library are updated based on the terminal interaction behavior data.
[0257] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a personalized scheme generation method on the server side.
[0258] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the user-side functions or steps of a personalized scheme generation method.
[0259] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0260] Collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes;
[0261] In the fusion processing module, based on the data relationship index, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector.
[0262] The comprehensive feature vector is input into the recognition model for recognition operation to generate a dynamic portrait of the object;
[0263] Based on the dynamic profile of the object, a set of personalized optimization paths is generated;
[0264] Extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the set of personalized optimization paths to generate a personalized solution.
[0265] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0266] Collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes;
[0267] In the fusion processing module, based on the data relationship index, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector.
[0268] The comprehensive feature vector is input into the recognition model for recognition operation to generate a dynamic portrait of the object;
[0269] Based on the dynamic profile of the object, a set of personalized optimization paths is generated;
[0270] Extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the set of personalized optimization paths to generate a personalized solution.
[0271] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0272] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0273] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0274] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for generating personalized schemes, characterized in that, Includes the following steps: Collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes; In the fusion processing module, based on the data relationship index, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector. The comprehensive feature vector is input into the recognition model for recognition operation to generate a dynamic portrait of the object; Based on the dynamic profile of the object, a set of personalized optimization paths is generated; Extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the set of personalized optimization paths to generate a personalized solution; In the fusion processing module, based on the data relationship index, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector, including: Based on the data relationship index, text data, image data, voice data, and numerical data of the same object are extracted from the preprocessed data set; Perform cross-modal feature alignment on the text and image data to generate text-image alignment features; Perform cross-modal feature interaction operations on the speech data and numerical data to generate speech-numerical interaction features; Extract sentiment characteristics from the text image alignment features; Extract decision-making style features from the voice-numerical interaction features; An attention-weighted operation is performed on the sentiment tendency features to generate weighted sentiment tendency features; An attention-weighted operation is performed on the decision style features to generate weighted decision style features; By integrating the weighted sentiment tendency features and weighted decision style features, a comprehensive feature vector is generated; The process of fusing the weighted sentiment characteristics and weighted decision style characteristics to generate a comprehensive feature vector includes: The weighted sentiment characteristics and the weighted decision style characteristics are subjected to dimension alignment and scale alignment. The weighted sentiment features and the weighted decision style features, after being processed by dimension alignment and scale alignment, are fused to generate a comprehensive feature vector; The comprehensive feature vector is input into the recognition model for recognition operations to generate a dynamic portrait of the object, including: The comprehensive feature vector is input into the demand identification model to perform a demand identification operation and generate a demand identification result. Extract the demand intensity term from the demand identification results; The comprehensive feature vector is input into the preference recognition model to perform preference recognition operations and generate preference recognition results. Extract preference weight terms from the preference identification results; The comprehensive feature vector is input into the emotion recognition model to perform emotion recognition operations and generate emotion recognition results. Extract the emotion state item from the emotion recognition results; Based on the demand intensity item, preference weight item, and emotional state item, a profile integration operation is performed to generate a dynamic profile of the object.
2. The personalized scheme generation method as described in claim 1, characterized in that, Collect multimodal data and perform preprocessing operations to generate a preprocessed dataset and data associations, including: Construct a multi-source data acquisition interface to collect multimodal data; Perform a cleaning operation on the multimodal data to generate a cleaned data set; Perform a format conversion operation on the image data in the cleaned dataset to generate matrix format image data; Perform a format conversion operation on the speech data in the cleaned dataset to generate spectrogram speech data; Normalization is performed on the numerical data in the cleaned dataset to generate normalized numerical data; The text data in the cleaned dataset is vectorized to generate vectorized text data; Extract unique identifiers for objects from the cleaned dataset; The matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data are bound to the unique identifier of the object to generate a data relationship index; The matrix-formatted image data, spectrogram speech data, normalized numerical data, and vectorized text data are merged to generate a preprocessed data set.
3. The personalized scheme generation method as described in claim 1, characterized in that, Based on the dynamic profile of the object, a set of personalized optimization paths is generated, including: Load the scene strategy entry set and resource entry library; Extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object; Based on the demand intensity item and the emotional state item, the target strategy item is matched from the set of scenario strategy items; Based on the demand intensity item and preference weight item, target resource items are matched from the resource item library; Generate a product matching optimization path based on the target resource entries; Generate an interaction strategy optimization path based on the target strategy items and emotion state items; The product matching optimization path and the interaction strategy optimization path are combined to generate a set of personalized optimization paths.
4. The personalized scheme generation method as described in claim 1, characterized in that, Extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the personalized optimization path set to generate a personalized solution, including: Separate the product matching optimization path and the interaction strategy optimization path from the set of personalized optimization paths; Based on the demand intensity item in the dynamic profile of the object and the product matching optimization path, a product combination scheme is determined; Based on the aforementioned demand intensity, a follow-up timeline plan is determined; Based on the preference weight items in the dynamic profile of the object and the product matching optimization path, the key content of the product introduction is determined. Based on the emotional state item in the dynamic profile of the object and the interaction strategy optimization path, communication skills are determined. Combine the aforementioned product portfolio, key product introduction content, communication skills, and follow-up timeline to generate a personalized solution.
5. The personalized scheme generation method as described in claim 1, characterized in that, After extracting the demand intensity, preference weight, and emotional state items from the dynamic profile of the object, and combining them with the personalized optimization path set to generate a personalized solution, the process further includes: The personalized plan is pushed through the terminal interface; Monitor emotional state changes during the implementation of the personalized plan; When the emotional state change data reaches the resistance threshold, a strategy adjustment prompt is generated. Record the execution response time of the strategy adjustment prompt; Record terminal interaction behavior data; The emotional state change data, execution response time, and terminal interaction behavior data are combined to generate execution process feedback data; The incremental training operations of the demand recognition model, preference recognition model, and emotion recognition model are performed using the feedback data of the execution process, and the demand recognition model parameters, preference recognition model parameters, and emotion recognition model parameters in the model parameter set are updated. Prioritize the strategy entries in the set of strategy entries for the scenario optimization based on the execution response time; The weights of resource entries in the resource entry library are updated based on the terminal interaction behavior data.
6. A personalized solution generation device, characterized in that, The personalized scheme generation device includes: The preprocessing module is used to collect multimodal data and perform preprocessing operations to generate a preprocessed data set and data relationship indexes; A cross-modal fusion processing module is used to perform cross-modal interaction processing and attention interaction processing on the preprocessed data set based on the data relationship index in the fusion processing module, and generate a comprehensive feature vector; The multi-dimensional feature recognition module is used to input the comprehensive feature vector into the recognition model for recognition operations and generate a dynamic portrait of the object. The optimized path generation module is used to generate a set of personalized optimized paths based on the dynamic profile of the object. The personalized solution generation module is used to extract the demand intensity item, preference weight item, and emotional state item from the dynamic profile of the object, and combine them with the personalized optimization path set to generate a personalized solution. In the fusion processing module, based on the data relationship index, cross-modal interaction processing and attention interaction processing are performed on the preprocessed data set to generate a comprehensive feature vector, including: Based on the data relationship index, text data, image data, voice data, and numerical data of the same object are extracted from the preprocessed data set; Perform cross-modal feature alignment on the text and image data to generate text-image alignment features; Perform cross-modal feature interaction operations on the speech data and numerical data to generate speech-numerical interaction features; Extract sentiment characteristics from the text image alignment features; Extract decision-making style features from the voice-numerical interaction features; An attention-weighted operation is performed on the sentiment tendency features to generate weighted sentiment tendency features; An attention-weighted operation is performed on the decision style features to generate weighted decision style features; By integrating the weighted sentiment tendency features and weighted decision style features, a comprehensive feature vector is generated; The process of fusing the weighted sentiment characteristics and weighted decision style characteristics to generate a comprehensive feature vector includes: The weighted sentiment characteristics and the weighted decision style characteristics are subjected to dimension alignment and scale alignment. The weighted sentiment features and the weighted decision style features, after being processed by dimension alignment and scale alignment, are fused to generate a comprehensive feature vector; The comprehensive feature vector is input into the recognition model for recognition operations to generate a dynamic portrait of the object, including: The comprehensive feature vector is input into the demand identification model to perform a demand identification operation and generate a demand identification result. Extract the demand intensity term from the demand identification results; The comprehensive feature vector is input into the preference recognition model to perform preference recognition operations and generate preference recognition results. Extract preference weight terms from the preference identification results; The comprehensive feature vector is input into the emotion recognition model to perform emotion recognition operations and generate emotion recognition results. Extract the emotion state item from the emotion recognition results; Based on the demand intensity item, preference weight item, and emotional state item, a profile integration operation is performed to generate a dynamic profile of the object.
7. A computer device, characterized in that, The computer device includes a memory, a processor, and a personalization scheme generation program stored in the memory and executable on the processor, wherein when executed by the processor, the personalization scheme generation program implements the steps of the personalization scheme generation method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The storage medium stores a personalized scheme generation program, which, when executed by a processor, implements the steps of the personalized scheme generation method as described in any one of claims 1-5.
Citation Information
Patent Citations
User personalized service strategy and system based on multi-modal large model
CN117235354A
Cross-modal data fusion user psychological portrait acquisition method and system
CN118885766A
Multi-modal abstract generation and output method and device, equipment and medium
CN120216723A