Multi-modal visual arrangement recommendation method and system
By combining multimodal intent parsing and a dynamic hybrid recommendation model with genetic algorithms and graph neural networks, the problem of insufficient flexibility and user experience in existing visual orchestration recommendation technologies is solved, and an efficient and personalized visual orchestration recommendation process is achieved.
Patent Information
- Application Number
- CN202511001458.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-11
AI Technical Summary
Existing visual orchestration and recommendation technologies suffer from static rule dependence, lack of multimodal parsing, and insufficient dynamic feedback mechanisms, resulting in poor template recommendation flexibility, incomplete understanding of user intent, rigid recommendation strategies, and inflexible layout generation.
Employing multimodal intent parsing, a dynamic hybrid recommendation model, and intelligent optimization algorithms, this approach integrates text, speech, and sketch features through a cross-attention mechanism. It combines genetic algorithms and graph neural networks for layout generation, constructing a three-tiered recommendation architecture of collaborative filtering, content matching, and reinforcement learning to achieve a closed-loop process of user intent → intelligent recommendation → feedback optimization.
It significantly improves the intelligence level and user experience of visual orchestration, accurately analyzes users' complex intentions, dynamically adjusts recommendation strategies, optimizes layout generation, reduces manual adjustments, and adapts to complex data scenarios and users' personalized needs.
Smart Images

Figure CN120929066A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and data visualization, specifically a multimodal visualization arrangement and recommendation method and system. Background Technology
[0002] Existing technologies have designed various tools and methods around "visual orchestration and recommendation," including data feature-driven template matching, knowledge recommendation and confidence ranking, and user behavior and collaborative filtering recommendations. However, they generally suffer from core problems such as static rule dependence, lack of multimodal parsing, and insufficient dynamic feedback mechanisms, which limit the flexibility, accuracy, and user experience of template recommendation and layout generation. Related key technical solutions are mainly implemented through open-source tools, patented technologies, and commercial platforms.
[0003] Specifically, existing visual template recommendation and generation technologies have the following shortcomings or areas for improvement:
[0004] Disadvantage 1: Static rule dependency, insufficient dynamic adaptation capability.
[0005] Existing patented technical solutions are based on matching data field types to predefined templates and generating recommendation lists through simple sorting rules, such as the number of fields and data type priority. The generated view templates rely on a predefined rule library and cannot dynamically adapt to complex data scenarios or changes in user intent. For example, when user data simultaneously contains time, geographic, and category fields, the system cannot automatically generate a comprehensive layout with multiple linked charts; it only recommends independent charts based on a single rule.
[0006] Its technical shortcomings lie in its rigid rules, relying on manually preset rules such as "time → line chart" and "geography → map," which cannot adapt to complex data scenarios; and its lack of dynamism, failing to incorporate user feedback or real-time data change response mechanisms, making it impossible to iteratively optimize recommendation results. This results in poor recommendation template flexibility, requiring repeated manual adjustments and leading to low efficiency.
[0007] Disadvantage 2: Incomplete understanding of user intent and insufficient support for multimodal input.
[0008] Existing patented technical solutions only support a single input modality, such as text commands or data fields, without integrating multimodal information such as sketches and voice. This limits the depth of user engagement and fails to apply AI-based intent recognition models, resulting in inaccurate and incomplete intent parsing. When the system only parses text commands, it easily overlooks spatial layout requirements, making it difficult for users to mark key areas, and the recommended templates may not match the user's ideal design.
[0009] The technical limitation lies in the simplistic input rules and restricted intent recognition. This leads to a significant discrepancy between recommended results and actual user needs, resulting in high interaction costs.
[0010] Disadvantage 3: The recommendation strategy is rigid and lacks real-time feedback for optimization.
[0011] Existing patented technical solutions employ offline clustering such as KMeans or caching strategies such as LRU, which cannot dynamically update recommendation strategies based on real-time user actions, such as user modifications to components or layout adjustments. They lack feedback and learning mechanisms and cannot analyze the implicit intentions behind user behavior, such as clicks and dwell time, leading to a disconnect between layout adjustments and user needs. For example, after a user deletes a chart, the system may still recommend similar templates without lowering its priority.
[0012] The technical shortcomings lie in the cold start limitation: when there are new users or no historical data, the recommendation results are highly random and inaccurate; the lack of a feedback mechanism means that user modifications to the recommendation results are not recorded, leading to repeated recommendations of invalid templates. This results in rigid recommendation results, high reliance on human intervention, incomplete intent parsing, and an inability to adapt to dynamic user needs.
[0013] Disadvantage 4: The layout generation is rigid and lacks multi-objective optimization.
[0014] Most existing patented solutions are applicable to simple view layouts, limiting the hierarchy and number of layouts and making it difficult to adapt to the generation needs of complex view templates. Furthermore, they only meet basic alignment requirements, failing to balance multiple objectives such as information density, aesthetics, and responsive adaptation. For example, component stacking can lead to information overload, or layout errors can occur when adapting to large screens.
[0015] The technical drawback lies in the existence of multiple conflicting objectives, making it difficult to balance or achieve an optimal solution for key factors such as data matching, layout aesthetics, and information density. This results in low layout usability, requiring manual secondary optimization. Furthermore, existing patents generally focus on the process details, without disclosing the algorithm selection and implementation logic, making it difficult to verify the effects or optimize the layout. Summary of the Invention
[0016] The technical objective of this invention is to address the above-mentioned shortcomings by providing a multimodal visual orchestration and recommendation method and system that can realize a closed-loop recommendation process of "user intent → intelligent recommendation → feedback optimization", significantly improving the intelligence level of visual orchestration and user experience.
[0017] The technical solution adopted by this invention to solve its technical problem is:
[0018] A multimodal visualization orchestration and recommendation method, based on multimodal input parsing, a dynamic hybrid recommendation model, and an intelligent optimization algorithm, achieves visualized orchestration and recommendation, including the following steps:
[0019] (1) Multimodal intent parsing:
[0020] Intelligent parsing of multimodal inputs is achieved by using a base model and a fine-tuning model together. Multimodal inputs include user text, voice, sketches, etc. Text, image, and voice features are fused through a cross-attention mechanism to improve the accuracy of complex intent recognition and parse the user's complex intents.
[0021] The multimodal intent parsing step integrates text, sketch, and voice input through a cross-attention mechanism, and uses a base model and a fine-tuning model in combination. That is, it combines a pre-trained language model and domain-adaptive fine-tuning technology to achieve high-precision intent classification and trigger dynamic prompt word recommendation, especially when the confidence is insufficient, it triggers multi-turn dialogue and dynamic prompt word recommendation.
[0022] (2) Intelligent layout generation:
[0023] A genetic algorithm is used for global optimization of component space allocation, combined with a constraint solver for business rule adaptation, and a graph neural network is used to model the interaction relationships between components.
[0024] (3) Dynamic hybrid recommendation:
[0025] Construct a three-level recommendation architecture that includes collaborative filtering, content matching, and reinforcement learning; adjust the recommendation strategy weights in real time based on user feedback; and output recommendation templates based on the scoring and sorting of the above rules.
[0026] (4) Closed-loop optimization:
[0027] We continuously optimize the parameters of the intent recognition model and recommendation model using user interaction behavior data.
[0028] To address the problems in existing technologies, this method proposes a multimodal visual orchestration recommendation approach based on AI and a dynamic hybrid recommendation model. In terms of multimodal intent parsing, it solves the problem of incomplete intent understanding in existing technologies. It improves the accuracy of complex intent recognition by fusing text, image, and speech features through a cross-attention mechanism. It breaks through the limitation of single input and accurately captures complex user intents. In terms of dynamic hybrid recommendation, it solves the problem of rigid recommendation strategies. It constructs a dynamic hybrid recommendation model, combining collaborative filtering, content matching, and reinforcement learning algorithms to achieve user group and behavior learning, and dynamically adjusts recommendation weights based on real-time feedback from data features. In terms of real-time feedback-driven multi-objective layout optimization, it solves the problems of rigid layout generation and lack of dynamic adaptation in existing technologies. It uses a genetic algorithm to generate a candidate layout population with multiple objectives such as information density, aesthetics, and responsiveness. Then, it injects business rules through the Cassowary algorithm to achieve user-operation-driven recommendation strategy iteration, balancing layout aesthetics and functionality, and reducing the time-consuming manual adjustments. In terms of closed-loop interaction logic and progressive optimization, it solves the problems of insufficient retention of interaction logic and one-way recommendation in existing technologies. By tracking user adjustment patterns and converting them into reinforcement learning reward signals, the recommendation model parameters are adjusted in real time based on user adoption or rejection behavior to optimize the recommendation strategy and improve user satisfaction.
[0029] Furthermore, for step (1), the dynamic prompt word recommendation is generated based on industry knowledge base and user historical behavior data;
[0030] The domain adaptation fine-tuning is implemented using LoRA technology.
[0031] Furthermore, for step (2), the fitness function of the genetic algorithm includes three optimization objectives: information density, overlap penalty, and distribution uniformity.
[0032] The constraint solver supports dynamic injection and forced application of business rules.
[0033] Furthermore, the dynamic hybrid recommendation step adopts a three-level architecture: the first level performs coarse screening based on business rules; the second level uses Faiss for vector similarity retrieval; and the third level performs fine ranking through the LambdaMART model.
[0034] The Q-Learning algorithm is used to dynamically adjust the weights of the recommendation strategy.
[0035] The model parameters are updated based on user clicks, dwell times, and deletion behaviors.
[0036] Furthermore, the specific implementation of recommendation matching and optimization is as follows:
[0037] Candidate set generation:
[0038] Multimodal features, including text intent and data features, are extracted from user input. Based on the extracted features, a rule engine is used to filter candidate templates that match the features from a predefined template library to generate a candidate set. Then, the Faiss library is used to retrieve similar templates to achieve efficient template matching.
[0039] In the candidate set generation stage, a multi-level filtering strategy is employed to achieve efficient and accurate template retrieval: The system first extracts structured features from the user's multimodal input, including key dimensions such as parsed business intent, data features, and interaction preferences; based on the parsed features, the rule engine performs the first round of hard screening, such as automatically filtering templates containing map components for the logistics field, and prioritizing layouts with risk indicator cards for financial risk control scenarios; after completing rule filtering, the system initiates similarity matching based on dense vector retrieval, where all templates are pre-encoded into 768-dimensional feature vectors through a deep representation learning model and a Faiss efficient index is constructed; during querying, the system projects user demand features onto the same vector space and quickly finds the Top-N most relevant templates through inner product similarity calculation. This process supports millisecond-level response to retrieval requests from a template library of hundreds of millions. This rule- and vector-driven retrieval mechanism simultaneously ensures business compliance and users' deep semantic needs, achieving efficient template matching;
[0040] Multi-dimensional rating:
[0041] In the collaborative filtering dimension, the template recommendation priority is calculated based on the user's historical behavior, and the dot product CF_score of the user's latent features and the template's latent features is calculated to represent the degree of preference of the target user and similar user groups for the template.
[0042] In the content matching dimension, template metadata, including component type and layout style, is encoded into a vector. The matching degree between template metadata and user input features is calculated, and the matching degree between data features extracted from user intent and components is injected to reflect the degree of matching between the template and the current requirements (CB_score), thus achieving content matching. At the same time, business rules are preset by the administrator for mandatory matching rules in various fields. Binary judgment is used, with Rule_score being 1 if the rule is fully met, and 0 otherwise. This dimension has a fixed weight of 0.2 to ensure key business requirements are met.
[0043] By taking real-time user actions such as adopting templates and adjusting layouts as input, the Q-Learning model is used to predict the immediate reward value RL_score, capture short-term changes in user preferences, and achieve real-time feedback.
[0044] In the real-time calculation phase, multiple dimensions of scores are calculated in parallel for each candidate template. The weights are adjusted according to the latest user behavior, and the weighted sum is used to obtain the final score.
[0045] Sorting optimization:
[0046] The LambdaMART model is employed, and the NDCG metric is optimized based on gradient boosting trees. This further optimizes the template recommendation order on top of the scoring model, ensuring that not only are individual template scores reasonable, but the overall recommendation list order also aligns with the user's actual browsing preferences.
[0047] Furthermore, the multimodal intent parsing specifically includes:
[0048] (1.1) Multimodal input processing: Intelligent parsing of three input methods: text, sketch and voice;
[0049] Text parsing (NLP): Qwen-72B-Instruct was chosen as the base model for general semantic understanding. To adapt to the professional needs of different vertical domains, LoRA was selected as the fine-tuning model, which was then fine-tuned to suit the specific domain. LoRA only adjusted 0.1% to 1% of the parameters of the base model. By loading different LoRA adapters, the same base model can support parsing needs in multiple domains. The inference cost of the fine-tuned model is almost the same as that of the base model, requiring no additional resources. The text parsing process includes two key sub-tasks: extracting key fields based on the BiLSTM-CRF model for entity recognition; and using the TextCNN model to determine the user's target type for intent classification.
[0050] Sketch Analysis (CV): The system integrates two computer vision models, SAM (Segment Anything Model) and ResNet-50. The SAM model is used to accurately segment the user's hand-drawn area, and the ResNet-50 model is used to classify and identify the segmented graphic elements, mapping them to specific visualization component types.
[0051] ASR (Automatic Speech Recognition): High-precision speech-to-text conversion is achieved using the Whisper-large-v3 model. This model is fine-tuned by Whisper to adapt to business terms and uses TensorRT to accelerate inference, ensuring real-time response.
[0052] (1.2) Multimodal feature fusion:
[0053] This multimodal feature fusion technique, based on the cross-attention mechanism, unifies the heterogeneous features of text, image, and speech into a joint representation. The core of this fusion mechanism lies in establishing a dynamic association model between text features and visual / speech features, where text features serve as query vectors, and image or speech features simultaneously serve as key and value vectors. In a 768-dimensional feature space, a linear transformation layer first maps various input features to a unified representation space. Then, the dot product similarity between the query vector and the key vector is calculated, and after softmax normalization, an attention weight matrix is generated. This weight matrix dynamically represents the correlation between text and visual / speech features. Finally, the fused joint feature representation is obtained by weighted aggregation of the value vectors. This fusion method automatically focuses on the semantic association regions between multimodal features, effectively solving the problem that traditional concatenation or average pooling methods struggle to capture deep cross-modal associations. The standardized design of the feature dimensions (default 768 dimensions) ensures compatibility between different modal features and controls the numerical stability of the attention weights through scaling factors, enabling the model to adaptively balance the contribution of each modal feature.
[0054] (1.3) Intent classification and confidence assessment:
[0055] A deep neural network architecture is used to perform semantic understanding and reliability determination on the fused multimodal features; a 768-dimensional joint feature vector from the feature fusion layer is received as input, and it is mapped to multiple preset intent category spaces through a fully connected layer. The number of categories is adjusted according to the actual application scenario.
[0056] The network first performs a linear transformation on the input features to generate unnormalized classification scores, and then uses a softmax function to transform them into probability distributions, representing the likelihood of various intentions. The confidence assessment mechanism is implemented by calculating the maximum value of the probability distribution, which intuitively reflects the model's certainty about the current classification result. When the confidence is lower than a preset threshold, the system automatically triggers a prompt word recommendation process, guiding the user to supplement information through multiple rounds of dialogue; otherwise, it directly outputs a high-confidence intent judgment result. This design ensures a fast response to regular inputs and effectively handles ambiguous or insufficient user commands, significantly improving the robustness of human-computer interaction. The parameters of the fully connected layers in the module learn the discrimination boundaries of different intentions during training, while softmax normalization ensures that the output has standard probabilistic interpretability, providing a reliable basis for subsequent decision-making processes.
[0057] (1.4) Dynamic prompt word recommendation:
[0058] The system employs a multi-source data fusion and hybrid retrieval strategy to achieve precise interactive guidance. The system's prompt word database data sources include industry knowledge bases and user behavior profiles. The industry knowledge base pre-configures standardized operation templates for vertical fields such as logistics and finance, while user profiles continuously record and analyze users' historical operation preferences. When the confidence level of intent recognition is insufficient, a two-layer retrieval mechanism is triggered: First, a keyword inverted index is constructed based on the TF-IDF algorithm. By calculating the weighted word frequency similarity between the user's input text and the prompt word trigger, a semantically relevant candidate set is quickly filtered. Then, a Faiss vector database is used for deep semantic matching. Both the user query and the prompt word are encoded into standard feature dimension vectors. Nearest neighbor search is achieved through inner product similarity calculation, ensuring that the retrieval results include both literal matching and semantically related prompt content.
[0059] The system has a built-in rule engine that maintains the dialogue logic for key scenarios. When missing information or ambiguous instructions are detected, prompt words are automatically triggered for guidance. Real-time feedback optimization depends on the weight of prompt words. After a user adopts a recommended prompt word, the weight of the corresponding prompt word increases, while the weight of other prompt words in the same category decreases accordingly. The weight is updated to form an adaptive learning feedback loop, and the weight value is adjusted according to the actual scenario.
[0060] The multi-turn dialogue manager controls the flow based on a context state machine and achieves a coherent interactive experience by maintaining a dialogue history stack. It integrates rule-driven and data-driven methods to ensure deterministic guidance in key scenarios and continuously adapt to users' personalized needs, reducing interaction friction caused by ambiguous commands.
[0061] (1.5) Data source and storage logic:
[0062] The structured parameter generation process is "user input → NLP parsing → JSON parameter generation". Data storage combines temporary caching and persistent storage. Parsed parameters are temporarily stored in Redis with a limited validity period for use in subsequent layout generation. Templates containing layout parameters after user confirmation are stored in MySQL to optimize the next personalized recommendation.
[0063] The prompt word library comes from a local pre-built library and a dynamically extended library. Industry-standard prompt words are stored in a local JSON file or a lightweight database, while user-defined prompt words are stored in MySQL. Role isolation ensures access security.
[0064] The AI-powered large-scale model Qwen is used only for parsing logic. The model itself does not store business data; it only generates structural parameters based on the input text or sketch. Data binding is completed in a later stage, thus decoupling the template from the data and simplifying the user experience.
[0065] Furthermore, the intelligent layout generation specifically includes:
[0066] (2.1) Space allocation optimization:
[0067] A grid system is used to divide the screen into logical grids. Components are assigned positions according to grid units, breakpoints are added, the grid base number is switched according to the screen width, and layout recalculation is triggered when a screen size change is detected. Each layout scheme is encoded into a gene sequence using a genetic algorithm to describe the position and size of the components. A fitness function is designed to evaluate the layout quality, maximize information density and reduce crowding. By setting target weights, the relationship between multiple objectives is balanced, and a near-optimal layout scheme is quickly found in the huge layout solution space, i.e., the possible combinations of component positions and sizes.
[0068] To accelerate genetic algorithm computation, the system uses a GPU to accelerate fitness calculation and caches the fitness values of evaluated genes; the fitness function is as follows:
[0069] Fitness=α·InfoScore-β·OverlapPenalty-γ·MaqrginViolation;
[0070] InfoScore: Calculated based on component weight, which is the product of component importance weight and component area ratio. Component importance weight is defined by the user or calculated based on historical click-through rates.
[0071] OverlapPenalty: Penalty for the area of overlapping regions;
[0072] MaqrginViolation: Penalty for violating minimum margin constraints;
[0073] (2.2) Constraint Solution:
[0074] The Cassowary algorithm is used to resolve constraints and correct the layout generated by the genetic algorithm. The constraint rules include hard constraints that must be met, such as business rules, and soft constraints that should be met as much as possible, such as aesthetic rules. The algorithm implementation process is to determine whether the candidate layout generated by the genetic algorithm meets all hard constraints. If it does, the layout is retained; if not, the Cassowary solver is called to correct it and output the corrected layout. When the user drags the component, the constraints are updated in real time and the solution is recalculated. To optimize the constraint solving effect, irrelevant constraints are grouped and solved independently, and incremental updates are performed, that is, only the constraints affected by user operations are recalculated.
[0075] (2.3) Component Relationship Modeling:
[0076] This paper adopts a Graph Neural Network (GNN) to replace the manual configuration of component linkage rules, realizing intelligent modeling and prediction of relationships between visualized components. Each visualized component is abstracted as a node in a graph structure, and the node features include a 64-dimensional feature vector containing component type, data binding relationship, etc. The interaction relationship between components is represented as an edge in the graph, and the relationship type and strength are characterized by a 16-dimensional feature vector. The model architecture adopts a multi-layer message passing mechanism, with each layer containing three core processing stages: First, the edge feature processor edge_mlp generates messages. This module receives the concatenated input of source node, target node features, and edge features, and generates a 128-dimensional intermediate representation through two fully connected layers and non-linear activation, which is finally compressed into a 64-dimensional message vector; Second, the neighborhood message is executed. The aggregation process uses the `index_add` operation to accumulate the message vectors received by each node according to the target node index. Finally, the node features are updated through a gated recurrent unit (GRU), and a new node representation is generated by combining historical states and aggregated messages. By stacking the above three processing units, the model can capture high-order interaction patterns between components and automatically derive complex linkage rules such as "map selection → chart filtering". In the specific implementation, the system uses 4 component nodes and 3 interaction edges as the smallest modeling unit, and multiple such topologies can be processed in parallel in each training batch. This design can effectively replace the traditional method of manually configuring rules, reduce the rule maintenance cost, and continuously learn and optimize from user historical behavior to achieve high component linkage prediction accuracy.
[0077] (2.4) Dynamic responsive adaptation:
[0078] Adopting an adaptive grid system, the display area is divided into dynamically adjustable logical units: 12 columns by default on PCs and 6 columns on mobile devices. A breakpoint detection module continuously tracks screen size changes, achieving optimal visualization across devices by monitoring device display characteristics in real time. When a display area size change reaches a preset threshold, the system automatically triggers a four-stage processing flow: First, the current device type and resolution range are identified through media query rules; then, the optimal space allocation scheme for each visualization component is recalculated, with primary components maintaining aspect ratio scaling and secondary components being vertically stacked or collapsed according to priority; next, the Cassowary constraint solver is applied to ensure the adjusted layout conforms to business rules and aesthetic requirements; finally, the rendering engine is updated in real time. For small-screen devices such as mobile devices, the system intelligently implements information hierarchy compression strategies, automatically folding detailed data and highlighting KPI indicators to ensure the readability of core information.
[0079] To support rapid prototyping during the layout design phase, an integrated intelligent mock data generator automatically synthesizes simulated datasets that conform to business semantics based on component semantic types. For example, the time series component generates normally distributed random values for 12 months, and the geographic information component simulates sales data for three major regions. All generated data strictly adheres to the data patterns of real business operations, such as using standard datetime format for timestamps and retaining two decimal places for numeric fields. The data binding engine automatically matches the mock data structure based on the component type; the line chart component binds to the x / y coordinates of the time series, and the bar chart component associates geographic regions with sales values, ensuring the realism and consistency of the demonstration. This dynamic data simulation mechanism allows users to quickly verify the layout adaptation effect on different devices without access to real data, significantly improving prototype iteration efficiency.
[0080] This invention also claims a multimodal visual orchestration and recommendation system, comprising:
[0081] A multimodal input parsing module is used to realize intent recognition and feature fusion of text, sketches, and speech;
[0082] The hybrid layout optimization module integrates a genetic algorithm engine, a constraint solver, and a graph neural network processor.
[0083] The dynamic recommendation engine module includes rule filters, vector retrieval systems, and ranking models.
[0084] The feedback learning module is used to collect user behavior data and update model parameters;
[0085] The system achieves multimodal visual orchestration and recommendation through the methods described above.
[0086] The present invention also claims a multimodal visual orchestration recommendation apparatus, comprising: at least one memory and at least one processor;
[0087] The at least one memory is used to store a machine-readable program;
[0088] The at least one processor is used to call the machine-readable program to implement the above method.
[0089] The present invention also claims a computer-readable medium storing computer instructions that, when executed by a processor, enable the implementation of the above-described method.
[0090] Compared with existing technologies, the multimodal visual orchestration and recommendation method and system of the present invention have the following advantages:
[0091] 1. In terms of multimodal intent parsing, this invention integrates multiple input methods such as text, voice, and sketches, and utilizes a cross-attention mechanism to fuse multimodal features. This overcomes the limitations of single input modes in existing technologies and solves the problem of incomplete intent understanding, achieving accurate parsing of complex user intents. Furthermore, it combines industry-standard prompts with user-personalized prompts, balancing the standardization and customization needs of users.
[0092] 2. Regarding intelligent layout generation, this invention utilizes a genetic algorithm combined with constraint solving, featuring global layout optimization, multi-objective support, and strong dynamic adaptability. It incorporates GNN component relationship modeling, automating relationship discovery and capturing high-order dependencies across components, replacing manual configuration of component linkage rules, significantly lowering the barrier to entry. It also supports automatic optimization of the relationship prediction model based on user historical behavior. The use of this invention will significantly improve the quality and practicality of layout generation, especially reducing the workload of manual adjustments when dealing with complex data scenarios, and is particularly suitable for multi-component, multi-constraint scenarios.
[0093] 3. Regarding the dynamic hybrid recommendation model, this invention employs a three-stage dynamic fusion strategy of collaborative filtering, content matching, and reinforcement learning. This strategy optimizes recommendation weights based on real-time user actions, achieving a closed-loop decision-making process for the recommendation system. The multi-strategy fusion avoids the limitations of a single recommendation strategy, covering more scenarios; reinforcement learning dynamically adjusts weights to adapt to personalized user needs, and allows administrators to flexibly adjust rule weights to suit different industry requirements.
[0094] 4. Regarding real-time feedback loop, this invention records user operation data and converts it into reinforcement learning reward signals to achieve online updates of the recommendation strategy. This not only solves the problem of insufficient retention of interaction logic in existing technologies, but also significantly improves the efficiency of iterative optimization of recommendation results, reduces the occurrence of repeated recommendations of invalid templates, and significantly improves user experience and the system's intelligence level. Attached Figure Description
[0095] Figure 1 This is a diagram illustrating the technical architecture of the multimodal visual orchestration and recommendation algorithm provided in this embodiment of the invention;
[0096] Figure 2 This is a flowchart illustrating the intent recognition and prompt word triggering implementation process provided in an embodiment of the present invention;
[0097] Figure 3 This is a flowchart illustrating the layout generation implementation process provided in an embodiment of the present invention;
[0098] Figure 4 This is a flowchart illustrating the recommended matching optimization implementation process provided in an embodiment of the present invention. Detailed Implementation
[0099] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0100] A multimodal visualization orchestration and recommendation method, based on multimodal input parsing, a dynamic hybrid recommendation model, and an intelligent optimization algorithm, achieves visualized orchestration and recommendation, including the following steps:
[0101] (1) Multimodal intent parsing:
[0102] Intelligent parsing of multimodal inputs is achieved by using a base model and a fine-tuning model together. Multimodal inputs include user text, voice, sketches, etc. Text, image, and voice features are fused through a cross-attention mechanism to improve the accuracy of complex intent recognition and parse the user's complex intents.
[0103] The multimodal intent parsing step integrates text, sketch, and voice input through a cross-attention mechanism, and jointly uses a base model and a fine-tuning model. That is, it combines a pre-trained language model and domain-adaptive fine-tuning techniques to achieve high-precision intent classification and trigger dynamic prompt word recommendations, especially when the confidence level is insufficient, it triggers multi-turn dialogue and dynamic prompt word recommendations.
[0104] (2) Intelligent layout generation:
[0105] A genetic algorithm is used for global optimization of component space allocation, a constraint solver is used for business rule adaptation, and a graph neural network is used to model the interaction relationships between components.
[0106] (3) Dynamic hybrid recommendation:
[0107] A three-level recommendation architecture is constructed, which includes collaborative filtering, content matching, and reinforcement learning. The recommendation strategy weights are adjusted in real time based on user feedback. Recommendation templates are output based on the scoring and ranking of the above rules.
[0108] (4) Closed-loop optimization:
[0109] We continuously optimize the parameters of the intent recognition model and recommendation model using user interaction behavior data.
[0110] In step (1), the dynamic prompt word recommendation is generated based on industry knowledge base and user historical behavior data;
[0111] The domain adaptation fine-tuning is implemented using LoRA technology.
[0112] For step (2), the fitness function of the genetic algorithm includes three optimization objectives: information density, overlap penalty, and distribution uniformity.
[0113] The constraint solver supports dynamic injection and forced application of business rules.
[0114] For step (3), the dynamic hybrid recommendation step adopts a three-level architecture: the first level performs coarse screening based on business rules; the second level uses Faiss for vector similarity retrieval; and the third level performs fine ranking through the LambdaMART model.
[0115] For step (4), the Q-Learning algorithm is used to dynamically adjust the weights of the recommendation strategy;
[0116] The model parameters are updated based on user clicks, dwell times, and deletion behaviors.
[0117] This method relies on the component library and business rule engine of the visualization system, and mainly innovates the design of multimodal intent recognition, hybrid layout optimization, and dynamic recommendation strategy. By integrating multimodal feature understanding with cross-attention mechanism, joint spatial optimization of genetic algorithm and graph neural network, and hybrid recommendation mechanism of collaborative filtering and reinforcement learning, a closed-loop recommendation process of "user intent → intelligent recommendation → feedback optimization" is realized, which significantly improves the intelligence level of visualization orchestration and user experience.
[0118] like Figure 1 As shown, the technical architecture of this visualization orchestration and recommendation method revolves around three core modules: intent recognition and prompt word triggering, layout generation algorithm, and recommendation matching optimization. This architecture achieves end-to-end generation from vague requirements to precise visualization solutions through a technical chain of "multimodal intent understanding → hybrid layout optimization → dynamic recommendation matching," exhibiting stronger adaptability and domain scalability compared to traditional solutions.
[0119] The specific implementation process of this method is as follows:
[0120] I. Intent recognition and prompt word triggering module.
[0121] 1. Multimodal input processing.
[0122] This method employs advanced deep learning technology in the multimodal input processing stage to achieve intelligent parsing of three input methods: text, sketch, and speech.
[0123] Text parsing (NLP): The Qwen-72B-Instruct model was chosen as the base model for general semantic understanding. To adapt to the specific needs of different vertical domains, LoRA was selected as the fine-tuning model, tailored to each domain. LoRA only adjusts 0.1% to 1% of the base model's parameters. By loading different LoRA adapters, the same base model can support parsing requirements across multiple domains. The inference cost of the fine-tuned model is almost identical to that of the base model, requiring no additional resources. The text parsing process includes two key sub-tasks: extracting key fields for entity recognition based on the BiLSTM-CRF model, and using the TextCNN model to determine the user's target type for intent classification.
[0124] Sketching Analysis (CV): The system integrates two computer vision models: SAM (Segment Anything Model) and ResNet-50. SAM is responsible for accurately segmenting the user's hand-drawn area, while ResNet-50 classifies and identifies the segmented graphic elements, mapping them to specific visualization component types.
[0125] ASR (Automatic Speech Recognition): The Whisper-large-v3 model is selected to achieve high-precision speech-to-text conversion. This model is fine-tuned by Whisper to adapt to business terms and uses TensorRT to accelerate inference, ensuring real-time response.
[0126] 2. Multimodal feature fusion.
[0127] This multimodal feature fusion technique, based on cross-attention, unifies the heterogeneous features of text, image, and speech into a joint representation. The core of this fusion mechanism lies in establishing a dynamic association model between text features and visual / speech features, where text features serve as query vectors, and image or speech features simultaneously serve as key and value vectors. In a 768-dimensional feature space, the system first maps various input features to a unified representation space through a linear transformation layer. Then, it calculates the dot product similarity between the query vector and the key vector, and generates an attention weight matrix after softmax normalization. This weight matrix dynamically represents the correlation between text and visual / speech features. Finally, the fused joint feature representation is obtained by weighted aggregation of the value vectors. This fusion method automatically focuses on the semantic association regions between multimodal features, effectively solving the problem that traditional concatenation or average pooling methods struggle to capture deep cross-modal associations. The standardized design of the feature dimensions (default 768 dimensions) ensures compatibility between different modal features and allows for scaling factors to be used. The numerical stability of the attention weights was controlled, enabling the model to adaptively balance the contribution of each modality feature.
[0128]
[0129] Q: Text feature vector;
[0130] K,V: Image / Speech feature vectors;
[0131] d: Feature dimension (default 768);
[0132] 3. Intent classification and confidence assessment.
[0133] A deep neural network architecture is employed to perform semantic understanding and reliability assessment on the fused multimodal features. This module receives a 768-dimensional joint feature vector from the feature fusion layer as input, which is then mapped to four preset intent category spaces through a fully connected layer. The number of categories can be adjusted according to the actual application scenario. The network first performs a linear transformation on the input features to generate unnormalized classification scores, and then uses a softmax function to transform them into probability distributions, representing the likelihood of each intent. The confidence assessment mechanism is implemented by calculating the maximum value of the probability distribution, which intuitively reflects the model's certainty about the current classification result. When the confidence level is lower than a preset threshold, the system automatically triggers a prompt word recommendation process, guiding the user to supplement information through multiple rounds of dialogue; otherwise, it directly outputs a high-confidence intent judgment result. This design ensures rapid response to regular inputs while effectively handling ambiguous or insufficiently informed user commands, significantly improving the robustness of human-computer interaction. The parameters of the fully connected layer in the module learn the discrimination boundaries of different intents during training, while softmax normalization ensures that the output has standard probabilistic interpretability, providing a reliable basis for subsequent decision-making processes.
[0134] 4. Dynamic prompt word recommendations.
[0135] A multi-source data fusion and hybrid retrieval strategy is employed to achieve precise interactive guidance. The system's prompt word database data sources include industry knowledge bases and user behavior profiles. The industry knowledge base pre-sets standardized operation templates for vertical fields such as logistics and finance, while user profiles continuously record and analyze users' historical operation preferences. When the confidence level of intent recognition is insufficient, a two-layer retrieval mechanism is triggered: First, a keyword inverted index is constructed based on the TF-IDF algorithm. By calculating the weighted word frequency similarity between the user's input text and the prompt word trigger, a semantically related candidate set is quickly filtered. Then, the Faiss vector database is used for deep semantic matching. Both the user query and the prompt word are encoded into standard feature dimension vectors. The nearest neighbor search is achieved through inner product similarity calculation, ensuring that the retrieval results include both literal matching and semantically related prompt content.
[0136] The system's built-in rule engine maintains the dialogue logic for key scenarios, automatically triggering prompts when missing information or ambiguous instructions are detected. Real-time feedback optimization relies on prompt weights. When a user adopts a recommended prompt, the corresponding prompt's weight increases, while the weights of other prompts in the same category decrease accordingly. This weight update forms an adaptive learning feedback loop, adjusting weight values according to the actual scenario.
[0137] The multi-turn dialogue manager uses a context state machine for flow control and maintains a dialogue history stack to achieve a coherent interactive experience. By integrating rule-driven and data-driven approaches, it can ensure deterministic guidance in key scenarios while continuously adapting to personalized user needs, reducing interaction friction caused by ambiguous commands.
[0138] 5. Data source and storage logic.
[0139] The structured parameter generation process is "user input → NLP parsing → JSON parameter generation". Data storage combines temporary caching and persistent storage. Parsed parameters are temporarily stored in Redis with a limited validity period for use in subsequent layout generation. Templates containing layout parameters after user confirmation are stored in MySQL to optimize the next personalized recommendation.
[0140] The prompt word library comes from a local pre-built library and a dynamically extended library. Industry-standard prompt words are stored in a local JSON file or a lightweight database, while user-defined prompt words are stored in MySQL. Role isolation ensures access security.
[0141] The AI-powered large-scale model Qwen is used only for parsing logic. The model itself does not store business data; it only generates structural parameters based on the input text or sketch. Data binding is completed in a later stage, thus decoupling the template from the data and simplifying the user experience.
[0142] The implementation process of intent recognition and prompt triggering technology is as follows: Figure 2 As shown.
[0143] II. Layout generation module.
[0144] The core objective of layout generation algorithms is to transform user needs after intent recognition into the spatial allocation of visual components and the relationships between components, so as to maximize information density, visual appeal, responsive adaptation, and compliance with business rule constraints.
[0145] 1. Space allocation optimization.
[0146] A grid system is used to divide the screen into logical grids. Components are assigned positions according to grid units, breakpoints are added, the grid cardinality is switched based on the screen width, and layout recalculation is triggered when a screen size change is detected. A genetic algorithm encodes each layout scheme as a gene sequence, describing the position and size of the components. A fitness function is designed to evaluate layout quality, maximizing information density and reducing crowding. By setting objective weights, multi-objective relationships are balanced, and a near-optimal layout scheme is quickly found from the vast layout solution space—the possible combinations of component positions and sizes. To accelerate the genetic algorithm computation, the system uses a GPU to accelerate fitness calculation and caches the fitness values of evaluated genes.
[0147] Fitness function:
[0148] Fitness=α·InfoScore-β·OverlapPenalty-γ·MaqrginViolation;
[0149] InfoScore: Calculated based on component weight, which is the product of component importance weight and component area ratio. Component importance weight is defined by the user or calculated based on historical click-through rates.
[0150] OverlapPenalty: Penalty for overlapping area.
[0151] MaqrginViolation: Penalty for violating the minimum margin constraint.
[0152] 2. Constraint solution.
[0153] The Cassowary algorithm is used to resolve constraints and correct the layout generated by the genetic algorithm. Constraint rules include hard constraints that must be met, such as business rules, and soft constraints that should be met as much as possible, such as aesthetic rules. The algorithm implementation process involves determining whether the candidate layout generated by the genetic algorithm satisfies all hard constraints. If it does, the layout is retained; otherwise, the Cassowary solver is called to correct it, and the corrected layout is output. When the user drags components, the constraints are updated in real time and the solution is recalculated. To optimize the constraint solving effect, irrelevant constraints are grouped and solved independently, with incremental updates, meaning only constraints affected by user actions are recalculated.
[0154] 3. Component relationship modeling.
[0155] This paper employs a Graph Neural Network (GNN) to replace manual configuration of component linkage rules, enabling intelligent modeling and prediction of relationships between visualized components. Each visualized component is abstracted as a node in a graph structure, with node features including a 64-dimensional feature vector representing component type, data binding relationships, etc. Interactions between components are represented as edges in the graph, characterized by 16-dimensional feature vectors depicting relationship type and strength. The model architecture uses a multi-layer message passing mechanism, with each layer containing three core processing stages: First, the edge feature processor `edge_mlp` generates messages. This module receives concatenated inputs of source node, target node features, and edge features, processes them through two fully connected layers and non-linear activation to generate a 128-dimensional intermediate representation, and finally compresses it into a 64-dimensional message vector. Second, neighborhood message aggregation is performed, using the `index_add` operation to accumulate the message vectors received by each node according to the target node index. Finally, a gated recurrent unit (GRU) updates node features, combining historical states and aggregated messages to generate new node representations. By stacking these three processing units, the model can capture high-order interaction patterns between components and automatically derive complex linkage rules such as "map selection → chart filtering". In its implementation, the system uses four component nodes and three interaction edges as the smallest modeling unit, and multiple such topologies can be processed in parallel per training batch. This design effectively replaces the traditional method of manually configuring rules, reducing rule maintenance costs, and can continuously learn and optimize from users' historical behavior to achieve high component linkage prediction accuracy.
[0156] 4. Dynamic responsive adaptation.
[0157] Employing an adaptive grid system, the display area is divided into dynamically adjustable logical units: 12 columns by default on PCs and 6 columns on mobile devices. A breakpoint detection module continuously tracks screen size changes, achieving optimal visualization across devices by monitoring device display characteristics in real time. When a display area size change reaches a preset threshold, the system automatically triggers a four-stage processing flow: First, it identifies the current device type and resolution range using media query rules; then, it recalculates the optimal space allocation scheme for each visualization component, where primary components maintain aspect ratio scaling, and secondary components are vertically stacked or collapsed according to priority; next, it applies the Cassowary constraint solver to ensure the adjusted layout conforms to business rules and aesthetic requirements; finally, it completes real-time updates of the rendering engine. For small-screen devices such as mobile devices, the system intelligently implements information hierarchy compression strategies, automatically collapsing detailed data and highlighting KPI indicators to ensure the readability of core information.
[0158] To support rapid prototyping during the layout design phase, an integrated intelligent mock data generator automatically synthesizes simulated datasets that conform to business semantics based on component semantic types. For example, the time series component generates normally distributed random values for 12 months, and the geographic information component simulates sales data for three major regions. All generated data strictly adheres to the data patterns of real business operations, such as using standard datetime format for timestamps and retaining two decimal places for numeric fields. The data binding engine automatically matches the mock data structure based on the component type; the line chart component binds to the x / y coordinates of the time series, and the bar chart component associates geographic regions with sales values, ensuring the realism and consistency of the demonstration. This dynamic data simulation mechanism allows users to quickly verify the layout adaptation effect on different devices without access to real data, significantly improving prototype iteration efficiency.
[0159] The layout generation technology implementation process is as follows: Figure 3 As shown.
[0160] III. Recommended Matching Optimization Module.
[0161] The core objective of the recommendation matching optimization module is to filter and sort the most suitable solutions from a massive pool of candidate layout templates based on user intent, historical behavior, and business rules, thereby improving user orchestration efficiency. Its core functions include generating a template candidate set, scoring templates across multiple dimensions, dynamically adjusting weights based on user feedback, and optimizing the recommendation strategy.
[0162] 1. Candidate set generation.
[0163] Multimodal features, including text intent and data features, are extracted from user input. Based on the extracted features, a rule engine filters candidate templates that match the features from a predefined template library to generate a candidate set. Then, the Faiss library is used to retrieve similar templates to achieve efficient template matching.
[0164] This method employs a multi-level filtering strategy in the candidate set generation stage to achieve efficient and accurate template retrieval. The system first extracts structured features from the user's multimodal input, including key dimensions such as parsed business intent, data features, and interaction preferences. Based on these features, the rule engine performs an initial hard screening, automatically filtering templates containing map components for the logistics field, and prioritizing layouts with risk indicator cards for financial risk control scenarios. After rule filtering, the system initiates similarity matching based on dense vector retrieval. All templates are pre-encoded into 768-dimensional feature vectors using a deep representation learning model, and a Faiss efficient index is constructed. During queries, the system projects user demand features onto the same vector space and quickly finds the Top-N most relevant templates through inner product similarity calculation. This process supports millisecond-level response to retrieval requests from a library of hundreds of millions of templates. This rule- and vector-driven retrieval mechanism simultaneously ensures business compliance and meets users' deep semantic needs, achieving efficient template matching.
[0165] 2. Multi-dimensional scoring model.
[0166] First, in the collaborative filtering dimension, template recommendation priority is calculated based on user historical behavior. The dot product (CF_score) of the user's latent features and the template's latent features is calculated, representing the degree of preference of the target user and similar user groups for that template. Next, in the content matching dimension, template metadata such as component type and layout style are encoded into vectors. The matching degree between template metadata and user input features is calculated, and the matching degree between data features extracted from user intent and components is injected, reflecting the degree of matching between the template and the current needs (CB_score), thus achieving content matching. Simultaneously, business rules are preset by the administrator for mandatory matching rules in various domains, using binary judgment. A Rule_score of 1 is awarded if the rule is fully met, otherwise it is 0. This dimension has a fixed weight of 0.2 to ensure key business requirements are met. Finally, real-time user actions such as template adoption and layout adjustments are used as input. The Q-Learning model predicts the immediate reward value (RL_score) to capture short-term user preference changes and achieve real-time feedback. The weight coefficients and adjustment rules are shown in Table 1 below.
[0167] Table 1 - Weighting Coefficients and Adjustment Rules
[0168]
[0169]
[0170] In the real-time calculation phase, scores for each candidate template across four dimensions are calculated in parallel. The weights are adjusted based on the latest user behavior, and the weighted sum is used to obtain the final score.
[0171] 3. Sorting optimization.
[0172] The LambdaMART model is employed, and the NDCG metric is optimized based on gradient boosting trees. This further optimizes the template recommendation order on top of the scoring model, ensuring that not only are individual template scores reasonable, but the overall recommendation list order also aligns with the user's actual browsing preferences.
[0173] The implementation process of the recommended matching optimization technology is as follows: Figure 4 As shown.
[0174] This invention also provides a multimodal visual orchestration and recommendation system, comprising:
[0175] A multimodal input parsing module is used to realize intent recognition and feature fusion of text, sketches, and speech;
[0176] The hybrid layout optimization module integrates a genetic algorithm engine, a constraint solver, and a graph neural network processor.
[0177] The dynamic recommendation engine module includes rule filters, vector retrieval systems, and ranking models.
[0178] The feedback learning module is used to collect user behavior data and update model parameters.
[0179] The hybrid layout optimization module includes a visual component gene encoder, which converts component attributes into numerical representations that can be processed by the optimization algorithm; and a layout renderer, which converts the optimization results into a visual interface.
[0180] The dynamic recommendation engine module includes a permission verification engine: which verifies role permissions in real time when a user interacts with a template, and prohibits access if unauthorized; and an operation log module: which records the executor and timestamp of template creation, referencing, and deletion operations.
[0181] The system implements multimodal visual orchestration and recommendation through the multimodal visual orchestration and recommendation method described in the above embodiments.
[0182] The multimodal input parsing module achieves intelligent parsing of multimodal inputs through the combined use of a base model and a fine-tuning model. Multimodal inputs include user text, speech, and sketches. A cross-attention mechanism is used to fuse text, image, and speech features, improving the accuracy of complex intent recognition and parsing complex user intents. The multimodal intent parsing steps fuse text, sketches, and speech inputs through a cross-attention mechanism, jointly using a base model and a fine-tuning model. This combines a pre-trained language model with domain-adaptive fine-tuning techniques to achieve high-precision intent classification and trigger dynamic prompt word recommendations, especially when confidence levels are insufficient, triggering multi-turn dialogues and dynamic prompt word recommendations.
[0183] The hybrid layout optimization module uses a genetic algorithm for global optimization of component space allocation, combines a constraint solver for business rule adaptation, and utilizes a graph neural network to model the interaction relationships between components.
[0184] The dynamic recommendation engine module constructs a three-level recommendation architecture that includes collaborative filtering, content matching, and reinforcement learning. It adjusts the recommendation strategy weights in real time based on user feedback and outputs recommendation templates based on the scoring and sorting of the above rules.
[0185] The feedback learning module continuously optimizes the parameters of the intent recognition model and recommendation model using user interaction behavior data.
[0186] The multimodal input parsing module includes:
[0187] (1) Multimodal input processing.
[0188] This method employs advanced deep learning technology in the multimodal input processing stage to achieve intelligent parsing of three input methods: text, sketch, and speech.
[0189] Text parsing (NLP): The Qwen-72B-Instruct model was chosen as the base model for general semantic understanding. To adapt to the specific needs of different vertical domains, LoRA was selected as the fine-tuning model, tailored to each domain. LoRA only adjusts 0.1% to 1% of the base model's parameters. By loading different LoRA adapters, the same base model can support parsing requirements across multiple domains. The inference cost of the fine-tuned model is almost identical to that of the base model, requiring no additional resources. The text parsing process includes two key sub-tasks: extracting key fields for entity recognition based on the BiLSTM-CRF model, and using the TextCNN model to determine the user's target type for intent classification.
[0190] Sketching Analysis (CV): The system integrates two computer vision models: SAM (Segment Anything Model) and ResNet-50. SAM is responsible for accurately segmenting the user's hand-drawn area, while ResNet-50 classifies and identifies the segmented graphic elements, mapping them to specific visualization component types.
[0191] ASR (Automatic Speech Recognition): The Whisper-large-v3 model is selected to achieve high-precision speech-to-text conversion. This model is fine-tuned by Whisper to adapt to business terms and uses TensorRT to accelerate inference, ensuring real-time response.
[0192] (2) Multimodal feature fusion.
[0193] This multimodal feature fusion technique, based on the cross-attention mechanism, unifies the heterogeneous features of text, image, and speech into a joint representation. The core of this fusion mechanism lies in establishing a dynamic association model between text features and visual / speech features, where text features serve as query vectors, and image or speech features simultaneously serve as key and value vectors. In a 768-dimensional feature space, the system first maps various input features to a unified representation space through a linear transformation layer. Then, it calculates the dot product similarity between the query vector and the key vector, and generates an attention weight matrix after softmax normalization. This weight matrix dynamically represents the correlation between text and visual / speech features. Finally, the fused joint feature representation is obtained by weighted aggregation of the value vectors. This fusion method automatically focuses on the semantic association regions between multimodal features, effectively solving the problem that traditional concatenation or average pooling methods struggle to capture deep cross-modal associations. The standardized feature dimension design (default 768 dimensions) ensures compatibility between different modal features and controls the numerical stability of the attention weights through scaling factors, enabling the model to adaptively balance the contribution of each modality feature.
[0194] (3) Intent classification and confidence assessment.
[0195] A deep neural network architecture is employed to perform semantic understanding and reliability assessment on the fused multimodal features. This module receives a 768-dimensional joint feature vector from the feature fusion layer as input, which is then mapped to four preset intent category spaces through a fully connected layer. The number of categories can be adjusted according to the actual application scenario. The network first performs a linear transformation on the input features to generate unnormalized classification scores, and then uses a softmax function to transform them into probability distributions, representing the likelihood of each intent. The confidence assessment mechanism is implemented by calculating the maximum value of the probability distribution, which intuitively reflects the model's certainty about the current classification result. When the confidence level is lower than a preset threshold, the system automatically triggers a prompt word recommendation process, guiding the user to supplement information through multiple rounds of dialogue; otherwise, it directly outputs a high-confidence intent judgment result. This design ensures rapid response to regular inputs while effectively handling ambiguous or insufficiently informed user commands, significantly improving the robustness of human-computer interaction. The parameters of the fully connected layer in the module learn the discrimination boundaries of different intents during training, while softmax normalization ensures that the output has standard probabilistic interpretability, providing a reliable basis for subsequent decision-making processes.
[0196] (4) Dynamic prompt word recommendation.
[0197] A multi-source data fusion and hybrid retrieval strategy is employed to achieve precise interactive guidance. The system's prompt word database data sources include industry knowledge bases and user behavior profiles. The industry knowledge base pre-sets standardized operation templates for vertical fields such as logistics and finance, while user profiles continuously record and analyze users' historical operation preferences. When the confidence level of intent recognition is insufficient, a two-layer retrieval mechanism is triggered: First, a keyword inverted index is constructed based on the TF-IDF algorithm. By calculating the weighted word frequency similarity between the user's input text and the prompt word trigger, a semantically related candidate set is quickly filtered. Then, the Faiss vector database is used for deep semantic matching. Both the user query and the prompt word are encoded into standard feature dimension vectors. The nearest neighbor search is achieved through inner product similarity calculation, ensuring that the retrieval results include both literal matching and semantically related prompt content.
[0198] The system's built-in rule engine maintains the dialogue logic for key scenarios, automatically triggering prompts when missing information or ambiguous instructions are detected. Real-time feedback optimization relies on prompt weights. When a user adopts a recommended prompt, the corresponding prompt's weight increases, while the weights of other prompts in the same category decrease accordingly. This weight update forms an adaptive learning feedback loop, adjusting weight values according to the actual scenario.
[0199] The multi-turn dialogue manager uses a context state machine for flow control and maintains a dialogue history stack to achieve a coherent interactive experience. By integrating rule-driven and data-driven approaches, it can ensure deterministic guidance in key scenarios while continuously adapting to personalized user needs, reducing interaction friction caused by ambiguous commands.
[0200] (5) Data source and storage logic.
[0201] The structured parameter generation process is "user input → NLP parsing → JSON parameter generation". Data storage combines temporary caching and persistent storage. Parsed parameters are temporarily stored in Redis with a limited validity period for use in subsequent layout generation. Templates containing layout parameters after user confirmation are stored in MySQL to optimize the next personalized recommendation.
[0202] The prompt word library comes from a local pre-built library and a dynamically extended library. Industry-standard prompt words are stored in a local JSON file or a lightweight database, while user-defined prompt words are stored in MySQL. Role isolation ensures access security.
[0203] The AI-powered large-scale model Qwen is used only for parsing logic. The model itself does not store business data; it only generates structural parameters based on the input text or sketch. Data binding is completed in a later stage, thus decoupling the template from the data and simplifying the user experience.
[0204] The hybrid layout optimization module includes:
[0205] (1) Space allocation optimization.
[0206] A grid system is used to divide the screen into logical grids. Components are assigned positions according to grid units, breakpoints are added, the grid cardinality is switched based on the screen width, and layout recalculation is triggered when a screen size change is detected. A genetic algorithm encodes each layout scheme as a gene sequence, describing the position and size of the components. A fitness function is designed to evaluate layout quality, maximizing information density and reducing crowding. By setting objective weights, multi-objective relationships are balanced, and a near-optimal layout scheme is quickly found from the vast layout solution space—the possible combinations of component positions and sizes. To accelerate the genetic algorithm computation, the system uses a GPU to accelerate fitness calculation and caches the fitness values of evaluated genes.
[0207] (2) Constraint solution.
[0208] The Cassowary algorithm is used to resolve constraints and correct the layout generated by the genetic algorithm. Constraint rules include hard constraints that must be met, such as business rules, and soft constraints that should be met as much as possible, such as aesthetic rules. The algorithm implementation process involves determining whether the candidate layout generated by the genetic algorithm satisfies all hard constraints. If it does, the layout is retained; otherwise, the Cassowary solver is called to correct it, and the corrected layout is output. When the user drags components, the constraints are updated in real time and the solution is recalculated. To optimize the constraint solving effect, irrelevant constraints are grouped and solved independently, with incremental updates, meaning only constraints affected by user actions are recalculated.
[0209] (3) Component relationship modeling.
[0210] This paper employs a Graph Neural Network (GNN) to replace manual configuration of component linkage rules, enabling intelligent modeling and prediction of relationships between visualized components. Each visualized component is abstracted as a node in a graph structure, with node features including a 64-dimensional feature vector representing component type, data binding relationships, etc. Interactions between components are represented as edges in the graph, characterized by 16-dimensional feature vectors depicting relationship type and strength. The model architecture uses a multi-layer message passing mechanism, with each layer containing three core processing stages: First, the edge feature processor `edge_mlp` generates messages. This module receives concatenated inputs of source node, target node features, and edge features, processes them through two fully connected layers and non-linear activation to generate a 128-dimensional intermediate representation, and finally compresses it into a 64-dimensional message vector. Second, neighborhood message aggregation is performed, using the `index_add` operation to accumulate the message vectors received by each node according to the target node index. Finally, a gated recurrent unit (GRU) updates node features, combining historical states and aggregated messages to generate new node representations. By stacking these three processing units, the model can capture high-order interaction patterns between components and automatically derive complex linkage rules such as "map selection → chart filtering". In its implementation, the system uses four component nodes and three interaction edges as the smallest modeling unit, and multiple such topologies can be processed in parallel per training batch. This design effectively replaces the traditional method of manually configuring rules, reducing rule maintenance costs, and can continuously learn and optimize from users' historical behavior to achieve high component linkage prediction accuracy.
[0211] (4) Dynamic responsive adaptation.
[0212] Employing an adaptive grid system, the display area is divided into dynamically adjustable logical units: 12 columns by default on PCs and 6 columns on mobile devices. A breakpoint detection module continuously tracks screen size changes, achieving optimal visualization across devices by monitoring device display characteristics in real time. When a display area size change reaches a preset threshold, the system automatically triggers a four-stage processing flow: First, it identifies the current device type and resolution range using media query rules; then, it recalculates the optimal space allocation scheme for each visualization component, where primary components maintain aspect ratio scaling, and secondary components are vertically stacked or collapsed according to priority; next, it applies the Cassowary constraint solver to ensure the adjusted layout conforms to business rules and aesthetic requirements; finally, it completes real-time updates of the rendering engine. For small-screen devices such as mobile devices, the system intelligently implements information hierarchy compression strategies, automatically collapsing detailed data and highlighting KPI indicators to ensure the readability of core information.
[0213] To support rapid prototyping during the layout design phase, an intelligent Mock data generator is integrated to automatically synthesize simulated datasets that conform to business semantics based on the semantic type of the components.
[0214] The dynamic recommendation engine module includes:
[0215] (1) Candidate set generation.
[0216] Multimodal features, including text intent and data features, are extracted from user input. Based on the extracted features, a rule engine filters candidate templates that match the features from a predefined template library to generate a candidate set. Then, the Faiss library is used to retrieve similar templates to achieve efficient template matching.
[0217] In the candidate set generation stage, a multi-level filtering strategy is employed to achieve efficient and accurate template retrieval. The system first extracts structured features from the user's multimodal input, including key dimensions such as parsed business intent, data features, and interaction preferences. Based on these features, the rule engine performs an initial hard screening, automatically filtering templates containing map components for the logistics field, and prioritizing layouts with risk indicator cards for financial risk control scenarios. After rule filtering, the system initiates similarity matching based on dense vector retrieval. All templates are pre-encoded into 768-dimensional feature vectors using a deep representation learning model, and a Faiss efficient index is constructed. During queries, the system projects user demand features onto the same vector space and quickly finds the Top-N most relevant templates through inner product similarity calculation. This process supports millisecond-level response to retrieval requests from a library of hundreds of millions of templates. This rule- and vector-driven retrieval mechanism simultaneously ensures business compliance and meets users' deep semantic needs, achieving efficient template matching.
[0218] (2) Multidimensional scoring model.
[0219] First, in the collaborative filtering dimension, template recommendation priority is calculated based on user historical behavior. The dot product (CF_score) of the user's latent features and the template's latent features is calculated, representing the target user's and similar user groups' preference for that template. Next, in the content matching dimension, template metadata such as component type and layout style are encoded into vectors. The matching degree between template metadata and user input features is calculated, and the matching degree between data features extracted from user intent and components is injected, reflecting the template's matching degree with current needs (CB_score), thus achieving content matching. Simultaneously, business rules are preset by administrators with mandatory matching rules for each domain. Binary judgment is used; a Rule_score of 1 is awarded for complete compliance with the rule, and 0 otherwise. This dimension has a fixed weight of 0.2 to ensure key business requirements are met. Finally, real-time user actions such as template adoption and layout adjustments are used as input. The Q-Learning model predicts the immediate reward value (RL_score) to capture short-term user preference changes and achieve real-time feedback. In the real-time calculation phase, scores across four dimensions are calculated in parallel for each candidate template. Weights are adjusted based on the latest user behavior, and a weighted sum is obtained to obtain the final score.
[0220] The feedback learning module employs the LambdaMART model, which optimizes the NDCG metric based on gradient boosting trees. This further optimizes the template recommendation order on top of the scoring model, ensuring that not only are individual template scores reasonable, but the overall recommendation list's order also aligns with the user's actual browsing preferences.
[0221] This invention also provides a multimodal visual orchestration and recommendation device, comprising: at least one memory and at least one processor;
[0222] The at least one memory is used to store a machine-readable program;
[0223] The at least one processor is used to call the machine-readable program to implement the multimodal visual orchestration and recommendation method described in the above embodiments.
[0224] This invention also provides a computer-readable medium storing computer instructions. When executed by a processor, the computer instructions cause the processor to perform the multimodal visual orchestration recommendation method described in the above embodiments. Specifically, a system or apparatus equipped with a storage medium storing software program code that implements the functions of any of the above embodiments can be provided, and the computer (or CPU or MPU) of the system or apparatus can read and execute the program code stored in the storage medium.
[0225] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.
[0226] Examples of storage media used to provide program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.
[0227] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0228] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0229] The present invention has been shown and described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above embodiments, those skilled in the art will know that more embodiments of the present invention can be obtained by combining the code review methods in the different embodiments. These embodiments are also within the protection scope of the present invention.
Claims
1. A multimodal visual orchestration and recommendation method, characterized in that, Visualized orchestration and recommendation is achieved based on multimodal input parsing, dynamic hybrid recommendation models, and intelligent optimization algorithms, including the following steps: (1) Multimodal intent parsing: Intelligent parsing of multimodal inputs is achieved by using a base model and a fine-tuning model together. The multimodal inputs include user text, voice, and sketch inputs. The cross-attention mechanism is used to fuse text, image, and voice features to parse the user's complex intent. By combining pre-trained language models and domain-adaptive fine-tuning techniques, high-precision intent classification is achieved, and dynamic prompt word recommendations are triggered, especially when confidence is insufficient, multi-turn dialogue and dynamic prompt word recommendations are triggered. (2) Intelligent layout generation: A genetic algorithm is used for global optimization of component space allocation, combined with a constraint solver for business rule adaptation, and a graph neural network is used to model the interaction relationships between components. (3) Dynamic hybrid recommendation: Construct a three-level recommendation architecture that includes collaborative filtering, content matching, and reinforcement learning; adjust the recommendation strategy weights in real time based on user feedback; and output recommendation templates based on the scoring and sorting of the above rules. (4) Closed-loop optimization: We continuously optimize the parameters of the intent recognition model and recommendation model using user interaction behavior data.
2. The multimodal visual orchestration and recommendation method according to claim 1, characterized in that, For step (1), the dynamic prompt word recommendation is generated based on industry knowledge base and user historical behavior data; The domain adaptation fine-tuning is implemented using LoRA technology.
3. The multimodal visual orchestration and recommendation method according to claim 1, characterized in that, For step (2), the fitness function of the genetic algorithm includes three optimization objectives: information density, overlap penalty, and distribution uniformity. The constraint solver supports dynamic injection and forced application of business rules.
4. The multimodal visual orchestration and recommendation method according to claim 1, characterized in that, The dynamic hybrid recommendation process adopts a three-level architecture: the first level performs coarse screening based on business rules; the second level uses Faiss for vector similarity retrieval; and the third level performs fine ranking through the LambdaMART model. The Q-Learning algorithm is used to dynamically adjust the weights of the recommendation strategy. The model parameters are updated based on user clicks, dwell times, and deletion behaviors.
5. A multimodal visual orchestration and recommendation method according to claim 1 or 4, characterized in that, The specific implementation of recommendation matching and optimization is as follows: Candidate set generation: Multimodal features are extracted from user input, including text intent and data features. Based on the extracted features, a rule engine is used to filter candidate templates that match the features from a predefined template library to generate a candidate set. Then, the Faiss library is used to retrieve similar templates to achieve template matching. In the candidate set generation stage, a multi-level filtering strategy is adopted to achieve efficient and accurate template retrieval: The system first extracts structured features from the user's multimodal input, including the parsed business intent, data features, and interaction preference dimensions; based on the parsed features, the rule engine performs the first round of hard screening. After completing the rule filtering, the system starts similarity matching based on dense vector retrieval, where all templates are pre-encoded into 768-dimensional feature vectors through a deep representation learning model and a Faiss efficient index is constructed; during the query, the system projects the user's demand features onto the same vector space and quickly finds the Top-N most relevant templates through inner product similarity calculation; Multi-dimensional rating: In the collaborative filtering dimension, the template recommendation priority is calculated based on the user's historical behavior, and the dot product CF_score of the user's latent features and the template's latent features is calculated to represent the degree of preference of the target user and similar user groups for the template. In the content matching dimension, template metadata, including component type and layout style, is encoded into a vector. The matching degree between template metadata and user input features is calculated, and the matching degree between data features extracted from user intent and components is injected to reflect the degree of matching between the template and the current requirements (CB_score), thus achieving content matching. At the same time, business rules are preset by the administrator for mandatory matching rules in various fields. Binary judgment is used, with Rule_score being 1 if the rule is fully met, and 0 otherwise. This dimension has a fixed weight of 0.2 to ensure key business requirements are met. By taking real-time user actions such as adopting templates and adjusting layouts as input, the Q-Learning model is used to predict the immediate reward value RL_score, capture short-term changes in user preferences, and achieve real-time feedback. In the real-time calculation phase, multiple dimensions of scores are calculated in parallel for each candidate template. The weights are adjusted according to the latest user behavior, and the weighted sum is used to obtain the final score. Sorting optimization: The LambdaMART model is used to optimize the NDCG metric based on gradient boosting trees.
6. The multimodal visual orchestration and recommendation method according to claim 1, characterized in that, The multimodal intent parsing specifically includes: (1.1) Multimodal input processing: Intelligent parsing of three input methods: text, sketch and voice; Text parsing: Qwen-72B-Instruct was selected as the base model for general semantic understanding; LoRA was selected as the fine-tuning model to adapt to vertical domains; the text parsing process includes: extracting key fields based on the BiLSTM-CRF model for entity recognition; and using the TextCNN model to determine the user's target type for intent classification. Sketch Analysis: The user's hand-drawn area is accurately segmented using the SAM model, and the segmented graphic elements are classified and identified using the ResNet-50 model, mapping them to specific visualization component types; Speech parsing: High-precision speech-to-text conversion is achieved using the Whisper-large-v3 model. This model is fine-tuned by Whisper to adapt to business terms and uses TensorRT to accelerate inference, ensuring real-time response. (1.2) Multimodal feature fusion: A dynamic association model is established between text features and visual / speech features, where text features serve as query vectors and image or speech features serve as both key and value vectors. In a 768-dimensional feature space, various input features are first mapped to a unified representation space through a linear transformation layer. Then, the dot product similarity between the query vector and the key vector is calculated. After softmax normalization, an attention weight matrix is generated. This weight matrix dynamically represents the correlation between text and visual / speech features. Finally, the fused joint feature representation is obtained by weighted aggregation of the value vectors. (1.3) Intent classification and confidence assessment: A deep neural network architecture is used to perform semantic understanding and reliability determination on the fused multimodal features; a 768-dimensional joint feature vector from the feature fusion layer is received as input, and it is mapped to multiple preset intent category spaces through a fully connected layer. The number of categories is adjusted according to the actual application scenario. The network first performs a linear transformation on the input features to generate unnormalized classification scores, and then uses a softmax function to transform them into probability distributions, representing the likelihood of various intentions. The confidence assessment mechanism is implemented by calculating the maximum value of the probability distribution, which intuitively reflects the model's certainty about the current classification result. When the confidence is lower than a preset threshold, the system automatically triggers a prompt word recommendation process, guiding the user to supplement information through multiple rounds of dialogue; otherwise, it directly outputs a high-confidence intent judgment result. (1.4) Dynamic prompt word recommendation: A multi-source data fusion and hybrid retrieval strategy is adopted to achieve accurate interactive guidance. When the confidence of intent recognition is insufficient, a two-layer retrieval mechanism is triggered: First, a keyword inverted index is constructed based on the TF-IDF algorithm. By calculating the weighted word frequency similarity between the user input text and the prompt word trigger, a semantically related candidate set is quickly filtered. Then, the Faiss vector database is used for deep semantic matching. Both the user query and the prompt word are encoded into standard feature dimension vectors. The nearest neighbor search is achieved through inner product similarity calculation to ensure that the retrieval results include both literal matching and semantically related prompt content. The rules engine maintains the dialogue logic for key scenarios. When missing information or ambiguous instructions are detected, prompt words are automatically triggered for guidance. Real-time feedback optimization depends on the weight of the prompt words. After a user adopts a recommended prompt word, the weight of the corresponding prompt word increases, while the weight of other prompt words in the same category decreases accordingly. The weight is updated to form an adaptive learning feedback loop. A multi-turn dialogue manager controls the flow of dialogue based on a context state machine, and maintains the dialogue history stack to achieve a coherent interactive experience. (1.5) Data source and storage logic: The structured parameter generation process is as follows: user input → NLP parsing → JSON parameter generation; the parsed parameters are temporarily stored in Redis with a limited validity period for use in subsequent layout generation; the template containing layout parameters after user confirmation is stored in MySQL to optimize the next personalized recommendation. The prompt word library comes from a local pre-built library and a dynamically extended library. Industry-standard prompt words are stored in a local JSON file or a lightweight database, while user-defined prompt words are stored in MySQL. Role isolation ensures access security. The large AI model Qwen is only used for parsing logic.
7. The multimodal visual orchestration and recommendation method according to claim 1, characterized in that, The intelligent layout generation specifically includes: (2.1) Space allocation optimization: A grid system is used to divide the screen into logical grids. Components are assigned positions according to grid units, breakpoints are added, the grid base number is switched according to the screen width, and layout recalculation is triggered when a screen size change is detected. Each layout scheme is encoded into a gene sequence using a genetic algorithm to describe the position and size of the components. A fitness function is designed to evaluate the layout quality, maximize information density and reduce crowding. By setting target weights, the relationship between multiple objectives is balanced, and a near-optimal layout scheme is quickly found in the huge layout solution space, i.e., the possible combinations of component positions and sizes. The fitness calculation is accelerated using a GPU, and the fitness values of the evaluated genes are cached; the fitness function is as follows: Fitness=α·InfoScore-β·OverlapPenalty-γ·MaqrginViolation; InfoScore: Calculated based on component weight, which is the product of component importance weight and component area ratio. Component importance weight is defined by the user or calculated based on historical click-through rates. OverlapPenalty: Penalty for the area of overlapping regions; MaqrginViolation: Penalty for violating minimum margin constraints; (2.2) Constraint Solution: The Cassowary algorithm is used to resolve constraints and correct the layout generated by the genetic algorithm. The constraint rules include hard constraints that must be satisfied and soft constraints that should be satisfied as much as possible. The algorithm implementation process is to determine whether the candidate layout generated by the genetic algorithm satisfies all hard constraints. If it does, the layout is retained; if not, the Cassowary solver is called to correct it and output the corrected layout. When the user drags the component, the constraints are updated in real time and the solution is recalculated. To optimize the constraint solving effect, irrelevant constraints are grouped and solved independently, and incremental updates are performed, that is, only the constraints affected by user operations are recalculated. (2.3) Component Relationship Modeling: This paper adopts a Graph Neural Network (GNN) to replace the manual configuration of component linkage rules, realizing intelligent modeling and prediction of relationships between visualized components. Each visualized component is abstracted as a node in a graph structure, and the node features contain 64-dimensional feature vectors. The interaction relationship between components is represented as an edge in the graph, and the relationship type and strength are characterized by 16-dimensional feature vectors. The model architecture adopts a multi-layer message passing mechanism, with each layer containing three core processing stages: First, the edge feature processor edge_mlp generates messages. This module receives the concatenated input of source node, target node features, and edge features, and generates a 128-dimensional intermediate representation through two fully connected layers and non-linear activation, which is finally compressed into a 64-dimensional message vector. Second, neighborhood message aggregation is performed, and the index_add operation is used to accumulate the message vectors received by each node according to the target node index. Finally, the node features are updated through a gated recurrent unit (GRU), and a new node representation is generated by combining historical states and aggregated messages. (2.4) Dynamic responsive adaptation: Adopting an adaptive grid system, the display area is divided into dynamically adjustable logical units, and a breakpoint detection module is established to continuously track screen size changes. By monitoring device display characteristics in real time, optimal visualization is achieved across terminals. When a display area size change is detected to reach a preset threshold, the system automatically triggers a four-stage processing flow: First, the current device type and resolution range are identified through media query rules; then, the optimal space allocation scheme for each visualization component is recalculated, where the main components maintain aspect ratio scaling, and secondary components are vertically stacked or collapsed according to priority; next, the Cassowary constraint solver is applied to ensure that the adjusted layout meets business rules and aesthetic requirements; finally, the rendering engine is updated in real time; for small-screen devices, the system intelligently implements information hierarchy compression strategies, automatically collapses detailed data, and highlights KPI indicators. An integrated intelligent mock data generator automatically synthesizes simulated datasets that conform to business semantics based on the semantic type of the components.
8. A multimodal visual orchestration and recommendation system, characterized in that, include: A multimodal input parsing module is used to realize intent recognition and feature fusion of text, sketches, and speech; The hybrid layout optimization module integrates a genetic algorithm engine, a constraint solver, and a graph neural network processor. The dynamic recommendation engine module includes rule filters, vector retrieval systems, and ranking models. The feedback learning module is used to collect user behavior data and update model parameters; The system achieves multimodal visual orchestration recommendation through the method described in any one of claims 1 to 7.
9. A multimodal visual orchestration and recommendation device, characterized in that, include: At least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to invoke the machine-readable program to implement the method according to any one of claims 1 to 7.
10. A computer-readable medium, characterized in that, The computer-readable medium stores computer instructions that, when executed by a processor, enable the implementation of the method described in any one of claims 1 to 7.
Citation Information
Cited By
Intelligent chart generation method and system suitable for multi-modal data and storage medium
CN121117072A
Dynamic instrument panel generation method, device and equipment based on user behavior analysis
CN121166996A
AIGC-based interactive large-screen real-time drawing method and system
CN121213737A
Scenarized decision interface dynamic generation and continuous optimization method based on AI
CN121255153A
AI-based scenario-based decision interface dynamic generation and continuous optimization method
CN121255153B