An AI large model-based cross-modal agent base deployment system and method

By constructing a cross-modal intelligent agent foundation system, utilizing multi-source heterogeneous databases and multimodal knowledge graphs for data fusion, and combining feature analysis and intelligent routing decisions, the system optimizes hierarchical retrieval and feedback mechanisms, thus solving the problems of low data processing efficiency, limited model performance, and unnatural human-computer interaction in the cross-modal intelligent agent foundation system, and achieving efficient intelligent decision-making and natural interaction.

CN121052282BActive Publication Date: 2026-04-24HUADIAN ELECTRIC POWER SCI INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUADIAN ELECTRIC POWER SCI INST CO LTD
Filing Date
2025-10-28
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing cross-modal intelligent agent deployments suffer from low data processing efficiency, limited model performance, poor adaptability, and unnatural human-computer interaction, failing to meet the needs of real-time response and intelligent decision-making in complex environments.

Method used

A cross-modal intelligent agent foundation system based on a large AI model is constructed, including a data layer, a processing layer, and an optimization layer. Data fusion is achieved through external multi-source heterogeneous databases and multimodal dynamic knowledge graphs. Efficient matching is performed using a multimodal feature parser and an intelligent routing decision engine. Optimization is achieved by combining a hierarchical retrieval enhancement module and a dynamic feedback optimization module.

Benefits of technology

It achieves improved response speed without sacrificing accuracy, enhances the adaptability of the intelligent agent and the naturalness of human-computer interaction, and meets the real-time requirements of multimodal data processing and intelligent decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052282B_ABST
    Figure CN121052282B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and discloses a cross-modal intelligent agent base deployment system and method based on an AI large model, which comprises the following: a data layer comprising an externally connected multi-source heterogeneous database and a multi-modal dynamic knowledge graph, a cross-modal knowledge base is constructed through unified coding and joint retrieval; a processing layer comprising a multi-modal feature parser, an intelligent routing decision engine and a three-level dynamic model matrix, the multi-modal feature parser is used for analyzing data, the intelligent routing decision engine selects a target model from the three-level dynamic model matrix for processing to generate a routing decision; and an optimization layer comprising a hierarchical retrieval enhancement module and a dynamic feedback optimization module, which optimizes the data layer and the processing layer, the application effectively fuses different modal data, selects a model for processing through the intelligent routing decision engine, improves the response speed without losing accuracy, the optimization layer improves the overall performance, the cross-modal intelligent agent base efficiently processes multi-modal data, and intelligent decision-making is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a cross-modal intelligent agent base deployment system and method based on a large AI model. Background Technology

[0002] With the rapid development of artificial intelligence technology, the research and application of intelligent agent platforms have gradually become a hot topic. Intelligent agent platforms enable intelligent agents to perform analysis based on data and algorithms, giving them powerful reasoning and decision-making capabilities and providing strong support for decision-making. Against this backdrop, how to deploy intelligent agent platforms has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, the present invention provides a cross-modal intelligent agent base deployment system and method based on a large AI model to solve the problem of how to deploy the intelligent agent base.

[0004] In a first aspect, the present invention provides a cross-modal intelligent agent base deployment system based on a large AI model. This system includes a data layer, a processing layer, and an optimization layer, wherein...

[0005] The data layer includes external multi-source heterogeneous databases and multimodal dynamic knowledge graphs. External multi-source heterogeneous databases are used to mount data resources and to build structured knowledge networks based on multimodal data. External multi-source heterogeneous databases and multimodal dynamic knowledge graphs build cross-modal knowledge bases through unified encoding and joint retrieval.

[0006] The processing layer includes a multimodal feature parser, an intelligent routing decision engine, and a three-level dynamic model matrix. The multimodal feature parser is used to parse multimodal data from business instructions and the data layer. The three-level dynamic model matrix includes lightweight, general-purpose, and professional-grade AI large models. The intelligent routing decision engine is used to select the target model from the three-level dynamic model matrix to process the input data and generate routing decisions.

[0007] The optimization layer includes a hierarchical retrieval enhancement module and a dynamic feedback optimization module. The hierarchical retrieval enhancement module is used to optimize the data in the data layer, while the dynamic feedback optimization module is used to push optimization parameters to the data layer and the processing layer to optimize them.

[0008] This invention constructs a data layer using an external multi-source heterogeneous database and a multimodal dynamic knowledge graph, effectively integrating data from different modalities and breaking through the bottleneck of single-modal data. This provides data support for decision-making. A large AI model is deployed through a three-level dynamic model matrix, and the intelligent routing decision engine selects appropriate models from this matrix for processing, achieving efficient matching between "problem" and "model." This improves response speed without sacrificing accuracy. An optimization layer is constructed through a hierarchical retrieval enhancement module and a dynamic feedback optimization module, optimizing both the data and processing layers to improve overall performance. This enables the cross-modal intelligent agent platform to efficiently process multimodal data and achieve intelligent decision-making.

[0009] In one optional implementation, the dynamic feedback optimization module employs a dynamic feedback optimization mechanism consisting of a fast loop, a slow loop, and a validation loop. The fast loop is used to adjust the routing strategy parameters, the slow loop is used to update the knowledge graph embedding, and the validation loop is used to perform incremental training on the model matrix.

[0010] This invention improves the operational stability of the agent by constructing a dynamic feedback optimization mechanism consisting of a fast loop, a slow loop, and a verification loop, which optimizes routing strategies, knowledge graphs, and model capabilities, covering the overall optimization of the data layer and the processing layer.

[0011] Secondly, the present invention provides a method for deploying a cross-modal intelligent agent platform based on an AI large model, applicable to a cross-modal intelligent agent platform deployment system based on an AI large model, the method comprising:

[0012] Unified encoding and joint retrieval of external multi-source heterogeneous databases and multimodal dynamic knowledge graphs are used to construct a cross-modal knowledge base;

[0013] A multimodal feature parser is used to parse business instructions and multimodal data to extract feature representations;

[0014] The intelligent routing decision engine selects the target model from the three-level dynamic model matrix and uses the target model to process the input data to generate routing decisions.

[0015] This invention constructs a cross-modal knowledge base by uniformly encoding and jointly retrieving external multi-source heterogeneous databases and multimodal dynamic knowledge graphs. This breaks down the information barriers of traditional single-modal data, providing a data foundation for intelligent agents. A multimodal feature parser is used to parse the data, obtaining feature representations suitable for subsequent model processing. An intelligent routing decision engine selects a suitable model from a three-level dynamic model matrix for decision-making, enabling the cross-modal intelligent agent base to efficiently process multimodal data and achieve intelligent decision-making.

[0016] In one optional implementation, unified encoding and joint retrieval are performed on external multi-source heterogeneous databases and multimodal dynamic knowledge graphs, including:

[0017] By extracting the logical relationships between data from external multi-source heterogeneous databases using cross-modal association technology, performing unified encoding, and constructing a multimodal dynamic knowledge graph;

[0018] Similarity retrieval is performed in a multimodal dynamic knowledge graph using the encoding vectors of each component.

[0019] This invention uses cross-modal association technology to uniformly encode data from external multi-source heterogeneous databases, constructs a multimodal dynamic knowledge graph, effectively integrates data from different modalities, and performs similarity retrieval in the multimodal dynamic knowledge graph to quickly match relevant cases, thereby improving knowledge retrieval efficiency and problem-solving speed.

[0020] In one optional implementation, a multimodal feature parser is used to parse the business instructions and multimodal data, extracting feature representations, including:

[0021] If the multimodal data types of business instructions and data layers are text, then word vector technology is used to map the text to a low-dimensional vector space to obtain the vector representation of the text;

[0022] If the multimodal data type of the business instructions and data layer is an image, then the image is used to extract features through a CNN architecture to obtain a high-level semantic feature map of the image, and then the high-level semantic feature map of the image is converted into a feature vector.

[0023] If the multimodal data types of business instructions and data layers are time-series data, then feature extraction is performed on the time-series data according to the meaning of the data to form a feature vector of the time-series data.

[0024] This invention employs corresponding feature extraction methods for different data types. It uses word vector technology to extract features from text data, accurately capturing the semantic information of the text. It uses a CNN architecture to extract features from images, preserving the spatial structure and abstract semantics of the images. It also extracts features from time-series data based on the meaning of the data, preserving the trend and periodicity of the time-series data.

[0025] In one alternative implementation, a target model is selected from a three-level dynamic model matrix using an intelligent routing decision engine, including:

[0026] If the problem complexity is simple, then choose a lightweight model as the target model;

[0027] If the problem complexity is moderate, then a general-purpose model should be selected as the target model.

[0028] If the problem has a high degree of complexity, then a professional-grade model should be selected as the target model.

[0029] If the problem is a single-modal problem, then choose a lightweight model or a general-purpose model as the target model;

[0030] If the problem is multimodal, the target model should be selected based on the requirements of domain specialization, real-time requirements, and accuracy requirements.

[0031] This invention achieves a balance between resource consumption and problem handling by selecting the corresponding model as the target model based on the problem complexity and modality type, thereby improving the adaptability of the target model to the application scenario.

[0032] In an alternative implementation, after processing the input data using the target model to generate a routing decision, the method further includes:

[0033] Calculate the confidence level of the response output and determine whether the confidence level is less than the preset confidence threshold;

[0034] If the confidence level is less than the preset confidence threshold, optimization will be carried out through a dynamic feedback optimization mechanism.

[0035] This invention optimizes and adjusts the response output result through a dynamic feedback mechanism when the confidence level does not meet the conditions, thereby avoiding the use of results below the preset confidence threshold for decision-making and preventing decision-making errors. This ensures the reliability of the decision support provided by the intelligent agent base.

[0036] In an alternative implementation, after processing the input data using the target model to generate a routing decision, the method further includes:

[0037] The intelligent routing decision engine is used to determine whether professional processing of routing decisions is required.

[0038] If specialized processing of routing decisions is required, then determine whether there are domain entities in the multimodal dynamic knowledge graph;

[0039] If domain entities exist in the multimodal dynamic knowledge graph, a professional-grade model is activated, and the retrieval is enhanced.

[0040] If no domain entities exist in the multimodal dynamic knowledge graph, it is downgraded to a general-level model.

[0041] This invention switches the target model based on the results of whether professional processing of routing decisions is required and whether domain entities exist, selects the target model that matches the multimodal dynamic knowledge graph, improves the adaptability of the model, and optimizes resource allocation.

[0042] In one alternative implementation, the method further includes:

[0043] If no specialized processing of routing decisions is required, then downgrade to a lightweight model.

[0044] This invention avoids redundant computing power and waste of resources by downgrading the model when no professional processing of routing decisions is required.

[0045] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the deployment method of the cross-modal intelligent agent base based on the AI ​​large model described in the second aspect or any corresponding embodiment thereof. Attached Figure Description

[0046] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating a cross-modal intelligent agent base deployment system based on an AI large model according to an embodiment of the present invention.

[0048] Figure 2 This is a flowchart illustrating the deployment method of a cross-modal intelligent agent base based on an AI large model according to an embodiment of the present invention;

[0049] Figure 3 This is a logical diagram of the intelligent routing decision engine according to an embodiment of the present invention;

[0050] Figure 4 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] Currently, the deployment of cross-modal intelligent agent bases is mainly reflected in the following aspects:

[0053] (1) In the process of processing and fusing multimodal data, existing methods usually process data of different modalities independently, such as using convolutional neural networks to process visual data, using recurrent neural networks to process text data, and then integrating the processed data through simple splicing or weighted fusion.

[0054] (2) When building intelligent agents based on large AI models, existing technologies generally rely on pre-trained models and then fine-tuning them according to specific tasks. The fine-tuning process is mainly based on predefined loss functions and optimization algorithms, using limited training data to adjust model parameters to adapt to specific tasks.

[0055] (3) The deployment of large AI models needs to take into account the limitations of computing and storage resources. Existing deployment methods often use model compression, quantization and other techniques to reduce the model size and reduce the demand for resources, while improving the model running efficiency through distributed computing and parallel processing.

[0056] (4) In terms of human-computer interaction, cross-modal intelligent agents mainly focus on voice interaction and simple text interaction. They can recognize and understand the voice commands and text inputs of human users and generate corresponding voice or text responses.

[0057] However, the relevant technologies have the following drawbacks in deploying cross-modal intelligent agent bases:

[0058] (1) Low data processing efficiency: The use of independent processing modules and simple fusion methods results in low data processing efficiency, which cannot respond to changes in multimodal information in real time and limits the perception and decision-making capabilities of intelligent agents in dynamic environments. For example, in autonomous driving scenarios, intelligent agents need to process visual, lidar and auditory data in real time and make decisions, which cannot meet the real-time requirements;

[0059] (2) Limited model performance: Existing fine-tuning methods do not fully utilize actual operational feedback information, resulting in limited performance improvement. Although model compression and quantization reduce the scale, they also sacrifice some performance, affecting the accuracy of the agent. For example, in complex medical diagnosis scenarios, insufficient model performance may lead to inaccurate diagnosis;

[0060] (3) Poor adaptability of intelligent agents: Existing intelligent agents lack full utilization of feedback information, making it difficult to maintain efficient operation in complex and ever-changing environments, thus limiting their application scope. For example, in intelligent customer service scenarios, facing diverse user problems and changing needs, poorly adaptable intelligent agents may not be able to provide effective solutions;

[0061] (4) Unnatural human-computer interaction: It mainly relies on voice and text interaction and lacks a deep understanding of the multimodal interaction intentions of human users, resulting in unnatural and inefficient human-computer interaction and failure to meet the personalized service needs of users. For example, in the intelligent education scenario, when teachers interact with intelligent agents in a multimodal manner, the unnatural human-computer interaction will affect the teaching effect and user experience.

[0062] To address the aforementioned problems, embodiments of the present invention provide a cross-modal intelligent agent foundation deployment system based on a large AI model, such as... Figure 1 As shown, the system comprises three layers: a data layer, a processing layer, and an optimization layer.

[0063] At the data layer, an external multi-source heterogeneous database and a multimodal dynamic knowledge graph are designed. The external multi-source heterogeneous database loads various professional knowledge according to business needs, providing rich data resources, and forms a multimodal dynamic knowledge graph through cross-modal parallel connection. The external multi-source heterogeneous database and the multimodal dynamic knowledge graph are managed through unified encoding and joint retrieval, constructing a dynamic cross-modal knowledge base, which constitutes the knowledge foundation of the intelligent agent.

[0064] Compared to static, single-modal knowledge bases in related technologies, the data layer deployed in this embodiment of the invention can better support knowledge acquisition and updating for cross-modal intelligent agents.

[0065] At the processing layer, a multimodal feature parser, an intelligent routing decision engine, and a three-level dynamic model matrix are introduced. The multimodal feature parser is used to parse business instructions and multimodal data from the data layer. The three-level dynamic model matrix includes three levels of AI large models: lightweight, general-purpose, and professional. The intelligent routing decision engine selects the most suitable model from the three levels of AI large models in the three-level dynamic model matrix as the target model for processing based on the characteristics and requirements of the input problem, thereby achieving efficient matching of "problem-model".

[0066] The processing layer deployed in this embodiment of the invention can improve response speed and reduce energy consumption without sacrificing accuracy. Furthermore, the three-level dynamic model matrix in the processing layer can be updated according to the development of AI technology, enabling the parallel deployment of AI models of various scales, improving flexibility and scalability. This solves the problem of performance loss that may occur due to model compression and quantization in existing technologies, ensuring that the agent can always use the most advanced and suitable model for processing, thereby improving generalization ability.

[0067] At the optimization layer, the entire system is dynamically optimized through the hierarchical retrieval enhancement module and the dynamic feedback optimization module. The hierarchical retrieval enhancement module optimizes the data in the data layer by optimizing the retrieval algorithm and strategy, thereby improving the efficiency and accuracy of information retrieval from the data layer. The dynamic feedback optimization module is used to dynamically push optimized data to the data layer and the processing layer, perform multi-level optimization adjustments, realize the optimization of the data layer and the processing layer, and thus continuously improve the performance and efficiency of the intelligent agent.

[0068] The hierarchical retrieval enhancement module processes the response output to obtain key information. After the key information is judged as useful by the dynamic feedback optimization module, it is fed back to the joint retrieval system of the data layer as the basis for adjusting and optimizing the retrieval algorithm and strategy.

[0069] The retrieval algorithms employed by the joint retrieval system include, but are not limited to, inverted index algorithms, vector retrieval algorithms, and graph traversal algorithms. Retrieval strategies include, but are not limited to, multi-level retrieval, semantic association retrieval, and dynamically adjusted retrieval. The optimization mechanisms employed by the dynamic feedback optimization module include, but are not limited to, feature-weighted optimization, index structure optimization, and caching mechanism optimization.

[0070] Furthermore, the optimization layer deploys online learning algorithms. The underlying principle is that the optimization layer feeds back information to the data layer for re-retrieval, and then the processing layer reprocesses and generates new content. The basic principle is that if the output of the processing layer does not meet the confidence threshold of the optimization layer, the dynamic feedback module organizes the entire process record information triggered by the business instruction and feeds it back to the joint retrieval system. This information is then broken down into more detailed retrieval instructions, expanding the breadth and depth of the database query. The re-retrieved information is then supplied to the processing layer, which integrates the previously low-confidence processing information for reprocessing. Reprocessing is necessary because it involves more information, and the system automatically determines whether to invoke a higher-level model based on the actual situation. This process is repeated until the confidence of the output content meets the requirements. Then, a corresponding knowledge graph is generated to record the responses to the business input. The next time a similar business input is encountered, the knowledge graph is prioritized for processing the recorded information and corresponding parameters at each level, directly processing from the optimal parameters, thereby achieving learning and optimization of the underlying processing capabilities.

[0071] The optimization layer deployed in this embodiment of the invention continuously monitors and adjusts its operating status to ensure that the intelligent agent adapts to constantly changing task requirements and maintains optimal performance.

[0072] The cross-modal intelligent agent base deployment system based on AI large model provided in this embodiment uses an external multi-source heterogeneous database and a multimodal dynamic knowledge graph to form a data layer, effectively integrating data from different modalities, breaking the bottleneck of single-modal data, and providing data support for decision-making. It deploys the AI ​​large model through a three-level dynamic model matrix, and the intelligent routing decision engine selects appropriate models from the matrix for processing, achieving efficient matching between "problem-model" and improving response speed without sacrificing accuracy. An optimization layer is constructed through a hierarchical retrieval enhancement module and a dynamic feedback optimization module to optimize the data layer and processing layer, improving overall performance. This enables the cross-modal intelligent agent base to efficiently process multimodal data and achieve intelligent decision-making.

[0073] Specifically, in the optimization layer, the dynamic feedback optimization module employs a dynamic feedback optimization mechanism consisting of a fast loop, a slow loop, and a validation loop. The fast loop adjusts routing strategy parameters at the minute level, the slow loop updates the knowledge graph embedding at the day level, and the validation loop is used for incremental training of the model matrix. This optimization mechanism dynamically pushes optimization parameters to the data layer and processing layer based on information such as accuracy, latency, and resource consumption from the end-to-end processing logs, thereby continuously improving the performance and efficiency of the agent.

[0074] By constructing a dynamic feedback optimization mechanism consisting of a fast loop, a slow loop, and a verification loop, optimizations are made from the aspects of routing strategy, knowledge graph, and model capabilities, covering the overall optimization of the data layer and the processing layer, thereby improving the operational stability of the agent.

[0075] This AI-based cross-modal intelligent agent platform is used in industrial digital systems to assist in handling specialized business needs. Examples include analyzing diagnostic information and optimizing maintenance recommendations for wind turbines, as well as data analysis and interpretation of various indicators.

[0076] Human-computer interaction mainly involves the input of business commands, which are then decomposed into different granularities by a multimodal feature parser, and processed by a routing decision engine that calls large models of different scales.

[0077] Different business needs are addressed by calling different levels of models, thereby enabling the understanding of the user's interactive intent in response to multimodal business commands. At the same time, because different levels of models are called to handle different problems, various business needs can receive appropriate and natural responses.

[0078] By incorporating a multi-granularity semantic extraction and interaction mechanism into the human-computer interaction module, and utilizing a hierarchical semantic parsing network and a multi-granularity interaction module, the system can gain a deeper understanding of the intentions used for multimodal interaction. This enables the intelligent agent to better grasp the personalized needs of users, provide more natural and considerate services, improve the experience and effectiveness of human-computer interaction, and overcome the shortcomings of existing technologies in human-computer interaction that are unnatural and inefficient.

[0079] According to an embodiment of the present invention, a method for deploying a cross-modal intelligent agent base based on a large AI model is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0080] This embodiment provides a method for deploying a cross-modal intelligent agent platform based on a large AI model, which is applied to the aforementioned cross-modal intelligent agent platform deployment system based on a large AI model. Figure 2 This is a flowchart of a cross-modal intelligent agent base deployment method based on an AI large model according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0081] Step S201: Unify the encoding and joint retrieval of external multi-source heterogeneous databases and multimodal dynamic knowledge graphs to construct a cross-modal knowledge base.

[0082] In this embodiment of the invention, a dynamic cross-modal knowledge base is constructed by managing the external multi-source heterogeneous databases and multimodal dynamic knowledge graphs of the data layer through a unified coding and joint retrieval system.

[0083] Step S202: Use a multimodal feature parser to parse the business instructions and multimodal data, and extract feature representations.

[0084] In this embodiment of the invention, a multimodal feature parser in the processing layer is used to parse business instructions and multimodal data from the data layer, and extract their feature representations as the data basis for subsequent processing.

[0085] Step S203: Use the intelligent routing decision engine to select the target model from the three-level dynamic model matrix, and use the target model to process the input data to generate a routing decision.

[0086] In this embodiment of the invention, the intelligent routing decision engine selects the most suitable model from the three-level dynamic model matrix as the target model to process the input data based on the characteristics and requirements of the input problem, and generates a routing decision or answer.

[0087] The cross-modal intelligent agent base deployment method based on AI large model provided in this embodiment constructs a cross-modal knowledge base by uniformly encoding and jointly retrieving external multi-source heterogeneous databases and multimodal dynamic knowledge graphs. This breaks down the information barriers of traditional single-modal data, provides a data foundation for intelligent agents, and uses a multimodal feature parser to parse the data to obtain feature representations suitable for subsequent model processing. The intelligent routing decision engine selects a suitable model from the three-level dynamic model matrix for decision-making, enabling the cross-modal intelligent agent base to efficiently process multimodal data and achieve intelligent decision-making.

[0088] This embodiment provides a method for deploying a cross-modal intelligent agent platform based on a large AI model. The process includes the following steps:

[0089] Step S301: Unify the encoding and joint retrieval of external multi-source heterogeneous databases and multimodal dynamic knowledge graphs to construct a cross-modal knowledge base.

[0090] Specifically, step S301 includes:

[0091] Step S3011: Extract the logical relationships between data in the external multi-source heterogeneous database through cross-modal association technology, perform unified encoding, and construct a multimodal dynamic knowledge graph.

[0092] Step S3012: Similarity retrieval is performed in the multimodal dynamic knowledge graph using the encoding vectors of each component.

[0093] In this embodiment of the invention, data from various databases within an external multi-source heterogeneous database is first integrated and preprocessed to ensure accuracy, completeness, and consistency. Then, cross-modal association technology is used to extract logical association information such as entities, relationships, and attributes from data of different modalities, and this information is then classified and vectorized. Next, corresponding association links are established based on specific business expertise, and unified encoding is implemented. The association information and encoded information are integrated into a knowledge graph. Through the unified structured representation of the graph, the fusion and storage of multimodal data are achieved, thereby constructing a multimodal dynamic knowledge graph that comprehensively and accurately reflects the complex relationships between multi-source heterogeneous data.

[0094] Specifically, encoding is based on Graph Neural Network (GNN). Taking a wind turbine system as an example, the various components of the wind turbine and their multimodal features are constructed as a graph structure. Each component is a node, and the connections between the components (such as mechanical connections, electrical connections, etc.) are edges. This graph structure intuitively reflects the connection relationships between the various parts of the wind turbine.

[0095] Each component node is assigned a fused feature vector, which is then used as the node feature input into a graph neural network (GNN). Through GNN's graph convolution operations and information transfer mechanism, a graph neural network encoding representation of each component is obtained. Specifically, feature vectors are extracted from multimodal data such as text, images, and time-series data, and then fused using simple concatenation and an attention-based mechanism to obtain the fused feature vector.

[0096] This coding method not only takes into account the multimodal characteristics of each component, but also incorporates the structural relationship information between components, which can better capture the overall characteristics of the wind turbine system and the mutual influence between components.

[0097] The fused feature vector obtained through multimodal feature fusion is used as the core attribute of each component node in the wind turbine system and stored in the multimodal dynamic knowledge graph. For example, for the wind turbine blade node, its encoding vector contains multimodal fused features such as the blade's text, image, and operating status. This information enriches the semantic expression of the node, enabling each component node in the knowledge graph to accurately reflect its actual characteristics and status.

[0098] When constructing a multimodal dynamic knowledge graph, the encoded nodes of different components are connected according to the actual connection relationships between them (such as mechanical connections, functional associations, data associations, etc.). For example, wind turbine blades and gearboxes are related in mechanical transmission; corresponding edges are established in the knowledge graph, and the weights of the edges are determined by calculating the similarity between the encoded representations, representing the tightness of the association. Simultaneously, information such as textual semantics and image features in the encoded representations is used to establish associations between components and conceptual nodes such as fault types and maintenance operations in the knowledge graph, forming a more complete structured knowledge network.

[0099] When multimodal data of wind turbine components is updated (e.g., new operating status data, maintenance records, component damage images, etc.), feature extraction and fusion encoding are re-performed, and the encoding attributes of the corresponding component nodes in the knowledge graph are updated in a timely manner. For example, when a new damage image appears on a wind turbine blade and is updated after encoding, the blade node encoding in the knowledge graph is updated accordingly to reflect the latest status change. Simultaneously, based on the correlation analysis between the new encoding and other nodes in the graph, the weights and relationships of relevant edges are dynamically adjusted to maintain the timeliness and accuracy of the knowledge graph.

[0100] When searching and querying in a multimodal dynamic knowledge graph, similarity searches are performed using the encoding vectors of each component. For example, when searching for cases similar to a specific faulty component, the similarity of the encoding vectors is compared to quickly locate component nodes with similar characteristics in the knowledge graph. Furthermore, the graph's relationships allow for the acquisition of relevant information such as fault causes and solutions, thereby improving the efficiency and accuracy of the search.

[0101] By using cross-modal association technology to uniformly encode data from external multi-source heterogeneous databases, a multimodal dynamic knowledge graph is constructed. This effectively integrates data from different modalities, enabling similarity retrieval within the multimodal dynamic knowledge graph to quickly match relevant cases, thereby improving knowledge retrieval efficiency and problem-solving speed.

[0102] Step S302: Use a multimodal feature parser to parse the business instructions and multimodal data, and extract feature representations.

[0103] Specifically, step S302 includes:

[0104] Step S3021: If the multimodal data type of the business instruction and the data layer is text, then the text is mapped to a low-dimensional vector space through word vector technology to obtain the vector representation of the text.

[0105] Step S3022: If the multimodal data type of the business instruction and the data layer is an image, then the image is used to extract features through a CNN architecture to obtain a high-level semantic feature map of the image, and the high-level semantic feature map of the image is converted into a feature vector.

[0106] Step S3023: If the multimodal data type of the business instruction and data layer is time series data, then the time series data is feature extracted according to the meaning of the data to form the feature vector of the time series data.

[0107] In this embodiment of the invention, the multimodal feature parser parses business instructions and multimodal data, which may be of data types such as text, images, videos, and audio. For the information input to the multimodal feature parser, the parser first identifies the data type and checks the data quality. Different tools are then used to preprocess different modal information, extracting key information and converting it into feature vectors to prepare for subsequent processing.

[0108] For cases where business instructions and data layers use text as the multimodal data type, word vector technology is used to map the text to a low-dimensional vector space to obtain the vector representation of the text. Word vector technology includes, but is not limited to, Word2Vec, GloVe, and fastText.

[0109] For cases where the multimodal data type of business instructions and data layers is an image, features are extracted from the image using a CNN (Convolutional Neural Network) architecture to obtain a high-level semantic feature map of the image, which is then converted into a feature vector. CNN structures include, but are not limited to, VGGNet, ResNet, and InceptionNet.

[0110] For cases where the multimodal data types of business instructions and data layers are time-series data, feature extraction is performed on the time-series data according to the specific meaning of the data, including statistical features (mean, variance, maximum value, etc.), frequency features (main frequency, amplitude, etc.), trend features, periodic features, etc., to form the feature vector of the time-series data.

[0111] For different data types, corresponding feature extraction methods are used for feature extraction. Text data is extracted using word vector technology to accurately capture the semantic information of the text. Images are extracted using a CNN architecture to preserve the spatial structure and abstract semantics of the images. Time series data is extracted based on the meaning of the data to preserve the trend and periodicity of the time series data.

[0112] Step S303: Use the intelligent routing decision engine to select the target model from the three-level dynamic model matrix, and use the target model to process the input data to generate a routing decision.

[0113] Specifically, step S303 includes:

[0114] Step S3031: If the problem complexity is simple, then select the lightweight model as the target model.

[0115] Step S3032: If the problem complexity is medium, then select the general-purpose model as the target model.

[0116] Step S3033: If the problem complexity is high complexity, then select the professional-grade model as the target model.

[0117] Step S3034: If the problem is a single-modal problem, then select a lightweight model or a general-purpose model as the target model.

[0118] Step S3035: If the problem is a multimodal problem, select the target model according to the domain-specific requirements, real-time requirements, and accuracy requirements.

[0119] In this embodiment of the invention, an intelligent routing decision engine is used to select the most suitable model as the target model from a three-level dynamic model matrix based on the characteristics and requirements of the input problem. The intelligent routing decision engine achieves efficient matching between "problem-model" by analyzing information such as the complexity of the input problem and the required modality.

[0120] Specifically, for problems involving simple information retrieval or single-step analysis (i.e., simple problem complexity), a lightweight model is sufficient. For problems requiring some analysis and multi-step reasoning, such as performance evaluation of a wind turbine component or fault diagnosis of common problems (i.e., medium complexity), a general-purpose model is chosen. General-purpose models can provide relatively accurate solutions while balancing computational resources and response time. For complex situations involving multi-component interaction analysis, complex root cause analysis, or deep fusion of multimodal data (i.e., high complexity), a professional-grade model is selected. Professional-grade models leverage their in-depth knowledge and powerful analytical capabilities in specific domains to ensure data accuracy and reliability.

[0121] The model selection is based on the required modality, and can be divided into two categories: single-modal problems and multi-modal problems. For single-modal problems, only one mode of data needs to be analyzed. For example, determining the name and model of a wind turbine component based solely on text descriptions, or determining whether a parameter is normal based solely on operating status data, a lightweight or general-purpose model can be selected. Lightweight or general-purpose models can efficiently handle single-modal tasks.

[0122] For multimodal problems, it is necessary to analyze data from multiple modalities. For example, when it is necessary to combine images, text and operational data to diagnose the cause of failure in wind turbine components, a professional-grade model should be selected.

[0123] Choose a model based on the domain's specific requirements. When the problem involves little or no domain-specific knowledge, a lightweight or general-purpose model will suffice. When the problem involves domain-specific knowledge, choose a professional-grade model.

[0124] For tasks with high real-time requirements, that is, requiring the input data to be processed and the results output in a short time, a lightweight model can be selected to meet the real-time requirements.

[0125] For tasks requiring high accuracy, such as life prediction of key wind turbine components and energy output prediction of wind farms, professional-grade models are selected. These models invest more computing resources and time in in-depth data analysis to provide more accurate and reliable prediction results.

[0126] For comprehensive tasks that require both real-time performance and accuracy, such as the daily inspection of wind turbines where it is necessary to quickly identify obvious faults and conduct relatively accurate analysis of suspicious situations, a general-purpose model can meet the requirements.

[0127] Based on the computing resources required by business needs, lightweight, general-purpose, and professional-grade models are selected layer by layer, from least to most, to achieve a match between model performance and resource consumption. In addition, the model selection can be adjusted appropriately based on feedback from the optimization layer.

[0128] By selecting the appropriate model as the target model based on the problem complexity and modality type, a balance is achieved between resource consumption and problem handling, thereby improving the adaptability of the target model to the application scenario.

[0129] In some alternative implementations, the method further includes:

[0130] Step S304: Calculate the confidence level of the response output and determine whether the confidence level is less than the preset confidence threshold.

[0131] Step S305: If the confidence level is less than the preset confidence threshold, then optimization and adjustment are carried out through a dynamic feedback optimization mechanism.

[0132] In this embodiment of the invention, the confidence level of the response output is calculated, and the calculation method can be selected according to different business needs, including but not limited to probability distribution, uncertainty estimation, model ensemble, and task-specific indicators.

[0133] Taking classification problems in probability distributions as an example, the probability distribution is used as the confidence threshold. The probability distribution represents the probability that the input data belongs to each category. For example, in wind turbine fault diagnosis, if the probability of judging a certain fault type is 0.95, then this probability value is used as the confidence threshold, meaning that the model has 95% confidence that the fault type is the correct answer.

[0134] Uncertainty estimation methods are mainly implemented through Bayesian neural networks and Monte Carlo methods. Taking Bayesian neural networks as an example, when predicting the remaining service life of wind turbines, Bayesian neural networks can not only output the predicted value, but also output the uncertainty interval. The size of this interval serves as a reference for the confidence level. The smaller the interval, the higher the confidence level.

[0135] The system checks if the confidence level is lower than a preset confidence threshold. If it is, the output result is marked and enters the hierarchical retrieval system. Through a dynamic feedback optimization mechanism, the results are fed back to the intelligent routing decision engine, the three-level dynamic model matrix, and the joint retrieval system at the data layer. Specifically, the intelligent routing decision engine adjusts its model's judgment parameters, and when a similar problem is encountered again, a higher-level model is invoked for processing. The three-level dynamic model matrix calls a higher-level model for reprocessing. The joint retrieval system at the data layer stores error cases and records the retrieval algorithm and strategy. Based on the feedback marking information, the retrieval algorithm and strategy are optimized.

[0136] By optimizing and adjusting the response output when the confidence level does not meet the conditions through a dynamic feedback optimization mechanism, the system avoids using results below the preset confidence threshold for decision-making, thus preventing decision-making errors and ensuring the reliability of the decision support provided by the intelligent agent base.

[0137] In some alternative implementations, the method further includes:

[0138] Step S306: Use the intelligent routing decision engine to determine whether professional processing of the routing decision is required.

[0139] Step S307: If professional processing of routing decisions is required, determine whether there are domain entities in the multimodal dynamic knowledge graph.

[0140] Step S308: If domain entities exist in the multimodal dynamic knowledge graph, then activate the professional-grade model and enhance the retrieval.

[0141] Step S309: If there are no domain entities in the multimodal dynamic knowledge graph, then downgrade to a general-level model.

[0142] In step S310, if no professional processing of routing decisions is required, the model is downgraded to a lightweight model.

[0143] In embodiments of the present invention, such as Figure 3 As shown, the intelligent routing decision engine determines whether specialized processing is required based on the characteristics and needs of the input problem, including its complexity, the types of modalities involved, and the domain to which the problem belongs. For multimodal data with high complexity, involving specialized domains, or requiring refined processing, the intelligent routing decision engine selects a professional-grade model for processing.

[0144] When the problem involves industry-specific terminology, complex image recognition tasks (such as wind power operation image diagnosis), or requires in-depth analysis of subtle features such as acoustic vibration, the intelligent routing decision engine determines that specialized processing is needed for routing decisions. Utilizing feature representations extracted by a multimodal feature parser, combined with a structured knowledge network in a multimodal dynamic knowledge graph, semantic matching and logical association analysis are performed within the multimodal dynamic knowledge graph to search for entities such as domain concepts, objects, or attributes related to the input problem, and to determine the existence of domain entities.

[0145] If a domain-specific entity highly relevant to the characteristics of the input question is found in the multimodal dynamic knowledge graph, such as a specific person, organization, or technical term, then the existence of a domain entity is confirmed. Furthermore, the external multi-source heterogeneous database stores rich professional knowledge, providing data support for determining the existence of domain entities and further verifying their existence.

[0146] If a domain entity exists, a professional-level model is selected from the three-level dynamic model matrix to invoke more professional knowledge for processing, ensuring the accuracy and professionalism of the results.

[0147] If no domain entity exists, the process is downgraded to a general-level model.

[0148] If specialized processing of routing decisions is required, then a lightweight model should be used for processing.

[0149] The cross-modal intelligent agent base deployment method based on the AI ​​large model provided in this embodiment switches the target model based on the judgment results of whether professional processing of routing decisions is required and whether domain entities exist. It selects the target model that matches the multimodal dynamic knowledge graph, improves the adaptability of the model, optimizes resource allocation, and degrades the model when professional processing of routing decisions is not required, thus avoiding computing power redundancy and resource waste.

[0150] This invention also provides a computer device; please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 4 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 4 Take a processor 10 as an example.

[0151] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0152] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0153] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0154] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0155] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.

[0156] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touch screen. Output device 40 may include a display device, etc.

[0157] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0158] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0159] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope of this application.

Claims

1. A cross-modal intelligent agent base deployment system based on an AI large model, characterized in that, The system includes a data layer, a processing layer, and an optimization layer, wherein... The data layer includes an external multi-source heterogeneous database and a multimodal dynamic knowledge graph. The external multi-source heterogeneous database is used to mount data resources and to construct a structured knowledge network based on multimodal data. The external multi-source heterogeneous database and the multimodal dynamic knowledge graph construct a cross-modal knowledge base through unified encoding and joint retrieval. The processing layer includes a multimodal feature parser, an intelligent routing decision engine, and a three-level dynamic model matrix. The multimodal feature parser is used to parse multimodal data from business instructions and the data layer. The three-level dynamic model matrix includes lightweight, general-purpose, and professional-grade AI large models. The intelligent routing decision engine is used to select a target model from the three-level dynamic model matrix to process the input data and generate routing decisions. The optimization layer includes a hierarchical retrieval enhancement module and a dynamic feedback optimization module. The hierarchical retrieval enhancement module is used to optimize the data in the data layer, and the dynamic feedback optimization module is used to push optimization parameters to the data layer and the processing layer to optimize the data layer and the processing layer. The dynamic feedback optimization module adopts a dynamic feedback optimization mechanism consisting of a fast loop, a slow loop, and a validation loop. The fast loop is used to adjust the routing strategy parameters, the slow loop is used to update the knowledge graph embedding, and the validation loop is used to perform incremental training on the model matrix. The optimization layer also deploys an online learning algorithm. When the result output by the processing layer does not meet the confidence threshold requirement of the optimization layer, the online learning algorithm feeds back the full-process record information triggered by the business instruction to the joint retrieval system for re-retrieval until the result output by the processing layer meets the confidence threshold requirement of the optimization layer.

2. A method for deploying a cross-modal intelligent agent base based on a large AI model, characterized in that, The method, applied to a cross-modal intelligent agent deployment system based on a large AI model, includes: Unified encoding and joint retrieval of external multi-source heterogeneous databases and multimodal dynamic knowledge graphs are used to construct a cross-modal knowledge base; A multimodal feature parser is used to parse business instructions and multimodal data to extract feature representations; The intelligent routing decision engine selects the target model from the three-level dynamic model matrix and uses the target model to process the input data to generate routing decisions. Dynamic feedback optimization is performed using a dynamic feedback optimization module. The dynamic feedback optimization module adopts a dynamic feedback optimization mechanism consisting of a fast loop, a slow loop, and a validation loop. The fast loop is used to adjust the routing strategy parameters, the slow loop is used to update the knowledge graph embedding, and the validation loop is used to perform incremental training on the model matrix. When the results output by the processing layer do not meet the confidence threshold requirements of the optimization layer, the full-process record information triggered by the business instruction is fed back to the joint retrieval system for re-retrieval until the results output by the processing layer meet the confidence threshold requirements of the optimization layer.

3. The method according to claim 2, characterized in that, The unified encoding and joint retrieval of external multi-source heterogeneous databases and multimodal dynamic knowledge graphs include: By extracting the logical relationships between data from external multi-source heterogeneous databases using cross-modal association technology, performing unified encoding, and constructing a multimodal dynamic knowledge graph; Similarity retrieval is performed in a multimodal dynamic knowledge graph using the encoding vectors of each component.

4. The method according to claim 2, characterized in that, The process of using a multimodal feature parser to parse business instructions and multimodal data and extract feature representations includes: If the multimodal data types of business instructions and data layers are text, then word vector technology is used to map the text to a low-dimensional vector space to obtain the vector representation of the text; If the multimodal data type of the business instructions and data layer is an image, then the image is used to extract features through a CNN architecture to obtain a high-level semantic feature map of the image, and then the high-level semantic feature map of the image is converted into a feature vector. If the multimodal data types of business instructions and data layers are time-series data, then feature extraction is performed on the time-series data according to the meaning of the data to form a feature vector of the time-series data.

5. The method according to claim 2, characterized in that, The process of selecting a target model from a three-level dynamic model matrix using an intelligent routing decision engine includes: If the problem complexity is simple, then choose a lightweight model as the target model; If the problem complexity is moderate, then a general-purpose model should be selected as the target model. If the problem has a high degree of complexity, then a professional-grade model should be selected as the target model. If the problem is a single-modal problem, then choose a lightweight model or a general-purpose model as the target model; If the problem is multimodal, the target model should be selected based on the requirements of domain specialization, real-time requirements, and accuracy requirements.

6. The method according to claim 2, characterized in that, After processing the input data using the target model to generate routing decisions, the method further includes: Calculate the confidence level of the response output and determine whether the confidence level is less than the preset confidence threshold; If the confidence level is less than the preset confidence threshold, optimization will be carried out through a dynamic feedback optimization mechanism.

7. The method according to claim 2, characterized in that, After processing the input data using the target model to generate routing decisions, the method further includes: The intelligent routing decision engine is used to determine whether professional processing of routing decisions is required. If specialized processing of routing decisions is required, then determine whether there are domain entities in the multimodal dynamic knowledge graph; If domain entities exist in the multimodal dynamic knowledge graph, a professional-grade model is activated, and the retrieval is enhanced. If no domain entities exist in the multimodal dynamic knowledge graph, it is downgraded to a general-level model.

8. The method according to claim 7, characterized in that, The method further includes: If no specialized processing of routing decisions is required, then downgrade to a lightweight model.

9. A computer device, characterized in that, include: The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes the computer instructions to perform the cross-modal intelligent agent base deployment method based on any one of claims 2 to 8.

Citation Information

Patent Citations

  • Archive knowledge base construction and retrieval method and system based on multi-modal data fusion

    CN120407703A

  • Construction method and system of intelligent question-answering system based on planning knowledge graph database

    CN120670543A