Self-adaptive AI agent generation method for accurate calculation of knowledge base
By constructing a multi-source heterogeneous data acquisition interface matrix and a dynamic structural adjustment algorithm for the LSTM-GRU hybrid neural network, the problems of format differences and uneven quality in multi-source heterogeneous data acquisition and preprocessing are solved, efficient semantic alignment and entity disambiguation of heterogeneous ontologies are achieved, the adaptability and computational efficiency of the model are improved, resource allocation is optimized, and the problems of low computational efficiency and low resource utilization in existing technologies are solved.
Patent Information
- Application Number
- CN202510746016.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Existing technologies have problems with large format differences and uneven quality in the collection and preprocessing of multi-source heterogeneous data. The efficiency of heterogeneous ontology alignment in semantic network modeling is low, the accuracy of entity disambiguation and relationship reasoning is limited, the computing framework lacks hierarchical decomposition, resource allocation is unreasonable, and model structure adjustment relies on manual labor, resulting in low computing efficiency and low resource utilization.
Construct a multi-source heterogeneous data acquisition interface matrix, design a dynamic structure adjustment algorithm based on the LSTM-GRU hybrid neural network, use distributed message queues to eliminate data format differences, use the improved TransE model for knowledge embedding, establish a task hierarchical computing framework, and achieve adaptive adjustment through a double-layer attention mechanism and a semantic network modeling engine.
It improves the efficiency and accuracy of data collection and preprocessing, achieves efficient semantic alignment and entity disambiguation of heterogeneous ontologies, enhances the generalization ability and computing efficiency of the model, optimizes resource allocation, and improves overall processing performance.
Smart Images

Figure CN120633857A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for generating an adaptive AI agent for precise calculation of a knowledge base. Background Art
[0002] With the rapid development of information technology, knowledge bases, as the core infrastructure of intelligent applications, play a key role in supporting accurate computing and decision-making in complex scenarios. Accurate computing for knowledge bases requires integrating multi-source heterogeneous data, building semantic association networks, and implementing dynamic knowledge processing and reasoning through efficient algorithms. Existing technologies have made certain progress in data acquisition, knowledge representation, and model optimization. For example, data access is achieved through API interfaces, semantic modeling is performed using knowledge graphs, and complex computing tasks are processed with the help of neural networks. These technologies provide basic capabilities for scenarios such as intelligent question answering, data analysis, and decision support. However, with the explosive growth of data scale and the complexity of knowledge structures, traditional methods have gradually revealed limitations in processing dynamically changing multi-source data, adaptively adjusting model structures, and hierarchical task decomposition.
[0003] In practical applications, the collection and preprocessing of multi-source heterogeneous data face challenges such as large format differences and uneven quality. Existing interface designs struggle to efficiently handle the mixed access of streaming and batch data, and data cleaning and quality assessment mechanisms lack robustness, making it difficult to ensure the accuracy and completeness of knowledge input. In semantic network modeling, semantic alignment of heterogeneous ontologies is inefficient, and the accuracy of entity disambiguation and relationship reasoning is limited by a fixed model structure, making it impossible to dynamically optimize the embedding space based on the scale of knowledge. Furthermore, traditional computing frameworks lack effective support for hierarchical task decomposition, making it difficult to balance resource allocation for tasks of varying complexity. Furthermore, model structure adjustments rely on manual experience and are unable to respond in real time to the dynamic demands of data updates and task changes. These issues lead to challenges for existing methods in handling large-scale, highly dynamic knowledge base precision computation tasks, such as low computational efficiency, insufficient model generalization, and low resource utilization. To address this, we propose an adaptive AI agent generation method for knowledge base precision computation. Summary of the Invention
[0004] In order to solve the above technical problems, an adaptive AI agent generation method for precise calculation of knowledge base is provided. This technical solution solves the above-mentioned problems in multi-source heterogeneous data collection and preprocessing, such as the difficulty of interfaces in efficiently handling mixed access and insufficient robustness of mechanisms such as data cleaning; in semantic network modeling, the efficiency of heterogeneous ontology alignment is low, and the accuracy of entity disambiguation is limited by fixed models; the computing framework lacks hierarchical decomposition, resource allocation is unreasonable, and model structure adjustment relies on manual labor.
[0005] In order to achieve the above objects, the technical solution adopted by the present invention is: A method for generating an adaptive AI agent for accurate knowledge base computation includes the following steps: Establish a multi-source heterogeneous data collection interface matrix, build a semantic network modeling engine including an RDF triple parser, design a dynamic structure adjustment algorithm based on an LSTM-GRU hybrid neural network, and develop a task-level computing framework; The multi-source heterogeneous data acquisition interface matrix consists of an API gateway service cluster, a streaming data ingestion pipeline, and a batch data import channel. It implements data buffering through a distributed message queue and uses a pattern matching algorithm to eliminate data format differences. The semantic network modeling engine includes an ontology parsing module, an entity disambiguation module, and a relationship reasoning module. It uses an improved TransE model for knowledge embedding, and the embedding vector dimension is dynamically adjusted according to the logarithmic value of the total number of entities and the number of relationship types. The dynamic structure adjustment algorithm constructs a two-layer attention mechanism network. The upper layer network generates a complexity index based on the number of task nodes, semantic relevance, and data update frequency, and the lower layer network generates structure adjustment parameters based on the complexity index. The task-level computing framework adopts a semantic-driven workflow engine to decompose computing tasks into three submodules: concept layer computing, entity layer computing, and relation layer computing. Each submodule is equipped with an independent adaptive computing unit.
[0006] Preferably, the multi-source heterogeneous data acquisition interface matrix is specifically implemented as follows: Build a data collection terminal cluster that includes a RESTful API adapter, a WebSocket real-time listener, and an FTP batch downloader. Each terminal completes security authentication through the OAuth2.0 protocol. Establish a data quality assessment model that calculates a quality index by weighted summation of completeness, accuracy, and timeliness scores. When the quality index falls below a threshold, the data recollection process is triggered. Design a data cleaning module based on a hybrid parser of regular expressions and XPath. This module defines data templates through pattern expressions, processes abnormal data with cleaning functions, and calculates data confidence based on a weighted similarity algorithm. Deploy distributed cache middleware, adopt a dynamic elimination mechanism that combines the LRU algorithm and access frequency statistics, and optimize the cache strategy by real-time monitoring of cache hit rate.
[0007] Preferably, the workflow of the semantic network modeling engine includes: An ontology parser is built based on the OWL language. By defining an ontology mapping rule set that includes source concepts, target concepts, mapping relationships, and confidence parameters, the semantic alignment of heterogeneous ontologies is achieved. In the entity disambiguation stage, a hybrid similarity calculation method that integrates the improved edit distance algorithm and the enhanced prefix weight algorithm is adopted, and the contribution of the two algorithms is balanced by dynamically adjusting parameters; For relational reasoning tasks, we first calculate the product of the path weight and the relation reliability index to obtain the path confidence, and then select the path with the highest confidence as the reasoning result; In the process of knowledge embedding, the TransE model with type constraints is adopted to construct the loss function by optimizing the difference between the distance between positive and negative sample vectors, and dynamically adjust the negative sampling weight coefficient based on the frequency of relationship occurrence.
[0008] Preferably, the dynamic structure adjustment algorithm is specifically as follows: A bidirectional LSTM network is constructed to capture temporal features, where the number of hidden layer units is dynamically configured according to the logarithmic values of the input and output dimensions; Design a gated recurrent unit with a composite structure, control the retention ratio of historical states through the update gate, use the reset gate to adjust the candidate state generation process, and finally combine the two to output the optimized state vector; Establish a hierarchical attention mechanism network. The first-level attention layer calculates the weight distribution of task features, and the second-level attention layer generates structural parameter adjustment coefficients based on the feature weights. Based on the above mechanism, a dynamic structure regulator is constructed, which dynamically adjusts the number of hidden layers and the number of neurons in each layer of the neural network according to the complexity index of real-time calculation and the preset upper limit of the network size.
[0009] Preferably, the construction process of the task hierarchical computing framework is: In the concept layer operation module, the implication relationship between concepts is deduced based on the inference rules of description logic, and complex concept calculations are completed through rule chain expansion; The physical layer operation module adopts a graph neural network architecture to construct a normalized adjacency matrix and implements hierarchical propagation of node features through a multi-hop graph convolution layer; The relationship layer operation module designs a three-dimensional tensor decomposition model, combines the core tensor with multiple factor matrices to predict the probability of the existence of potential relationships; The task scheduler generates a priority queue by comprehensively calculating the time sensitivity and criticality of the tasks, and dynamically allocates computing resources based on the linear combination of the logarithm of the task complexity and the number of nodes.
[0010] Preferably, the data cleaning module is enhanced by: Construct a conditional random field model for sequence labeling and define a composite feature function set including word form features, part of speech features, and context window features; The bidirectional LSTM network is cascaded with the conditional random field model. The deep features of the sequence are extracted by LSTM and then input into the CRF layer to calculate the optimal labeling sequence. Develop an incremental rule learning mechanism that triggers the rule optimization process to generate new rule combinations when the applicability index of the cleaning rules is detected to be lower than the set threshold; At the same time, a timestamp-data summary mapping table based on the SHA-256 hash algorithm is established to realize version tracking of data changes and abnormal status rollback functions.
[0011] Preferably, the optimization steps of the knowledge embedding process are: Add a compatibility constraint term between the entity type vector and the concept set to the TransE model loss function, and optimize the embedding space distribution through type matching error; Design an adaptive negative sampling strategy based on relationship frequency, dynamically adjusting the probability distribution of negative sample generation according to the smoothed statistical value of the relationship frequency; Construct a hierarchical embedding space architecture, establish a projection matrix between the concept layer and the entity layer, and realize an interpretable mapping from concept vectors to entity vectors; Finally, a dynamic dimension expansion mechanism is introduced. When it is detected that the entity scale growth exceeds a threshold, the embedding vector dimension is automatically expanded according to a logarithmic function relationship.
[0012] Preferably, the improvement of the attention mechanism includes: The standard attention mechanism is expanded to a multi-head structure, which divides the feature vector according to the preset ratio between the model dimension and the number of attention heads. The sine and cosine functions are used in the position encoding layer to generate relative position vectors, enhancing the model's ability to model sequence position relationships. A hybrid network of content attention and structural attention is constructed, the fusion ratio of the two attention weights is dynamically adjusted through trainable parameters, and a sparse strategy based on a logarithmic function is implemented to only retain key connections with attention weights higher than the dynamic threshold to improve computational efficiency.
[0013] Preferably, the optimization method of the physical layer operation module includes: Design an edge-type-sensitive attention network and calculate the attention weight coefficient by concatenating the node feature vector and the edge type vector; A stratified sampling strategy is adopted to dynamically adjust the neighbor sampling range by combining the node degree centrality index and random walk path length; for newly added entity nodes, the initial embedding vector is generated by aggregating the weighted average of the features of its direct neighbor nodes; Block calculation optimization is implemented in the matrix operation layer, and the matrix block size is dynamically adjusted according to the logarithm of the current number of nodes to improve parallel efficiency.
[0014] Preferably, the overall optimization strategy includes: Construct a composite loss function that includes task loss, structural complexity loss, and L2 regularization, and balance model performance and computational overhead through multi-objective optimization. Implemented a hot restart training mechanism to save the current best model and reset the optimizer parameters when the validation loss does not decrease for three consecutive epochs; Design a multi-granularity knowledge distillation framework to extract knowledge at the concept level, entity level, and relationship level into lightweight models. Develop a flexible resource scheduling algorithm that combines task priority scores and resource utilization decay coefficients to dynamically allocate GPU memory and computing core resources.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The adaptive AI agent generation method proposed in the present invention effectively solves the problems of large differences in data formats and uneven quality by constructing a multi-source heterogeneous data acquisition interface matrix, improves the efficiency and accuracy of data acquisition and preprocessing, and utilizes the semantic network modeling engine to achieve efficient semantic alignment of heterogeneous ontologies, as well as improved accuracy of entity disambiguation and relationship reasoning, providing strong support for dynamic processing and reasoning of knowledge. The dynamic structure adjustment algorithm based on the LSTM-GRU hybrid neural network can adaptively adjust the model structure according to the complexity of the task and the frequency of data update, thereby improving the generalization ability and computing efficiency of the model. The introduction of the task hierarchical computing framework realizes the fine decomposition of computing tasks and the rational allocation of resources, further improving the overall processing performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is a diagram of the key technical nodes of the present invention. DETAILED DESCRIPTION
[0017] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0018] Reference Figure 1 As shown, a method for generating an adaptive AI agent for accurate computation of a knowledge base includes the following steps: Establish a multi-source heterogeneous data collection interface matrix, build a semantic network modeling engine including an RDF triple parser, design a dynamic structure adjustment algorithm based on an LSTM-GRU hybrid neural network, and develop a task hierarchical computing framework. This layered architecture design can effectively cope with data diversity and improve system scalability.
[0019] The multi-source heterogeneous data acquisition interface matrix consists of an API gateway service cluster, a streaming data ingestion pipeline, and a batch data import channel. It implements data buffering through a distributed message queue and uses a pattern matching algorithm to eliminate data format differences. This design ensures data stability and format uniformity in a high-throughput environment.
[0020] The semantic network modeling engine includes an ontology parsing module, an entity disambiguation module, and a relationship reasoning module. It uses an improved TransE model for knowledge embedding. The embedding vector dimension is dynamically adjusted based on the logarithmic value of the total number of entities and the number of relationship types. This dynamic adjustment mechanism effectively controls computational complexity while maintaining the model's expressiveness. The dynamic structure adjustment algorithm constructs a two-layer attention mechanism network. The upper layer of the network generates a complexity index based on the number of task nodes, semantic relevance, and data update frequency. The lower layer of the network generates structure adjustment parameters based on the complexity index. This two-layer mechanism realizes intelligent adaptation of network structure and task characteristics. The task-level computing framework adopts a semantic-driven workflow engine to decompose computing tasks into three sub-modules: conceptual layer operations, entity layer operations, and relational layer operations. Each sub-module is configured with an independent adaptive computing unit. This decoupling design improves the system's parallel processing capabilities and module reusability.
[0021] Reference Figure 2 As shown, the multi-source heterogeneous data acquisition interface matrix is specifically implemented as follows: building a data acquisition terminal cluster including a RESTful API adapter, a WebSocket real-time listener, and an FTP batch downloader, wherein each terminal completes security authentication through the OAuth2.0 protocol. The support of multiple protocols ensures the wide compatibility and access security of data sources; A data quality assessment model was established. This model calculates a quality index by weighted summation of the completeness score, accuracy score, and timeliness score. When the quality index falls below a threshold, a data recollection process is triggered. This quantitative assessment mechanism ensures the reliability of data input. A data cleaning module based on a hybrid parser of regular expressions and XPath was designed. This module defines data templates through pattern expressions, processes abnormal data through cleaning functions, and calculates data confidence based on a weighted similarity algorithm. The hybrid parsing strategy improves the processing accuracy of unstructured data. Deploy distributed cache middleware, employing a dynamic eviction mechanism that combines the LRU algorithm with access frequency statistics. Furthermore, the cache strategy is optimized through real-time monitoring of cache hit rates. This hybrid eviction strategy achieves a better balance between memory utilization and access efficiency. A multi-protocol acquisition terminal cluster achieves broad compatibility and security authentication. A data quality assessment model ensures input reliability, and a hybrid parser improves the accuracy of unstructured data processing. The distributed cache dynamic eviction strategy balances memory and efficiency, enhancing overall system compatibility, reliability, processing accuracy, and operational efficiency.
[0022] The workflow of the semantic network modeling engine includes: building an ontology parser based on the OWL language, achieving semantic alignment of heterogeneous ontologies by defining an ontology mapping rule set including source concepts, target concepts, mapping relationships, and confidence parameters, and standardizing mapping rules to reduce the semantic gap in knowledge fusion; In the entity disambiguation stage, a hybrid similarity calculation method is used that combines an improved edit distance algorithm and an enhanced prefix weight algorithm. By dynamically adjusting parameters to balance the contributions of the two algorithms, this composite algorithm can maintain high accuracy in both short and long text scenarios. For relational reasoning tasks, we first calculate the product of the path weight and the relationship reliability index to obtain the path confidence, and then select the path with the highest confidence as the reasoning result. This path screening mechanism effectively improves the interpretability of the reasoning results. In the knowledge embedding process, the TransE model with type constraints is adopted. The loss function is constructed by optimizing the difference between the distance between positive and negative sample vectors, and the negative sampling weight coefficient is dynamically adjusted according to the frequency of relationship occurrence. This loss function maintains the translation characteristics and strengthens the type constraint. The loss function expression is: ; Where, is the embedding vector of the head entity (HeadEntity), is the embedding vector of the relation, is the embedding vector of the tail entity (TailEntity), Embedding vector for the head entity in the negative sample (generated by replacing the real head entity), Embedding vector for the tail entity in the negative sample (generated by replacing the real tail entity), is the distance function (usually L1 or L2 norm, used to measure the distance between vectors), is the interval hyperparameter (used to control the minimum safe interval between positive and negative samples), is a piecewise function, take and 0 (i.e., the ReLU function), is the regularization parameter (used to balance the weights of the main loss term and the constraint term), EntityTypeVector is the entity type vector, which indicates the concept type to which the entity belongs. is the concept set vector (ConceptSetVector), which represents the semantic features of the target concept; The dynamic structure adjustment algorithm specifically involves: constructing a bidirectional LSTM network to capture time series features, where the number of hidden layer units is dynamically configured based on the logarithmic values of the input and output dimensions. This adaptive configuration avoids insufficient model capacity or excessive redundancy; A gated recurrent unit with a composite structure is designed. The update gate controls the proportion of historical states retained, and the reset gate is used to adjust the candidate state generation process. Finally, the two are combined to output the optimized state vector. This dual gating mechanism enhances the model's ability to model long-range dependencies. A hierarchical attention mechanism network is established. The first-level attention layer calculates the weight distribution of task features, and the second-level attention layer generates structural parameter adjustment coefficients based on the feature weights. The hierarchical attention mechanism realizes the granularity conversion from features to structures. Based on the above mechanism, a dynamic structural regulator is constructed. This regulator dynamically adjusts the number of hidden layers and the number of neurons in each layer of the neural network according to the complexity index of real-time calculation and the preset upper limit of the network scale. This dynamic adjustment capability enables the model to adapt to computing tasks of different scales.
[0023] The construction process of the task-level computation framework is as follows: in the concept-level computation module, the implication relationship between concepts is deduced based on the inference rules of description logic, and complex concept computation is completed through rule chain expansion. The formalized reasoning mechanism ensures the logical rigor of concept computation; The entity layer operation module adopts a graph neural network architecture to construct a normalized adjacency matrix and implements hierarchical propagation of node features through a multi-hop graph convolutional layer. The graph structure modeling effectively captures the topological relationship between entities. The relational layer operation module designs a three-dimensional tensor decomposition model, combining the core tensor with multiple factor matrices to predict the probability of the existence of potential relationships. The tensor decomposition method demonstrates better mathematical expression capabilities in relational prediction tasks. The task scheduler generates a priority queue by comprehensively calculating the time sensitivity and criticality of the tasks, and dynamically allocates computing resources based on a linear combination of the logarithm of the task complexity and the number of nodes. This scheduling strategy optimizes the global utilization of system resources.
[0024] The enhancement method of the data cleaning module is as follows: constructing a conditional random field model for sequence labeling, defining a composite feature function set including word form features, part-of-speech features and context window features, and enriching the feature set to enhance the context perception ability of sequence labeling; cascading a bidirectional LSTM network with a conditional random field model, extracting deep sequence features through LSTM and then inputting them into a CRF layer to calculate the optimal labeling sequence. This cascade architecture combines the advantages of deep learning and probabilistic graph models; developing an incremental rule learning mechanism. When it is detected that the applicability index of a cleaning rule is lower than a set threshold, the rule optimization process is triggered to generate a new rule combination. The adaptive learning mechanism enables the cleaning rule to continuously adapt to changes in data distribution; at the same time, a timestamp-data summary mapping table based on the SHA-256 hash algorithm is established to realize version tracking of data changes and abnormal state rollback functions. The hash summary technology ensures the verifiability and integrity of the data version.
[0025] The optimization steps of the knowledge embedding process are as follows: adding a compatibility constraint term between the entity type vector and the concept set to the TransE model loss function, optimizing the embedding space distribution through type matching error. This constraint enhances the semantic consistency of the embedding vector; designing an adaptive negative sampling strategy based on relationship frequency, dynamically adjusting the probability distribution of negative sample generation according to the smoothed statistical value of the relationship occurrence frequency. This strategy alleviates the impact of the long-tail distribution on model training; constructing a hierarchical embedding space architecture, establishing a projection matrix between the concept layer and the entity layer, and realizing an interpretable mapping from concept vector to entity vector. This hierarchical design supports cross-granularity reasoning of knowledge; finally, introducing a dynamic dimension expansion mechanism. When it is detected that the entity scale growth exceeds a threshold, the embedding vector dimension is automatically expanded according to a logarithmic function relationship. This elastic expansion strategy adapts to the dynamic growth characteristics of the knowledge base.
[0026] The improvements to the attention mechanism include: expanding the standard attention mechanism into a multi-head structure, dividing the feature vector according to a preset ratio between the model dimension and the number of attention heads, using sine-cosine functions in the position encoding layer to generate relative position vectors, enhancing the model's ability to model sequence position relationships, and the multi-head mechanism enables the model to focus on feature patterns in different subspaces in parallel; constructing a hybrid network of content attention and structural attention, dynamically adjusting the fusion ratio of the two attention weights through trainable parameters, and implementing a sparseness strategy based on a logarithmic function, retaining only key connections whose attention weights are higher than a dynamic threshold to improve computational efficiency. The hybrid attention mechanism reduces computational overhead while maintaining the focus of attention.
[0027] The optimization method of the entity layer operation module includes: designing an edge type-sensitive attention network, calculating the attention weight coefficient by splicing the node feature vector and the edge type vector. This design integrates edge semantic information into the attention calculation process; adopting a layered sampling strategy, dynamically adjusting the neighbor sampling range in combination with the node degree centrality index and the random walk path length. The layered strategy balances the capture requirements of local features and global features; for newly added entity nodes, the initial embedding vector is generated by aggregating the weighted average of the features of its direct neighbor nodes. The cold start processing method improves the system's adaptability to new entities; implementing block calculation optimization at the matrix operation layer, dynamically adjusting the matrix block size according to the logarithm of the current number of nodes to improve parallel efficiency. The block strategy fully utilizes the parallel computing characteristics of modern computing hardware.
[0028] The overall optimization strategy of the system includes: constructing a composite loss function that includes task loss term, structural complexity loss term, and L2 regularization term, and balancing model performance and computational overhead through multi-objective optimization. This composite objective function achieves the joint optimization of model effect and efficiency. The composite objective function expression is: ; Where, is the weight parameter (used to balance the importance of different loss terms), is the task loss term (TaskLoss), which measures the prediction error of the model on a specific task. Structure Complexity Loss is the structural complexity loss term, which controls the complexity of the model structure (such as the number of neural network layers, the number of parameters, etc.). L2RegularizationTerm is the L2 regularization term, which is used to prevent overfitting and penalize the sum of the absolute squares of the model parameters.
[0029] A hot restart training mechanism is implemented. When the verification loss does not decrease for three consecutive epochs, the current best model is saved and the optimizer parameters are reset. This mechanism effectively avoids the local optimal trap and accelerates model convergence. A multi-granularity knowledge distillation framework is designed to extract knowledge from the concept layer, entity layer, and relationship layer into lightweight models respectively. The hierarchical distillation strategy retains key semantic information during the model compression process. A flexible resource scheduling algorithm is developed. It combines task priority scores and resource utilization attenuation coefficients to dynamically allocate GPU memory and computing core resources. This scheduling algorithm achieves efficient utilization of hardware resources.
[0030] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for generating an adaptive AI agent for accurate knowledge base calculation, characterized in that: The following steps are involved: Establish a multi-source heterogeneous data collection interface matrix, build a semantic network modeling engine including an RDF triple parser, design a dynamic structure adjustment algorithm based on an LSTM-GRU hybrid neural network, and develop a task-level computing framework; The multi-source heterogeneous data acquisition interface matrix consists of an API gateway service cluster, a streaming data ingestion pipeline, and a batch data import channel. It implements data buffering through a distributed message queue and uses a pattern matching algorithm to eliminate data format differences. The semantic network modeling engine includes an ontology parsing module, an entity disambiguation module, and a relationship reasoning module. It uses an improved TransE model for knowledge embedding, and the embedding vector dimension is dynamically adjusted according to the logarithmic value of the total number of entities and the number of relationship types. The dynamic structure adjustment algorithm constructs a two-layer attention mechanism network. The upper layer network generates a complexity index based on the number of task nodes, semantic relevance, and data update frequency, and the lower layer network generates structure adjustment parameters based on the complexity index. The task-level computing framework adopts a semantic-driven workflow engine to decompose computing tasks into three submodules: concept layer computing, entity layer computing, and relation layer computing. Each submodule is equipped with an independent adaptive computing unit.
2. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 1, characterized in that: The multi-source heterogeneous data acquisition interface matrix is specifically implemented as follows: Build a data collection terminal cluster that includes a RESTful API adapter, a WebSocket real-time listener, and an FTP batch downloader. Each terminal completes security authentication through the OAuth2.0 protocol. Establish a data quality assessment model that calculates a quality index by weighted summation of completeness, accuracy, and timeliness scores. When the quality index falls below a threshold, the data recollection process is triggered. Design a data cleaning module based on a hybrid parser of regular expressions and XPath. This module defines data templates through pattern expressions, processes abnormal data with cleaning functions, and calculates data confidence based on a weighted similarity algorithm. Deploy distributed cache middleware, adopt a dynamic elimination mechanism that combines the LRU algorithm and access frequency statistics, and optimize the cache strategy by real-time monitoring of cache hit rate.
3. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 1, characterized in that: The workflow of the semantic network modeling engine includes: An ontology parser is built based on the OWL language. By defining an ontology mapping rule set that includes source concepts, target concepts, mapping relationships, and confidence parameters, the semantic alignment of heterogeneous ontologies is achieved. In the entity disambiguation stage, a hybrid similarity calculation method that integrates the improved edit distance algorithm and the enhanced prefix weight algorithm is adopted, and the contribution of the two algorithms is balanced by dynamically adjusting parameters; For relational reasoning tasks, we first calculate the product of the path weight and the relation reliability index to obtain the path confidence, and then select the path with the highest confidence as the reasoning result; In the process of knowledge embedding, the TransE model with type constraints is adopted to construct the loss function by optimizing the difference between the distance between positive and negative sample vectors, and dynamically adjust the negative sampling weight coefficient based on the frequency of relationship occurrence.
4. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 1, characterized in that: The dynamic structure adjustment algorithm is specifically as follows: A bidirectional LSTM network is constructed to capture temporal features, where the number of hidden layer units is dynamically configured according to the logarithmic values of the input and output dimensions; Design a gated recurrent unit with a composite structure, control the retention ratio of historical states through the update gate, use the reset gate to adjust the candidate state generation process, and finally combine the two to output the optimized state vector; Establish a hierarchical attention mechanism network. The first-level attention layer calculates the weight distribution of task features, and the second-level attention layer generates structural parameter adjustment coefficients based on the feature weights. Based on the above mechanism, a dynamic structure regulator is constructed, which dynamically adjusts the number of hidden layers and the number of neurons in each layer of the neural network according to the complexity index of real-time calculation and the preset upper limit of the network size.
5. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 1, characterized in that: The construction process of the task hierarchical computing framework is as follows: In the concept layer operation module, the implication relationship between concepts is deduced based on the inference rules of description logic, and complex concept calculations are completed through rule chain expansion; The physical layer operation module adopts a graph neural network architecture to construct a normalized adjacency matrix and implements hierarchical propagation of node features through a multi-hop graph convolution layer; The relationship layer operation module designs a three-dimensional tensor decomposition model, combines the core tensor with multiple factor matrices to predict the probability of the existence of potential relationships; The task scheduler generates a priority queue by comprehensively calculating the time sensitivity and criticality of the tasks, and dynamically allocates computing resources based on the linear combination of the logarithm of the task complexity and the number of nodes.
6. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 2, characterized in that: The enhancement method of the data cleaning module is: Construct a conditional random field model for sequence labeling and define a composite feature function set including word form features, part of speech features, and context window features; The bidirectional LSTM network is cascaded with the conditional random field model. The deep features of the sequence are extracted by LSTM and then input into the CRF layer to calculate the optimal labeling sequence. Develop an incremental rule learning mechanism that triggers the rule optimization process to generate new rule combinations when the applicability index of the cleaning rules is detected to be lower than the set threshold; At the same time, a timestamp-data summary mapping table based on the SHA-256 hash algorithm is established to realize version tracking of data changes and abnormal status rollback functions.
7. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 3, characterized in that: The optimization steps of the knowledge embedding process are: Add a compatibility constraint term between the entity type vector and the concept set to the TransE model loss function, and optimize the embedding space distribution through type matching error; Design an adaptive negative sampling strategy based on relationship frequency, dynamically adjusting the probability distribution of negative sample generation according to the smoothed statistical value of the relationship frequency; Construct a hierarchical embedding space architecture, establish a projection matrix between the concept layer and the entity layer, and realize an interpretable mapping from concept vectors to entity vectors; Finally, a dynamic dimension expansion mechanism is introduced. When it is detected that the entity scale growth exceeds a threshold, the embedding vector dimension is automatically expanded according to a logarithmic function relationship.
8. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 4, characterized in that: Improvements to the attention mechanism include: The standard attention mechanism is expanded to a multi-head structure, which divides the feature vector according to the preset ratio between the model dimension and the number of attention heads. The sine and cosine functions are used in the position encoding layer to generate relative position vectors, enhancing the model's ability to model sequence position relationships. A hybrid network of content attention and structural attention is constructed, the fusion ratio of the two attention weights is dynamically adjusted through trainable parameters, and a sparse strategy based on a logarithmic function is implemented to only retain key connections with attention weights higher than the dynamic threshold to improve computational efficiency.
9. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 5, characterized in that: The optimization method of the physical layer operation module includes: Design an edge-type-sensitive attention network and calculate the attention weight coefficient by concatenating the node feature vector and the edge type vector; A stratified sampling strategy is adopted to dynamically adjust the neighbor sampling range by combining the node degree centrality index and random walk path length; for newly added entity nodes, the initial embedding vector is generated by aggregating the weighted average of the features of its direct neighbor nodes; Block calculation optimization is implemented in the matrix operation layer, and the matrix block size is dynamically adjusted according to the logarithm of the current number of nodes to improve parallel efficiency.
10. The method for generating an adaptive AI agent for accurate knowledge base calculation according to claim 1, characterized in that: The overall optimization strategy includes: Construct a composite loss function that includes task loss, structural complexity loss, and L2 regularization, and balance model performance and computational overhead through multi-objective optimization. Implemented a hot restart training mechanism to save the current best model and reset the optimizer parameters when the validation loss does not decrease for three consecutive epochs; Design a multi-granularity knowledge distillation framework to extract knowledge at the concept level, entity level, and relationship level into lightweight models. Develop a flexible resource scheduling algorithm that combines task priority scores and resource utilization decay coefficients to dynamically allocate GPU memory and computing core resources.
Citation Information
Patent Citations
Text mining method of technology transaction platform
CN116956228A
Dynamic gesture recognition method based on hand key point and double-layer bidirectional LSTM network
CN117576783A
Data collection and drug safety signal mining methods and agents for PMS
CN119786078A
Intelligent data production method and system based on graph neural network and adaptive learning
CN119988647A
Deep neural network optimization system for machine learning model scaling
US20220036194A1
Cited By
Policy field-oriented reasoning generation method, apparatus and device, and storage medium
CN121279459A