Information retrieval intelligent knowledge service electronic equipment
Through BERT+GNN hybrid parsing and AutoML technology, combined with the SWRL rule engine, the reasoning path is dynamically adjusted, which solves the problem of insufficient semantic understanding of traditional information retrieval systems and achieves efficient knowledge services and improved user experience.
Patent Information
- Application Number
- CN202510726043.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional information retrieval intelligent knowledge service systems lack semantic understanding, resulting in low retrieval efficiency, inability to effectively associate deep concepts, and wasting user time.
It uses BERT+GNN hybrid parsing technology, combined with the SWRL rule engine and reinforcement learning, to dynamically adjust the reasoning path according to user domain preferences, and uses AutoML technology to perform multimodal data governance, heterogeneous cleaning and dynamic storage, build an adaptive knowledge service system, and achieve accurate decision-making and efficient reuse of multimodal data pools.
It greatly improves the system's semantic understanding ability and retrieval efficiency, improves user experience and saves user time.
Smart Images

Figure CN120633792A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval intelligent knowledge service, and in particular to an information retrieval intelligent knowledge service electronic device. Background Art
[0002] The information retrieval intelligent knowledge service system was born at the intersection of the explosive growth of data and the upgrading of users' demand for efficient information acquisition. With the exponential expansion of scientific research data, the increasing complexity of business decisions and the increasing requirements for refined social governance, the system relies on the deep integration of artificial intelligence and virtual reality technologies. By constructing a distributed knowledge graph and a multimodal interactive environment, users can complete the accurate capture and correlation analysis of cross-domain knowledge in an immersive experience. Its core value lies in breaking through the limitations of single text retrieval, automatically updating the entity relationship network through dynamic knowledge graphs, supporting complex reasoning tasks, and using intelligent agent technology to automate multi-step retrieval processes, allowing users to deepen their understanding of cross-domain knowledge connections in the process of "learning by using";
[0003] Traditional information retrieval intelligent knowledge service systems suffer from a shallow understanding of semantics, leading to low retrieval efficiency. For example, a graduate student might want to search for "research on traffic flow prediction based on graph neural networks," but traditional systems only match the keywords "graph neural network" and "traffic flow prediction," returning a large number of results containing these two words but with no real connection, such as the application of neural networks in image processing or general traffic flow statistics reports. These systems are unable to understand the dependencies expressed by "based on" and are unable to relate deeper concepts such as "traffic flow prediction" to "spatiotemporal data modeling," significantly reducing retrieval efficiency and wasting user time. Therefore, an electronic information retrieval intelligent knowledge service device is proposed. Summary of the Invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose an information retrieval intelligent knowledge service electronic device.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] An information retrieval intelligent knowledge service electronic device, comprising:
[0007] Data Governance: This layer provides raw data to the cognitive computing layer. By continuously monitoring feedback from the human-computer interaction layer (e.g., declining user satisfaction triggers adjustments to collection strategies), it creates a closed-loop optimization of "data-service-experience." The data governance layer utilizes a three-stage intelligent pipeline consisting of multimodal collection, heterogeneous cleaning, and dynamic storage, combined with AutoML technology for full-link optimization.
[0008] Cognitive computing layer: Responsible for converting raw data into a semantic label system and a dynamic knowledge graph. Through a three-step collaborative mechanism, the cognitive computing layer builds an adaptive knowledge service system. It uses BERT+GNN hybrid parsing, with BERT processing text context and GNN parsing entity relationships in the knowledge graph. Joint representation, combined with the SWRL rule engine and reinforcement learning, dynamically adjusts the reasoning path based on user domain preferences and creates a three-dimensional portrait. It integrates behavioral sequences, cognitive styles, and environmental context to ensure that semantic parsing fits the user scenario. Simultaneously, through a metacognitive feedback loop (user expression / brainwave data triggers model parameter adjustments), the adaptive knowledge service system evolves adaptively, providing a multimodal data pool for the service orchestration layer to provide a precise decision-making basis.
[0009] Service Orchestration Layer: This layer is responsible for extracting data from multimodal data pools, solidifying and orchestrating complex query processes (such as historical conversation analysis, multi-source search, and result fusion) to achieve logical decoupling and efficient reuse. It also provides a visual interface to enable non-technical personnel to quickly configure new scenarios (such as in-depth search of academic literature), builds a monitoring system to track node performance in real time to optimize the system, and outputs structured answers and visual maps to the human-computer interaction layer.
[0010] Human-computer interaction layer: Through the structured answers and visual maps output by the service orchestration layer, a natural interaction bridge is built (such as proactive clarification of needs after ambiguous voice queries), personalized interface adaptation is performed based on historical conversations (such as pre-positioning the document entry for high-frequency academic users), and a user feedback closed loop is established (instant correction and simultaneous optimization of incorrect answers are marked), comprehensively improving the naturalness, intelligence and service accuracy of system interactions. At the same time, user behavior data and user usage feedback information are transmitted to the cognitive computing layer.
[0011] The above technical solution further includes:
[0012] Furthermore, the data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The multimodal acquisition includes the following steps:
[0013] By integrating heterogeneous data from multiple sources, including text, images, audio, and video, and leveraging AutoML technology for intelligent perception and dynamic decision-making, AutoML first trains a collection strategy network using reinforcement learning algorithms. This network can analyze user behavior logs and environmental context in real time, dynamically adjusting the weights of different modal data collections. For example, during public health emergencies, AutoML prioritizes medical images and social media sentiment. Secondly, AutoML introduces time series analysis and machine learning algorithms to dynamically adjust multimodal data in real time, adapting collection strategies to changing scenarios.
[0014] Time series feature extraction:
[0015] Fourier transform, extract frequency domain features, and identify periodic patterns:
[0016]
[0017] Wavelet packet decomposition, multi-resolution analysis, capture local features:
[0018]
[0019] Statistical features, calculate time series statistics;
[0020] Machine Learning Model Training:
[0021] LSTM encoding layer, using bidirectional LSTM to capture temporal dependencies:
[0022] h t =LSTM(x t ,h t-1 );
[0023] Calculate modality attention weights: where e i =W a h i
[0024]
[0025] Output the acquisition weight of each modality:
[0026] Real-time dynamic adjustment;
[0027] Online reasoning:
[0028] Receive new data every second and update the feature vector X new ;
[0029] Weight calculation:
[0030] Calculate the new weight W through the trained model new ;
[0031] Smooth transition:
[0032] Use exponential smoothing to avoid sudden changes:
[0033] W current =λW new +(1-λ)W previous ;
[0034] Finally, AutoML deploys online learning modules to establish a closed-loop feedback mechanism, monitors user satisfaction indicators in real time, and automatically triggers policy network retraining when service quality degrades, forming an intelligent closed loop of "collection-feedback-optimization."
[0035] Furthermore, the data governance layer uses a three-stage intelligent pipeline of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The heterogeneous cleaning includes the following steps:
[0036] By integrating data from different sources, formats, and structures, AutoML technology is used for intelligent cleaning and dynamic optimization. AutoML first uses a clustering analysis algorithm to dynamically analyze redundant, abnormal, and missing data in the multi-source heterogeneous data fusion table and automatically adjust the cleaning strategy:
[0037] Where: k is the preset number of clusters, Ci is the set of data points of the i-th cluster, μ i is the centroid (mean vector) of the i-th cluster, ||x-μ i || is the distance from data point x to centroid, μ i The Euclidean distance of
[0038] First randomly select k centroids, and then assign each data point to the cluster corresponding to the nearest centroid:
[0039] Finally, recalculate the centroid of each cluster and repeat until the centroid no longer changes or the maximum number of iterations is reached;
[0040] For example, in the healthcare field, when faced with unstructured medical records from different hospitals, AutoML uses a generative adversarial network (GAN) to simulate missing fields and, combined with a Transformer model, repairs incomplete data, significantly improving data integrity. AutoML also deploys an online learning module to monitor cleaning results in real time. If the noise recognition rate drops, it automatically triggers retraining of the cleaning strategy network. This closed-loop "analysis-cleaning-feedback" mechanism enables the continuous evolution of cleaning strategies.
[0041] Furthermore, the data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The dynamic storage includes the following steps:
[0042] Efficient resource allocation is achieved through intelligent prediction and closed-loop optimization. AutoML uses LSTM neural networks to predict data access frequency, build a popularity grading model, and dynamically adjust storage media allocation strategies.
[0043] Extract time series features from historical access records, including:
[0044] Time decay factor: γ t (γ=0.95, t is the time interval from the current time);
[0045] Frequency of visits: t: time window i: number of visits;
[0046] Access interval entropy: H(t) = -∑p(t)logp((t)) (reflects the randomness of the access pattern);
[0047] Convert the data into a supervised learning sequence: X = [f i,t-n ,...,f i,t -1]Y=f i,t ;
[0048] The LSTM model construction design includes a network with 2 layers of LSTM (128 units) and a fully connected layer;
[0049] Using mean square error (MSE) loss:
[0050] Based on the predicted frequency of visits Calculate the heat score:
[0051] in is the mean and standard deviation of the predicted frequencies;
[0052] Establish a three-level heat system:
[0053] Cold data: S<1; Warm data: 1≤S<3; Hot data: S≥3;
[0054] Design a joint optimization objective that includes access latency cost and storage cost:
[0055] Cost = α·Z + β·V, where Z is the access delay and V is the storage cost
[0056] Hot data: prioritize SSDs (accounting for >70%);
[0057] Warm data: Mixed allocation of HDD and SSD (ratio 4:6);
[0058] Cold data: compressed and stored in HDD (compression ratio > 65%);
[0059] For example, in financial trading scenarios, frequently accessed order flow data is identified in real time and automatically migrated to the SSD storage tier; cold backup data is compressed and stored on low-cost HDDs. Furthermore, AutoML uses reinforcement learning to train the storage policy network, dynamically adjusting the cache replacement algorithm based on real-time load. This closed-loop feedback mechanism continuously monitors storage efficiency indicators. When a drop in the prediction accuracy of hot data is detected, the policy network is automatically retrained, forming an intelligent closed-loop "prediction-adjustment-optimization" system that significantly improves the storage system's adaptability.
[0060] Furthermore, the hybrid parsing using BERT+GNN, where BERT processes the text context and GNN parses the entity relationships in the knowledge graph, and the joint representation includes the following steps:
[0061] BERT generates semantic vectors through text context generated by multi-layer Transformer encoder.
[0062] The text sequence X = [x 1 ,x 2 ,...,x n ]Converted to word embedding vector E token ∈R n×d , and add position code E pos and segment code E seg , get the initial input:
[0063] H0=E token +E pos +E seg ;
[0064] At each layer l of the Transformer, self-attention is calculated to capture long-range dependencies:
[0065] Q=H l-1 W Q K=H l-1 W K V=H l-1 W V
[0066] Among them, W V , W Q , W K ,∈R d×dk is a trainable parameter;
[0067] Perform a nonlinear transformation on the self-attention output:
[0068] FFN(x)=ReLU(xW1+b1)W2+b2;
[0069] The final output is:
[0070] H l =LayerNorm(H l-1 +FFN(LayerNorm(H l-1 +Attention(Q,K,V))))
[0071] Repeat the above steps L times to get the final text semantic vector H BERT ∈R n×d ;
[0072] GNN parses the entity-relationship output mapping in the knowledge graph and initializes the entities and relationships in the knowledge graph into low-dimensional vectors:
[0073]
[0074] At each layer l, node v aggregates neighbor information:
[0075] Where ⊕ is vector splicing;
[0076] Combine the current node features with the aggregate information to update the representation:
[0077] Where σ is the activation function;
[0078] Repeat the message passing and updating L times to get the final mapping H GNN ∈R ∣V∣×d ;
[0079] Then, bimodal fusion is performed, and the semantic vector and the mapping are input into the Transformer encoder to generate a joint representation vector.
[0080] Furthermore, the method combines the SWRL rule engine and reinforcement learning to dynamically adjust the reasoning path according to the user's domain preference, including the following steps:
[0081] The domain preference model is constructed by user behavior sequence, cognitive style and environmental context; let the user behavior sequence be S=[s 1 ,s 2 ,...,s m ]The cognitive style vector is c∈R k The environmental context vector is e∈R p Then the domain preference vector u∈R d It can be expressed as u=W S LSTM(S)+W C c+W E e, where W S , W C , W E is the trainable weight matrix;
[0082] The SWRL rule engine matches relevant rules according to the user domain preference u; let the entity relationship set in the knowledge graph be R, and the SWRL rule base be R={r 1 ,r 2 ,...,r n};
[0083] Each rule ri contains a premise and a conclusion. By calculating the similarity between the rule premise and the user preference: sim(r i ,u)=cosine(hri ,u) select the one with the highest similarity;
[0084] Matching-based rules R u , dynamically adjust the reasoning path in the knowledge graph; let the current query entity be v q , the target entity is vt, and the reasoning path P=[v q ,v1,...,v t ];
[0085] The transition probability is defined as: where r vi→vi+1 To connect v i and v i+1 The relationship between N(v i ) is the neighbor node of vi.
[0086] Furthermore, the solidified and orchestrated complex query process is logically decoupled and efficiently reused, including the following steps:
[0087] First, a visual process model is built based on the BPMN 2.0 standard, which breaks down complex queries into atomic service units of entity linking and relationship reasoning, and defines standardized input and output interfaces.
[0088] The Camunda process engine is then deployed to manage the lifecycle of process instances, and asynchronous communication between services is performed through the Kafka message queue, so that each service can be deployed and expanded independently.
[0089] Next, we established a process template library that includes parameterized configuration and version management. We then combined TF-IDF / BERT similarity calculation with a weighted routing algorithm based on historical success rates to intelligently distribute query requests.
[0090] Route(q)=argmax r (Sim(q,r desc )·Conf(r)), where Sim is the similarity between the query and the route description; Conf is the route confidence;
[0091] Ultimately, through dynamic scaling of service replicas and versioned deployment mechanisms, we can improve process reuse efficiency while performing logical decoupling.
[0092] Furthermore, the construction of a monitoring system to track node performance optimization system in real time includes the following steps:
[0093] Prometheus is used to collect key indicators such as CPU, memory, and response time, and an LSTM neural network is used to predict future trends based on historical load data: Among them L t is the node load at the current moment; n is the time window length;
[0094] At the same time, dynamic threshold algorithm is combined for anomaly detection: θ dynamic =μ history +k·σ history , where k is the adjustment coefficient.
[0095] Furthermore, the naturalness, intelligence, and service accuracy of system interactions are comprehensively improved by building a natural interaction bridge, performing personalized interface adaptation based on historical conversations, and establishing a user feedback closed loop, including the following steps:
[0096] Building a natural interaction bridge: By integrating speech recognition, speech synthesis, natural language processing, and gesture recognition technologies, combined with knowledge graph visualization, a multimodal interactive interface is created that supports voice question and answer, multi-round dialogue, and intuitive information exploration, forming a natural communication bridge.
[0097] Personalized interface adaptation (taking academic users as an example):
[0098] We built user profiles by analyzing domain keywords in historical conversations (e.g., "literature review" and "citation rate" appearing >5 times per week). We combined behavioral logs (database access duration and download history) to establish an academic preference model. Using front-end configuration tools, we moved the literature database entry from the third-level menu to the homepage quick entry area, dynamically adjusted the default search bar suggestions (e.g., "core journals" and "impact factor query"), and displayed the new interface to 10% of users. We analyzed the results using click heat maps and dwell time, and iteratively optimized the layout.
[0099] Establish a closed loop of user feedback:
[0100] Incorrect annotation identification: An "Incorrect" button is integrated into the result card. Clicking it triggers an annotation event, recording the user ID, session ID, and incorrect answer version number. The correction process then begins, pushing the annotated data to the correction queue for distribution to domain experts or initiating AI recalculation (e.g., invoking stricter knowledge graph validation rules). Service orchestration optimization synchronizes the correction results to the service orchestration layer, updates the answer priority for that scenario, and triggers parameter tuning for related service nodes (e.g., literature search API).
[0101] System iteration and upgrade: Interaction data sedimentation, storing each clarification dialogue, interface adjustment, and correction record in the interaction knowledge base, annotating timestamps and user characteristics, regularly (such as weekly) using new data to train the interaction model, optimize the clarification strategy generation algorithm and personalized recommendation weights; grayscale release verification, pushing the new version of the interaction logic to users in batches, monitoring key indicators (clarification rate, interface element usage rate, error labeling reduction rate) to evaluate the effect.
[0102] The present invention has the following beneficial effects:
[0103] In this invention, BERT+GNN hybrid analysis is used, BERT processes the text context, GNN parses the entity relationship in the knowledge graph, and jointly represents it. By combining the SWRL rule engine and reinforcement learning, the reasoning path is dynamically adjusted according to the user's domain preferences, and a three-dimensional portrait is added. The behavior sequence, cognitive style, and environmental context are integrated to make the semantic analysis fit the user scenario, greatly improving the system's semantic understanding ability, enhancing retrieval efficiency, saving users' time, and making the user experience better. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] Figure 1 This is a system block diagram of an information retrieval intelligent knowledge service electronic device proposed by the present invention. DETAILED DESCRIPTION
[0105] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0106] See also Figure 1 As shown, the present invention is an information retrieval intelligent knowledge service electronic device, comprising:
[0107] Data Governance: This layer provides raw data to the cognitive computing layer. By continuously monitoring feedback from the human-computer interaction layer (e.g., declining user satisfaction triggers adjustments to collection strategies), it creates a closed-loop optimization of "data-service-experience." The data governance layer utilizes a three-stage intelligent pipeline consisting of multimodal collection, heterogeneous cleaning, and dynamic storage, combined with AutoML technology for full-link optimization.
[0108] Cognitive computing layer: Responsible for converting raw data into a semantic label system and a dynamic knowledge graph. Through a three-step collaborative mechanism, the cognitive computing layer builds an adaptive knowledge service system. It uses BERT+GNN hybrid parsing, with BERT processing text context and GNN parsing entity relationships in the knowledge graph. Joint representation, combined with the SWRL rule engine and reinforcement learning, dynamically adjusts the reasoning path based on user domain preferences and creates a three-dimensional portrait. It integrates behavioral sequences, cognitive styles, and environmental context to ensure that semantic parsing fits the user scenario. Simultaneously, through a metacognitive feedback loop (user expression / brainwave data triggers model parameter adjustments), the adaptive knowledge service system evolves adaptively, providing a multimodal data pool for the service orchestration layer to provide a precise decision-making basis.
[0109] Service Orchestration Layer: This layer is responsible for extracting data from multimodal data pools, solidifying and orchestrating complex query processes (such as historical conversation analysis, multi-source search, and result fusion) to achieve logical decoupling and efficient reuse. It also provides a visual interface to enable non-technical personnel to quickly configure new scenarios (such as in-depth search of academic literature), builds a monitoring system to track node performance in real time to optimize the system, and outputs structured answers and visual maps to the human-computer interaction layer.
[0110] Human-computer interaction layer: Through the structured answers and visual maps output by the service orchestration layer, a natural interaction bridge is built (such as proactive clarification of needs after ambiguous voice queries), personalized interface adaptation is performed based on historical conversations (such as pre-positioning the document entry for high-frequency academic users), and a user feedback closed loop is established (instant correction and simultaneous optimization of incorrect answers are marked), comprehensively improving the naturalness, intelligence and service accuracy of system interactions. At the same time, user behavior data and user usage feedback information are transmitted to the cognitive computing layer.
[0111] In one embodiment, the data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The multimodal acquisition includes the following steps:
[0112] By integrating heterogeneous data from multiple sources, including text, images, audio, and video, and leveraging AutoML technology for intelligent perception and dynamic decision-making, AutoML first trains a collection strategy network using reinforcement learning algorithms. This network can analyze user behavior logs and environmental context in real time, dynamically adjusting the weights of different modal data collections. For example, during public health emergencies, AutoML prioritizes medical images and social media sentiment. Secondly, AutoML introduces time series analysis and machine learning algorithms to dynamically adjust multimodal data in real time, adapting collection strategies to changing scenarios.
[0113] Time series feature extraction:
[0114] Fourier transform, extract frequency domain features, and identify periodic patterns:
[0115]
[0116] Wavelet packet decomposition, multi-resolution analysis, capture local features:
[0117]
[0118] Statistical features, calculate time series statistics;
[0119] Machine Learning Model Training:
[0120] LSTM encoding layer, using bidirectional LSTM to capture temporal dependencies:
[0121] ht =LSTM(x t ,h t-1 );
[0122] Calculate modality attention weights: where e i =W a h i ;
[0123]
[0124] Output the acquisition weight of each modality:
[0125] Real-time dynamic adjustment;
[0126] Online reasoning:
[0127] Receive new data every second and update the feature vector X new ;
[0128] Weight calculation:
[0129] Calculate the new weight W through the trained model new ;
[0130] Smooth transition:
[0131] Use exponential smoothing to avoid sudden changes:
[0132] W current =λW new +(1-λ)W previous ;
[0133] Finally, AutoML deploys online learning modules to establish a closed-loop feedback mechanism, monitors user satisfaction indicators in real time, and automatically triggers policy network retraining when service quality degrades, forming an intelligent closed loop of "collection-feedback-optimization."
[0134] In one embodiment, the data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The heterogeneous cleaning process includes the following steps:
[0135] By integrating data from different sources, formats, and structures, AutoML technology is used for intelligent cleaning and dynamic optimization. AutoML first uses a clustering analysis algorithm to dynamically analyze redundant, abnormal, and missing data in the multi-source heterogeneous data fusion table and automatically adjust the cleaning strategy.
[0136] Where: k is the preset number of clusters, Ci is the set of data points of the i-th cluster, μ i is the centroid (mean vector) of the i-th cluster, ||x-μi || is the distance from data point x to centroid, μ i The Euclidean distance of
[0137] First randomly select k centroids, and then assign each data point to the cluster corresponding to the nearest centroid
[0138]
[0139] Finally, recalculate the centroid of each cluster and repeat until the centroid no longer changes or the maximum number of iterations is reached;
[0140] For example, in the healthcare field, when faced with unstructured medical records from different hospitals, AutoML uses a generative adversarial network (GAN) to simulate missing fields and, combined with a Transformer model, repairs incomplete data, significantly improving data integrity. AutoML also deploys an online learning module to monitor cleaning results in real time. If the noise recognition rate drops, it automatically triggers retraining of the cleaning strategy network. This closed-loop "analysis-cleaning-feedback" mechanism enables the continuous evolution of cleaning strategies.
[0141] In one embodiment, the data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The dynamic storage includes the following steps:
[0142] Efficient resource allocation through intelligent prediction and closed-loop optimization. AutoML uses LSTM neural networks to predict data access frequency, build a popularity classification model, and dynamically adjust storage media allocation strategies:
[0143] Extract time series features from historical access records, including:
[0144] Time decay factor: γ t (γ=0.95, t is the time interval from the current time);
[0145] Frequency of visits: t: time window, i: number of visits;
[0146] Access interval entropy: H(t) = -∑p(t)logp((t)) (reflects the randomness of the access pattern);
[0147] Convert the data into a supervised learning sequence: X = [f i,t-n ,...,f i,t -1]Y=f i,t ;
[0148] The LSTM model construction design includes a network with 2 layers of LSTM (128 units) and a fully connected layer;
[0149] Using mean square error (MSE) loss:
[0150] Based on the predicted frequency of visits Calculate the heat score:
[0151] in is the mean and standard deviation of the predicted frequencies
[0152] Establish a three-level heat system:
[0153] Cold data: S<1; Warm data: 1≤S<3; Hot data: S≥3;
[0154] Design a joint optimization objective that includes access latency cost and storage cost:
[0155] Cost = α·Z + β·V, where Z is the access latency and V is the storage cost;
[0156] Hot data: prioritize SSDs (accounting for >70%);
[0157] Warm data: Mixed allocation of HDD and SSD (ratio 4:6);
[0158] Cold data: compressed and stored in HDD (compression ratio > 65%);
[0159] For example, in financial trading scenarios, frequently accessed order flow data is identified in real time and automatically migrated to the SSD storage tier; cold backup data is compressed and stored on low-cost HDDs. Furthermore, AutoML uses reinforcement learning to train the storage policy network, dynamically adjusting the cache replacement algorithm based on real-time load. This closed-loop feedback mechanism continuously monitors storage efficiency indicators. When a drop in the prediction accuracy of hot data is detected, the policy network is automatically retrained, forming an intelligent closed-loop "prediction-adjustment-optimization" system that significantly improves the storage system's adaptability.
[0160] In one embodiment, the hybrid parsing using BERT+GNN, where BERT processes the text context and GNN parses the entity relationships in the knowledge graph, and the joint representation includes the following steps:
[0161] BERT generates a text context and a semantic vector through a multi-layer Transformer encoder, converting the text sequence X=[x 1 ,x 2 ,...,x n ]Converted to word embedding vector E token ∈R n×d , and add position code E pos and segment code E seg , get the initial input:
[0162] H0=E token +E pos +E seg ;
[0163] At each layer l of the Transformer, self-attention is calculated to capture long-range dependencies:
[0164] Q=H l-1 W Q K=H l-1 W K V=H l-1 W V
[0165] Among them, W V , W Q , W K ,∈R d×dk is a trainable parameter;
[0166] Perform a nonlinear transformation on the self-attention output:
[0167] FFN(x)=ReLU(xW1+b 1 )W 2 +b 2 ;
[0168] The final output is:
[0169] H l =LayerNorm(H l-1 +FFN(LayerNorm(H l-1 +Attention(Q,K,V))));
[0170] Repeat the above steps L times to get the final text semantic vector H BERT ∈R n×d ;
[0171] GNN parses the entity-relationship output mapping in the knowledge graph and initializes the entities and relationships in the knowledge graph into low-dimensional vectors:
[0172]
[0173] At each layer l, node v aggregates neighbor information:
[0174] Where ⊕ is vector concatenation
[0175] Combine the current node features with the aggregate information to update the representation:
[0176] Where σ is the activation function;
[0177] Repeat the message passing and updating L times to get the final mapping H GNN ∈R ∣V∣×d ;
[0178] Then, bimodal fusion is performed, and the semantic vector and the mapping are input into the Transformer encoder to generate a joint representation vector.
[0179] In one embodiment, the method of combining the SWRL rule engine and reinforcement learning to dynamically adjust the reasoning path based on user domain preferences includes the following steps:
[0180] The domain preference model is constructed by user behavior sequence, cognitive style and environmental context; let the user behavior sequence be S=[s 1 ,s 2 ,...,s m ]The cognitive style vector is c∈R k The environmental context vector is e∈R p Then the domain preference vector u∈R d It can be expressed as u=W S LSTM(S)+W C c+W E e, where W S , W C , W E is the trainable weight matrix;
[0181] The SWRL rule engine matches relevant rules according to the user domain preference u. Let the entity relationship set in the knowledge graph be R, and the SWRL rule base be R = {r 1 ,r 2 ,...,r n};
[0182] Each rule ri contains a premise and a conclusion. By calculating the similarity between the rule premise and the user preference: sim(r i ,u)=cosine(h ri ,u) select the one with the highest similarity;
[0183] Matching-based rules R u , dynamically adjust the reasoning path in the knowledge graph; let the current query entity be v q , the target entity is vt, and the reasoning path P = [vq, v1, ..., v t ];
[0184] The transition probability is defined as: where r vi→vi+1 To connect v i and v i+1 The relationship between N(vi ) is the neighbor node of vi.
[0185] In one embodiment, the solidified and orchestrated complex query process is logically decoupled and efficiently reused, including the following steps:
[0186] First, a visual process model is built based on the BPMN 2.0 standard, which breaks down complex queries into atomic service units of entity linking and relationship reasoning, and defines standardized input and output interfaces.
[0187] The Camunda process engine is then deployed to manage the lifecycle of process instances, and asynchronous communication between services is performed through the Kafka message queue, so that each service can be deployed and expanded independently.
[0188] Next, we established a process template library that includes parameterized configuration and version management. We then combined TF-IDF / BERT similarity calculation with a weighted routing algorithm based on historical success rates to intelligently distribute query requests.
[0189] Route(q)=argmax r (Sim(q,r desc )·Conf(r)), where Sim is the similarity between the query and the route description; Conf is the route confidence;
[0190] Ultimately, through dynamic scaling of service replicas and versioned deployment mechanisms, we can improve process reuse efficiency while performing logical decoupling.
[0191] In one embodiment, the construction of a monitoring system for real-time tracking of node performance optimization system includes the following steps:
[0192] Prometheus is used to collect key indicators such as CPU, memory, and response time, and an LSTM neural network is used to predict future trends based on historical load data: Among them L t is the node load at the current moment; n is the time window length;
[0193] At the same time, dynamic threshold algorithm is combined for anomaly detection: θ dynamic =μ history +k·σ history , where k is the adjustment coefficient.
[0194] In one embodiment, the process of comprehensively improving the naturalness, intelligence, and service accuracy of system interactions by building a natural interaction bridge, performing personalized interface adaptation based on historical conversations, and establishing a closed loop of user feedback includes the following steps:
[0195] Building a natural interaction bridge: By integrating speech recognition, speech synthesis, natural language processing, and gesture recognition technologies, combined with knowledge graph visualization, a multimodal interactive interface is created that supports voice question and answer, multi-round dialogue, and intuitive information exploration, forming a natural communication bridge.
[0196] Personalized interface adaptation (taking academic users as an example):
[0197] We built user profiles by analyzing domain keywords in historical conversations (e.g., "literature review" and "citation rate" appearing >5 times per week). We combined behavioral logs (database access duration and download history) to establish an academic preference model. Using front-end configuration tools, we moved the literature database entry from the third-level menu to the homepage quick entry area, dynamically adjusted the default search bar suggestions (e.g., "core journals" and "impact factor query"), and displayed the new interface to 10% of users. We analyzed the results using click heat maps and dwell time, and iteratively optimized the layout.
[0198] Establish a closed loop of user feedback:
[0199] Incorrect annotation identification: An "Incorrect" button is integrated into the result card. Clicking it triggers an annotation event, recording the user ID, session ID, and incorrect answer version number. The correction process then begins, pushing the annotated data to the correction queue for distribution to domain experts or initiating AI recalculation (e.g., invoking stricter knowledge graph validation rules). Service orchestration optimization synchronizes the correction results to the service orchestration layer, updates the answer priority for that scenario, and triggers parameter tuning for related service nodes (e.g., literature search API).
[0200] System iteration and upgrade: Interaction data sedimentation, storing each clarification dialogue, interface adjustment, and correction record in the interaction knowledge base, annotating timestamps and user characteristics, regularly (such as weekly) using new data to train the interaction model, optimize the clarification strategy generation algorithm and personalized recommendation weights; grayscale release verification, pushing the new version of the interaction logic to users in batches, monitoring key indicators (clarification rate, interface element usage rate, error labeling reduction rate) to evaluate the effect.
[0201] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An information retrieval intelligent knowledge service electronic device, characterized in that: include: Data Governance Layer: This layer provides raw data to the cognitive computing layer and performs closed-loop optimization by continuously monitoring feedback from the human-computer interaction layer. The data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology, for full-link optimization. Cognitive computing layer: Responsible for converting raw data into a semantic label system and a dynamic knowledge graph. Through a three-step collaborative mechanism, the cognitive computing layer builds an adaptive knowledge service system. It uses BERT+GNN hybrid parsing, with BERT processing text context and GNN parsing entity relationships in the knowledge graph. Joint representation, combined with the SWRL rule engine and reinforcement learning, dynamically adjusts the reasoning path based on user domain preferences and creates a three-dimensional portrait. It integrates behavioral sequences, cognitive styles, and environmental context to ensure that semantic parsing is tailored to user scenarios. Simultaneously, a metacognitive feedback loop enables the adaptive evolution of the adaptive knowledge service system, providing a multimodal data pool to inform decision-making at the service orchestration layer. Service Orchestration Layer: This layer is responsible for extracting data from multimodal data pools, solidifying and orchestrating complex query flows for logical decoupling and efficient reuse, providing a visual interface to support non-technical personnel in configuring new scenarios, building a monitoring system to track node performance optimization in real time, and outputting structured answers and visual graphs to the human-computer interaction layer. Human-computer interaction layer: Through the structured answers and visual maps output by the service orchestration layer, a natural interaction bridge is built, personalized interface adaptation is performed based on historical conversations, and a user feedback loop is established. At the same time, user behavior data and user usage feedback information are fed back to the cognitive computing layer.
2. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The multimodal acquisition includes the following steps: By integrating heterogeneous data from multiple sources, including text, images, audio, and video, and leveraging AutoML technology for intelligent perception and dynamic decision-making, AutoML first trains the collection strategy network through reinforcement learning algorithms, analyzing user behavior logs and environmental context in real time to dynamically adjust the collection weights for different modal data. Secondly, AutoML introduces time series analysis and machine learning algorithms to dynamically adjust multimodal data in real time, allowing collection strategies to adapt to changing scenarios. Time series feature extraction: Fourier transform, extract frequency domain features, and identify periodic patterns: Wavelet packet decomposition, multi-resolution analysis, capture local features Statistical features, calculate time series statistics; Machine Learning Model Training: LSTM encoding layer, using bidirectional LSTM to capture temporal dependencies: h t =LSTM(x t ,h t-1 ); Calculate modality attention weights: where e i =W a h i , Output the acquisition weights of each modality in real-time dynamic adjustment; Online reasoning: Receive new data every second and update the feature vector X new ; Weight calculation: Calculate the new weight W through the trained model new ; Smooth transition: Use exponential smoothing to avoid sudden changes: IN current =λW new +(1-λ)W previous ; Finally, AutoML deploys online learning modules, establishes a closed-loop feedback mechanism, monitors user satisfaction indicators in real time, and automatically triggers policy network retraining when service quality deteriorates.
3. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The heterogeneous cleaning process includes the following steps: By integrating data from different sources, formats, and structures, AutoML technology is used for intelligent cleaning and dynamic optimization. AutoML first uses a clustering analysis algorithm to dynamically analyze redundant, abnormal, and missing data in the multi-source heterogeneous data fusion table and automatically adjust the cleaning strategy: Where: k is the preset number of clusters, Ci is the set of data points of the i-th cluster, μ i is the centroid (mean vector) of the i-th cluster, ||x-μ i || is the distance from data point x to centroid, μ i The Euclidean distance of First randomly select k centroids, and then assign each data point to the cluster corresponding to the nearest centroid: Finally, recalculate the centroid of each cluster and repeat until the centroid no longer changes or the maximum number of iterations is reached; At the same time, AutoML deploys online learning modules to monitor cleaning results in real time.
4. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The data governance layer uses a three-stage intelligent pipeline consisting of multimodal acquisition, heterogeneous cleaning, and dynamic storage, combined with AutoML technology to perform full-link optimization. The dynamic storage includes the following steps: Through intelligent prediction and closed-loop optimization, AutoML achieves efficient resource allocation. It uses LSTM neural networks to predict data access frequency, build a popularity grading model, and dynamically adjust storage media allocation strategies. Extract time series features from historical access records, including: Time decay factor: γ t Frequency of visits: t: time window i: number of visits, Entropy of visit interval: H(t) = -∑p(t)logp(t); Convert the data into a supervised learning sequence: X = [f i,t-n ,...,f i,t-1 ]Y=f i,t ; The LSTM model construction design includes a network with 2 LSTM layers and a fully connected layer; Using mean square error loss: Based on the predicted frequency of visits Calculate the heat score: in is the mean and standard deviation of the predicted frequencies Establish a three-level heat system: Cold data: S<1; Temperature data: 1≤S<3; Thermal data: S ≥ 3; Design a joint optimization objective that includes access latency cost and storage cost: Cost = α·Z + β·V, where Z is the access delay and V is the storage cost Hot data: prioritize SSD allocation; Warm data: Mixed allocation of HDD+SSD; Cold data: compressed and stored in HDD; At the same time, the storage policy network is trained through reinforcement learning, and the cache replacement algorithm is dynamically adjusted according to the real-time load; AutoML establishes a closed-loop feedback mechanism to continuously monitor storage efficiency indicators, and retrains the policy network when it detects a decrease in the prediction accuracy of hot data.
5. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The BERT+GNN hybrid parsing method uses BERT to process text context and GNN to parse entity relationships in the knowledge graph. The joint representation includes the following steps: BERT generates semantic vectors through text context generated by multi-layer Transformer encoder. The text sequence X = [x 1 ,x 2 ,...,x n ]Converted to word embedding vector E token ∈R n×d , and add position code E pos and segment code E seg , get the initial input: H0=E token +E pos +E seg ; At each layer l of the Transformer, self-attention is calculated to capture long-range dependencies: Q=H l-1 W Q K=H l-1 W K V=H l-1 W V Among them, W V , W Q , W K ,∈R d×dk is a trainable parameter; Perform a nonlinear transformation on the self-attention output: FFN(x)=ReLU(xW1+b1)W2+b2; The final output is: H l =LayerNorm(H l-1 +FFN(LayerNorm(H l-1 +Attention(Q,K,V)))); Repeat the above steps L times to get the final text semantic vector H BERT ∈R n×d ; GNN parses the entity-relationship output mapping in the knowledge graph and initializes the entities and relationships in the knowledge graph into low-dimensional vectors: At each layer l, node v aggregates neighbor information: Where ⊕ is vector splicing; Combine the current node features with the aggregate information to update the representation: Where σ is the activation function; Repeat the message passing and updating L times to get the final mapping H GNN ∈R ∣V∣×d ; Then, bimodal fusion is performed, and the semantic vector and the mapping are input into the Transformer encoder to generate a joint representation vector.
6. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The method combines the SWRL rule engine and reinforcement learning to dynamically adjust the reasoning path according to the user's domain preferences, including the following steps: The domain preference model is constructed by user behavior sequence, cognitive style and environmental context; let the user behavior sequence be S=[s 1 ,s 2 ,...,s m ]The cognitive style vector is c∈R k The environmental context vector is e∈R p Then the domain preference vector u∈R d It can be expressed as u=W S LSTM(S)+W C c+W E e, where W S , W C , W E is the trainable weight matrix; The SWRL rule engine matches relevant rules according to the user domain preference u. Let the entity relationship set in the knowledge graph be R, and the SWRL rule base be R={r 1 ,r 2 ,...,r n }; Each rule ri contains a premise and a conclusion, and the similarity between the premise and the user preference is calculated: sim(r i ,u)=cosine(h ri ,u) select the one with the highest similarity; Matching-based rules R u , dynamically adjust the reasoning path in the knowledge graph; let the current query entity be v q , the target entity vt, generates the reasoning path P = [vq, v1, ..., vt] through rule-guided random walk; The transition probability is defined as: Where rvi→vi+1 is the connection v i and v i+1 The relationship between N(v i ) is the neighbor node of vi.
7. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The solidified and orchestrated complex query processes are logically decoupled and efficiently reused. The following steps are involved: First, a visual process model is built based on the BPMN 2.0 standard, which breaks down complex queries into atomic service units of entity linking and relationship reasoning, and defines standardized input and output interfaces. The Camunda process engine is then deployed to manage the lifecycle of process instances, and asynchronous communication between services is performed through the Kafka message queue, so that each service can be deployed and expanded independently. Next, we established a process template library that includes parameterized configuration and version management. We then combined TF-IDF / BERT similarity calculation with a weighted routing algorithm based on historical success rates to intelligently distribute query requests. Route(q)=argmax r (Sim(q,r d esc)·Conf(r)), where Sim is the similarity between the query and the route description; Conf is the route confidence; Ultimately, through dynamic scaling of service replicas and versioned deployment mechanisms, we can improve process reuse efficiency while performing logical decoupling.
8. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The construction of monitoring system to track node performance optimization system in real time, The following steps are involved: Prometheus is used to collect key indicators such as CPU, memory, and response time, and an LSTM neural network is used to predict future trends based on historical load data: Where Lt is the node load at the current moment; n is the time window length; At the same time, dynamic threshold algorithm is combined for anomaly detection: θ dynamic =μ history +k·σ histor y , where k is the adjustment coefficient.
9. The information retrieval intelligent knowledge service electronic device according to claim 1, characterized in that: The process of building a natural interaction bridge, personalizing the interface based on historical conversations, and establishing a closed loop of user feedback includes the following steps: Building a natural interaction bridge: By integrating speech recognition, speech synthesis, natural language processing, and gesture recognition technologies, combined with knowledge graph visualization, a multimodal interactive interface is created that supports voice question and answer, multi-round dialogue, and intuitive information exploration, forming a natural communication bridge. Personalized interface adaptation: Analyze domain keywords in historical conversations and build an academic preference model based on behavioral logs; dynamically adjust the default search bar suggestions through front-end configuration tools; show the new interface to 10% of users, analyze the effects through click heat maps and dwell time, and iteratively optimize the layout plan; Establish a closed loop of user feedback: Identify incorrect annotations and integrate an "Incorrect" button into the result card. Clicking this button triggers an annotation event, recording the user ID, session ID, and incorrect answer version number. The annotation data is then pushed to the correction queue for assignment to domain experts or AI recalculation. The correction results are synchronized to the service orchestration layer, updating the answer priority for that scenario and triggering parameter tuning for related service nodes. System iteration and upgrade: Each clarification conversation, interface adjustment, and correction record is stored in the interactive knowledge base, annotated with timestamps and user characteristics; new data is regularly used to train the interaction model, optimize the clarification strategy generation algorithm and personalized recommendation weights; the new version of the interaction logic is pushed to users in batches, and key indicators are monitored to evaluate the effect.
Citation Information
Cited By
Optical cable intelligent label full life cycle management method and system
CN121052276A
Intelligent information acquisition method and system based on scene design feedback
CN121328679A
Distribution network operation and business expansion project multi-source business data fusion method and system
CN121502702A
A power distribution network operation and industry expansion engineering multi-source service data fusion method and system
CN121502702B
Decision-making method and system for organic semiconductor luminescent material, terminal and medium
CN121709118A