Construction of knowledge graph, reasoning method and device based on knowledge graph

By constructing a multimodal knowledge graph through multimodal knowledge acquisition and weighted fusion, the problems of incomplete knowledge coverage and untimely updates in traditional knowledge bases are solved, enabling dynamic updates and high-quality knowledge graph construction, and significantly improving the coverage and update efficiency of the insurance industry's knowledge base.

CN122432345APending Publication Date: 2026-07-21PICC INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PICC INFORMATION TECH CO LTD
Filing Date
2026-03-03
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional knowledge bases rely on manual organization or single-modal data, making it difficult to effectively integrate multimodal information such as text, images, and audio. This results in incomplete knowledge coverage and untimely updates, which is particularly prominent in the insurance industry.

Method used

We employ a multimodal knowledge acquisition, weighted fusion, and spatiotemporal attribute construction approach. By generating adaptive extraction rules through a large model and combining BERT semantic understanding and cross-modal attention fusion, we construct a multimodal knowledge graph and achieve dynamic updates.

Benefits of technology

It improves the comprehensiveness and timeliness of knowledge base coverage, solves the problems of blind target selection and fragmented multimodal data in traditional methods, and provides a high-quality, structured multimodal data foundation for subsequent knowledge processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432345A_ABST
    Figure CN122432345A_ABST
Patent Text Reader

Abstract

The application discloses a knowledge graph construction method, a reasoning method and device based on the knowledge graph. The knowledge graph construction method comprises the following steps: acquiring multi-modal knowledge; extracting multi-modal features corresponding to the multi-modal knowledge; performing weighted fusion on the multi-modal features to generate structured knowledge triples; and constructing a multi-modal knowledge graph comprising space-time attributes according to the knowledge triples. The application automatically generates structured knowledge triples based on multi-modal knowledge and constructs a multi-modal knowledge graph comprising space-time attributes, which can realize dynamic updating of the multi-modal knowledge graph and improve the comprehensive coverage and timeliness of knowledge updating of the knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to the construction of a knowledge graph, a reasoning method and apparatus based on the knowledge graph. Background Technology

[0002] Currently, the construction and management of enterprise knowledge bases face numerous challenges. Traditional knowledge bases rely on manual organization or single-modal data, making it difficult to effectively integrate multimodal information such as text, images, and audio, resulting in incomplete knowledge coverage and untimely updates. These problems are particularly prominent in the insurance industry, as insurance business involves a large amount of unstructured data, complex clause interpretation, and dynamic risk assessment. Therefore, the industry urgently needs a knowledge base construction method that can meet the industry's requirements for real-time and comprehensive knowledge bases. Summary of the Invention

[0003] The purpose of this application is to provide a knowledge graph construction method and apparatus based on the knowledge graph, so as to solve the problems of incomplete knowledge coverage and untimely updates in traditional knowledge bases in related technologies.

[0004] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide a method for constructing a knowledge graph, comprising: acquiring multimodal knowledge; extracting multimodal features corresponding to the multimodal knowledge; performing weighted fusion on the multimodal features to generate structured knowledge triples; and constructing a multimodal knowledge graph including spatiotemporal attributes based on the knowledge triples.

[0005] Secondly, embodiments of this application provide a reasoning method based on a knowledge graph, comprising: obtaining a user query; obtaining a pre-constructed knowledge graph, wherein the knowledge graph is constructed according to the method described in the first aspect; and generating a corresponding reasoning result based on the user query and the knowledge graph.

[0006] Thirdly, embodiments of this application provide a knowledge graph construction apparatus, comprising: a first acquisition module for acquiring multimodal knowledge; an extraction module for extracting multimodal features corresponding to the multimodal knowledge; a fusion module for performing weighted fusion of the multimodal features to generate structured knowledge triples; and a construction module for constructing a multimodal knowledge graph including spatiotemporal attributes based on the knowledge triples.

[0007] Fourthly, embodiments of this application provide a knowledge graph-based reasoning device, comprising: a second acquisition module for acquiring a user query; a third acquisition module for acquiring a pre-constructed knowledge graph, wherein the knowledge graph is constructed by the device described in the third aspect; and a generation module for generating a corresponding reasoning result based on the user query and the knowledge graph.

[0008] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: In this embodiment, structured knowledge triples are automatically generated based on multimodal knowledge, and a multimodal knowledge graph including spatiotemporal attributes is constructed. This enables dynamic updating of the multimodal knowledge graph, improving the comprehensiveness and timeliness of knowledge coverage in the knowledge base. Attached Figure Description

[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a method for constructing a knowledge graph, provided as an embodiment of this application; Figure 2 A schematic diagram of the overall process of a knowledge graph-based reasoning method provided for one embodiment of this application; Figure 3 A schematic diagram of a knowledge graph construction apparatus provided in one embodiment of this application; Figure 4 A schematic diagram of a knowledge graph-based recommendation device provided for one embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, "and / or" in this application indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. It should be noted that all data involved in this application was obtained with the user's authorization.

[0012] The technical solutions provided in the embodiments of this application are described in detail below with reference to the accompanying drawings.

[0013] Figure 1 This is a flowchart illustrating a method for constructing a knowledge graph, as provided in one embodiment of this application. Figure 1 As shown, the knowledge graph construction method of this application embodiment may specifically include the following steps: S101, acquire multimodal knowledge.

[0014] In this embodiment of the application, the execution entity of the knowledge graph construction method is a knowledge graph construction device, which can be located in an electronic device. This electronic device can be a terminal device or a server. The terminal device can be a mobile phone, tablet computer, desktop computer, laptop, in-vehicle device, etc.; the server can be a standalone server or a server cluster composed of multiple servers. For example, in the insurance business field, the knowledge graph construction device can be located in an insurance business platform to build a knowledge base for the insurance field.

[0015] Acquire knowledge (e.g., web pages) from multiple modalities such as text, images, and audio from at least one data source (e.g., website A, website B).

[0016] As a feasible implementation method, step S101, "acquiring multimodal knowledge", may specifically include the following steps: selecting target data sources based on the relevance of each data source to the needs, content quality, and timeliness of the content; and acquiring multimodal knowledge from the target data sources according to extraction rules.

[0017] Specifically, the comprehensive credibility score of each data source can be calculated based on the relevance of each data source to the needs, content quality, and timeliness of the content. The data sources are then sorted in descending order of their scores, and a preset number (e.g., K or n%) of the top-ranked data sources are selected as the target data sources to achieve the screening of high-value data sources.

[0018] For example, a weighted scoring model that integrates demand relevance, content quality, and content timeliness can be used to calculate the overall credibility score of each data source. The formula for the weighted scoring model can be: ( ) Here, S represents the overall credibility score of the data source (e.g., a website), which consists of three parts; α, β, and γ represent the weighting coefficients of content quality (e.g., page quality), content timeliness (e.g., update frequency), and demand relevance (i.e., demand matching degree), respectively; TF-IDF(D) is used to calculate the keyword density score of the page, reflecting the quality of the content (i.e., professionalism). It constitutes the exponentially decaying term, where It is the time interval between page updates, and λ is the decay coefficient, which ensures that pages with strong timeliness score higher. ( The cosine similarity () represents the similarity between a user's query and the page content, reflecting the relevance of the page content to user needs. It can be calculated from semantic vectors generated by a Bidirectional Encoder Representations from Transformers (BERT) model. The above formula achieves a quantitative assessment of website value by weightedly integrating three dimensions: content quality, content timeliness, and relevance to user needs.

[0019] The pseudocode for calculating the semantic similarity (e.g., cosine similarity) between a user query and page content using the BERT model described above can be shown below: class SemanticScorer: def __init__(self): self.tokenizer =AutoTokenizer.from_pretrained('bert-base-chinese') self.model = AutoModel.from_pretrained('bert-base-chinese') def compute_similarity(self, query, content): inputs=self.tokenizer([query,content],padding=True,return_tensors='pt') outputs = self.model( inputs).last_hidden_state[:,0,:] return cosine_similarity(outputs[0].detach(), outputs[1].detach()) Among them, tokenizer and model are used to load the BERT tokenizer and model respectively to generate semantic vectors; inputs are used to store the query and content text after tokenization; outputs.last_hidden_state is used to obtain the last hidden state of BERT; [0] and [1] indices correspond to the [CLS] token vectors of the query and content respectively; the cosine_similarity function is used to calculate the cosine similarity between the two vectors (query vector and content vector).

[0020] Based on adaptive extraction rules (such as crawling rules) and combined with BERT's semantic understanding capabilities, multimodal knowledge can be accurately extracted from the target data source.

[0021] Extraction rules (such as crawling rules) can be generated using large models. The pseudocode for generating extraction rules using a large model is shown below: prompt_template = ''' Example of a known page structure: {sample_structure} Please generate XPath rules that can extract the following content: - Main text content - Related image links - Release time Output format: JSON dictionary ''' response = mm_llm.generate(prompt_template.format(sample _structure=dom_tree)) rules = json.loads(response) The module is structured as follows: `sample_structure` stores the DOM tree structure of the target webpage, serving as input for rule generation; `prompt_template` defines a prompt template containing the example structure and task requirements; `mm_llm.generate()` calls the multimodal large model interface to generate extraction rules based on the prompts; the returned response contains the extraction rules in JSON format output by the large model, which are parsed by `json.loads()` and stored in the `rules` dictionary. The extraction rules include XPath location rules for fields such as main text content, image links, and publication time. This rule generation module dynamically generates webpage-specific extraction rules, solving the problems of high maintenance costs and poor adaptability caused by manually writing rules in traditional web crawlers. The `rules` dictionary, as the final output, will be directly called by the subsequent crawler engine to perform precise content extraction.

[0022] Furthermore, the knowledge graph construction method of this application embodiment may also include the following rule verification feedback loop steps: calculating the accuracy of multimodal knowledge acquisition; when the accuracy is lower than a preset accuracy threshold, reconstructing the extraction rules.

[0023] Specifically, a rule verification feedback loop can be designed. When the accuracy of multimodal knowledge acquisition falls below a preset accuracy threshold, the extraction rule reconstruction is triggered, forming a closed-loop optimization system. The formula for calculating the accuracy of multimodal knowledge acquisition is as follows:

[0024] Here, Accuracy represents the accuracy rate; Correctly Extracted Fields indicates the number of correctly extracted fields, that is, the number of expected content that the crawler engine successfully extracted from the webpage using the generated XPath rules. For example, if the goal is to extract the title, body, and publication time of the webpage, and the crawler engine successfully extracted the title and body, then the number of correctly extracted fields is 2; Total Fields indicates the total number of fields, that is, the total number of all target content that needs to be extracted from the webpage. Continuing with the example above, if the goal is to extract the title, body, and publication time, then the total number of fields is 3.

[0025] As shown above, this application embodiment constructs an intelligent multi-source knowledge acquisition system. This system achieves full-process automation of knowledge acquisition through a dynamic optimization mechanism driven by a multimodal large model. The system employs a weighted scoring model that integrates content quality, timeliness, and relevance to demand to intelligently filter high-value data sources. It also utilizes a large model to generate adaptive crawling rules and combines BERT semantic understanding capabilities to achieve accurate content extraction. By introducing a rule verification feedback loop and a cross-modal unified indexing system, it possesses continuous evolution capabilities. While ensuring legality and compliance, it effectively solves the problems of blind target selection, rigid content parsing, and fragmented multimodal data inherent in traditional crawling methods, providing a high-quality, structured multimodal data foundation for subsequent knowledge processing.

[0026] S102, Extract the multimodal features corresponding to multimodal knowledge.

[0027] In this embodiment of the application, long texts including multimodal knowledge can be segmented to obtain multiple knowledge units that maintain contextual integrity, and the multimodal features corresponding to each knowledge unit can be extracted.

[0028] For example, a dynamic chunking algorithm based on semantic coherence (i.e., a hierarchical attention chunking model) can be used to analyze the semantic association strength between different sentences using attention weights, thereby intelligently segmenting long texts containing multimodal knowledge into multiple knowledge units that maintain contextual integrity.

[0029] The formula for the block decision function used in the above dynamic block partitioning algorithm can be expressed as follows:

[0030] The attention calculation employs a multi-head mechanism, and the formula is as follows: =Softmax( V in, This represents the optimal text segmentation result determined after optimization calculation. The segmentation scheme that maximizes the total attention score is selected by the argmax function. Represents the query vector With key vector The multi-head attention score reflects the strength of semantic association between different sentences; It is a key vector The dimension scaling factor is used to stabilize gradient calculation; the Softmax function is used to normalize the attention weights, ensuring semantic coherence within each block. The above formula achieves context-aware adaptive text segmentation by quantifying the semantic relationship strength between sentences.

[0031] The pseudocode for the hierarchical attention block model described above can be seen below: class HierarchicalChunker(nn.Module): def __init__(self): self.sentence_encoder = TransformerEncoder() self.chunk_scorer = nn.Linear(768, 1) def forward(self, text): sentences = split_into_sentences(text) embeddings = self.sentence_encoder(sentences) chunks = [] current_chunk = [] for i, emb in enumerate(embeddings): if len(current_chunk)>0: attention_scores = torch.matmul(emb, embeddings[current_chunk].T) if attention_scores.mean() < 0.5: # Adjustable threshold chunks.append(current_chunk) current_chunk = [] current_chunk.append(i) return chunks In the above model implementation, the variable `sentences` stores the list of sentences to be processed; `embeddings` records the vector representation of each sentence after Transformer encoding; `current_chunk` dynamically maintains the index of the text chunk currently being processed; `attention_scores` is used to calculate and store the attention correlation scores between sentences; `threshold` serves as a preset semantic coherence threshold, triggering chunk segmentation when the mean attention score falls below this value; and the final output chunks are a set of text chunks that conform to semantic integrity. These variables collectively support the hierarchical attention chunking model's ability to intelligently chunk long texts.

[0032] Pre-trained models such as BERT and CLIP can be used to generate multimodal features (i.e., multimodal deep semantic representations) corresponding to each knowledge unit.

[0033] S103 performs weighted fusion of multimodal features to generate structured knowledge triples.

[0034] In this embodiment, the fusion weight of each modal feature can be determined based on the semantic importance and contextual relevance of each modal feature; the modal features are then weighted and fused according to the fusion weight to generate knowledge triples.

[0035] It can be used in conjunction with a gated recurrent unit to achieve dynamic fusion of multimodal features through a cross-modal attention fusion mechanism. The final output is a structured knowledge triple with rich contextual information, providing high-quality structured data input for subsequent knowledge graph construction and significantly improving the accuracy of subsequent knowledge reasoning. During the multimodal feature fusion process, a multimodal entity disambiguation system can be constructed simultaneously. This system uses weighted fusion of cross-modal features such as text, images, and audio, combined with a temporal attention mechanism, to resolve entity referencing ambiguities.

[0036] The formula for the above cross-modal attention fusion mechanism can be expressed as follows:

[0037] The formula for calculating attention weights is as follows: Softmax

[0038] in, The entity feature vector after fusion is obtained by weighted summation of the features hm of each modality. This represents the attention weights of each modality normalized by softmax, reflecting the importance of different modalities such as text (t), image (i), and audio (a); It is the learnable projection matrix corresponding to each modality, used to map heterogeneous features to a unified semantic space; Represents the context encoding vector; v and v are the context transformation matrix and attention parameter vector, respectively, which jointly participate in the weight calculation; the tanh activation function ensures gradient stability, and finally realizes the dynamic fusion of cross-modal features, effectively solving the problem of entity referential ambiguity.

[0039] The pseudocode representation of the above cross-modal attention fusion mechanism can be as follows: class MultimodalDisambiguator: def __init__(self): self.text_encoder = BertModel() self.image_encoder = CLIP() self.attention_layer = nn.Linear(2560, 3)# 768+512+1280=2560 def fuse_modalities(self, text, image): t_emb = self.text_encoder(text) i_emb = self.image_encoder(image) combined = torch.cat([t_emb, i_emb], dim=-1) weights = F.softmax(self.attention_layer(combined), dim=-1) return weights[0] t_emb + weights[1] i_emb In this model, `text_encoder` and `image_encoder` are BERT and CLIP models, respectively, used to extract text and image features. The `combined` variable is used to concatenate text and image feature vectors to form a joint representation. The `attention_layer` is a linear layer that outputs the weight scores for each modality. `weights` represents the fusion coefficients of the text and image modalities after softmax normalization. The final weighted feature vector integrates multimodal evidence. The cross-modal attention fusion mechanism is implemented by automatically learning modal importance through a neural network. The `temperature` parameter `τ` controls the sharpness of the weight distribution, influencing the model's preference for the dominant modality.

[0040] S104, Construct a multimodal knowledge graph including spatiotemporal attributes based on knowledge triples.

[0041] In this embodiment, a spatiotemporal awareness mechanism can be introduced into multimodal knowledge fusion, and dynamic knowledge evolution modeling can be achieved through a spatiotemporal graph neural network (ST-GNN). Specifically, semantic modeling can be performed on knowledge triples to obtain a semantic model; a time encoding function is used to convert the time interval between nodes in the semantic model into periodic position encoding vectors; based on the periodic position encoding vectors, a time-aware message vector is generated through a multilayer perceptron; and a multimodal knowledge graph is constructed based on the message vectors.

[0042] The knowledge triples generated in step S103 can be semantically modeled using the Resource Description Framework (RDF) and cross-modal feature fusion can be achieved using node embedding formulas, where text, image, and audio features are encoded using BERT, CLIP, and Wav2Vec, respectively.

[0043] A time-aware message passing mechanism (corresponding to time-aware message functions and spatiotemporal graph neural networks) can be adopted to capture the timeliness of knowledge relationships through time encoding functions.

[0044] The time-aware message function described above can be represented as:

[0045] in, This represents the message vector from node i to node j at timestamp t, generated by a multilayer perceptron (MLP). and These represent the hidden states of node i and node j at the previous time step, respectively. It is a time difference coding function.

[0046] The above time encoding function can be expressed as:

[0047] in, The time encoding function is used to convert the time interval Δt into a periodic position encoding vector; [·||·] represents the vector concatenation operation, which enables the model to consider both topological structure and temporal features; the GNN layer updates the node representation by aggregating spatiotemporal messages, where edge time participates in message computation, ensuring that the timeliness of knowledge association is accurately modeled, and finally forming a dynamically evolving graph representation.

[0048] The pseudocode representation of a time-aware message passing mechanism can be as follows: class TemporalGNNConv(MessagePassing): def __init__(self): super().__init__(aggr='mean') self.time_encoder = nn.Linear(1, 16) self.msg_nn = nn.Sequential( nn.Linear(2 in_dim+16, 128), nn.ReLU() ) def forward(self, x, edge_index, edge_time): return self.propagate(edge_index, x=x, edge_time=edge_time) def message(self, x_i, x_j, edge_time): time_feat = self.time_encoder(edge_time.unsqueeze(-1)) return self.msg_nn(torch.cat([x_i, x_j, time_feat], dim=-1)) In the above code implementation, `time_encoder` maps the scalar time difference Δt to 16-dimensional temporal features; `msg_nn` acts as a message generation network, processing the concatenated source node features `x_i`, target node features `x_j`, and the temporal encoding `edge_time`; the `propagate` method performs message passing in the graph structure; `edge_index` stores the topological connectivity of the graph; `x` represents the feature matrix containing all nodes; and `edge_time` records the temporal attributes of each edge. This code implementation, by separating the temporal encoding and topology processing modules, supports parallel computation of dynamic knowledge relationships. The ReLU activation function ensures nonlinear modeling capability, and the mean aggregation method maintains balanced fusion of neighbor information.

[0049] It can construct multimodal knowledge graphs with spatiotemporal attributes based on the Neo4j graph database, optimize modality alignment using contrastive learning loss, support complex reasoning and dynamic updates, and solve the limitations of traditional static modeling of knowledge graphs.

[0050] For example, InfoNCE loss can be used for contrastive learning of multimodal alignment. The formula for InfoNCE loss can be expressed as:

[0051] In the above formula, This represents the contrastive learning loss function, used to optimize the alignment of multimodal representations; Calculate text feature vectors With image feature vectors The similarity score, of which and Generated by BERT and CLIP encoders respectively; τ is a temperature coefficient hyperparameter used to adjust the smoothness of the probability distribution; the summation term in the denominator The calculation includes negative sample pairs, which improves the model's discriminative ability by comparing positive and negative sample pairs. This formula, through the InfoNCE loss framework, brings related modal samples closer together and pushes away irrelevant samples in the feature space, thereby achieving alignment and fusion of heterogeneous modalities such as text and images in a unified semantic space, providing a consistent multimodal entity representation foundation for subsequent graph inference.

[0052] In summary, the knowledge graph construction method of this application automatically generates structured knowledge triples based on multimodal knowledge and constructs a multimodal knowledge graph including spatiotemporal attributes. This enables dynamic updating of the multimodal knowledge graph, improving the comprehensiveness and timeliness of knowledge base coverage. High-value data sources can be selected based on the relevance, content quality, and timeliness of each data source. Accurate extraction of multimodal knowledge can be achieved by combining adaptive extraction rules generated by a large model with BERT semantic understanding capabilities. By introducing a rule verification feedback loop and a cross-modal unified indexing system, it possesses continuous evolution capabilities. While ensuring legality and compliance, it effectively solves the problems of blind target selection, rigid content parsing, and fragmented multimodal data inherent in traditional crawling methods, providing a high-quality, structured multimodal data foundation for subsequent knowledge processing. Context-aware adaptive text segmentation is achieved by quantifying the semantic association strength between sentences. A cross-modal attention fusion mechanism was employed to achieve dynamic fusion of multimodal features, ultimately outputting structured knowledge triples with rich contextual information. This provides high-quality structured data input for subsequent knowledge graph construction, significantly improving the accuracy of subsequent knowledge reasoning. Dynamic knowledge evolution modeling was implemented using a spatiotemporal graph neural network (ST-GNN). The InfoNCE loss framework was used to shorten the distance between related modal samples and push away irrelevant samples in the feature space, thereby achieving alignment and fusion of heterogeneous modalities such as text and images in a unified semantic space. This provides a consistent multimodal entity representation foundation for subsequent graph reasoning.

[0053] This application also provides a reasoning method based on knowledge graphs. Figure 2 This is a flowchart illustrating a knowledge graph-based reasoning method as provided in one embodiment of this application. Figure 2 As shown, the knowledge graph-based reasoning method of this application embodiment may specifically include the following steps: S201, retrieve user query.

[0054] In this embodiment, the execution entity of the knowledge graph-based reasoning method is a knowledge graph-based reasoning device, which can be located in an electronic device. This electronic device can be a terminal device or a server. The terminal device can be a mobile phone, tablet computer, desktop computer, laptop, in-vehicle device, etc.; the server can be a standalone server or a server cluster composed of multiple servers. For example, in the insurance business field, the knowledge graph-based reasoning device can be located in an insurance business platform for reasoning based on an insurance knowledge base.

[0055] User queries can include questions or content recommendations.

[0056] S202, Obtain the pre-built knowledge graph, which is constructed according to the knowledge graph construction method.

[0057] In this embodiment of the application, the knowledge graph used for reasoning is constructed according to the knowledge graph construction method of any of the above embodiments.

[0058] S203, Generate corresponding reasoning results based on user queries and knowledge graphs.

[0059] In this embodiment of the application, a recursive multi-hop reasoning mechanism can be used to search for the optimal path in the knowledge graph based on the user query; and a reasoning result can be generated based on the optimal path.

[0060] By deeply integrating knowledge graphs and large language models, an intelligent service system with multi-hop reasoning capabilities is constructed. This system employs a dynamically decaying multi-hop reasoning (GraphRAG) mechanism, combining recursive path search with semantic similarity calculation to achieve deep mining of knowledge associations. Simultaneously, a knowledge-aware hybrid recommendation model is designed, collaboratively utilizing collaborative filtering and graph attention networks to jointly model user behavior data with structured knowledge. The system optimizes path search efficiency using priority queues, captures high-order graph relationships through a GAT graph attention layer, and integrates multimodal encoders such as BERT and CLIP to generate unified semantic representations. Ultimately, it achieves intelligent question answering and personalized recommendation services that are accurate, interpretable, and timely, significantly improving cognitive reasoning capabilities in complex scenarios.

[0061] The formula for recursively calculating the path score in the above multi-hop GraphRAG mechanism can be summarized as follows: (q, )+(1- )

[0062] In the above formula, Score(p) represents the comprehensive relevance score of path p, which is recursively calculated from the matching degree of the current node and the scores of subsequent paths; λ∈(0,1) is the decay coefficient (usually taken as 0.7), used to balance the weights of direct association and indirect inference; sim(q,n_1) measures the semantic similarity between the user query q and the first node n_1 in the path, calculated using cosine similarity; p_(2:k) represents the sub-path from the 2nd hop to the kth hop in the path, and its score is recursively calculated using the same mechanism. The above formula achieves deep propagation of query relevance by dynamically weighting and fusing information from multi-hop paths, enabling the intelligent service system to capture both direct associations and potential long-distance knowledge connections.

[0063] The pseudocode for the multi-hop GraphRAG mechanism can be represented as follows: def graph_rag_search(query, kg, max_hops=3): visited = set() frontier = PriorityQueue() frontier.put( (0, kg.match(query)) ) for _ in range(max_hops): current = frontier.get() if current.score > threshold: yield current.path for neighbor in kg.get_neighbors(current.node): new_score = 0.7 current.score + 0.3 cosine_sim(query, neighbor.embed) frontier.put( (new_score, neighbor) ) return sorted(frontier.items(), key=lambda x: -x[0])[:10] In the pseudocode above, `frontier` serves as a priority queue storing nodes to be explored and their cumulative scores; `visited` records processed nodes to prevent duplicate calculations; `threshold` sets a score threshold to terminate low-quality path searches early; `kg.match(query)` initializes seed nodes related to the query; `kg.get_neighbors()` retrieves the multi-hop connections of nodes; and the `cosine_sim` function calculates the real-time matching degree between the query and node features. This implementation expands the search path through a greedy strategy, combined with pruning optimization using a priority queue, controlling computational complexity while maintaining inference depth. The weighting coefficients of 0.7 and 0.3 reflect the cognitive decision preference of "nearest path priority".

[0064] The formula for the knowledge-aware hybrid recommendation model described above can be expressed as follows:

[0065] In the above formula, The predicted rating of user u for item i is composed of two weighted components: Matrix Factorization (MF) and Knowledge Graph Attention (KGAT). μ is a hyperparameter balancing the two methods. MF(u,i) models the collaborative filtering signal using the latent vector dot product of user and item. KGAT(u,i) aggregates multi-hop path information from the knowledge graph using a graph attention network. The GATConv layer learns the importance weights between nodes, path_embs stores the mean of entity embeddings along the path, and finally, average pooling integrates the semantic influence of different paths. The above formula explicitly combines user behavior data and structured information from the knowledge graph, simultaneously capturing collaborative filtering patterns and knowledge-driven reasoning patterns.

[0066] The pseudocode for a knowledge-aware hybrid recommendation model can be represented as follows: def __init__(self): self.embedding = nn.Embedding(num_entities, 64) self.gat_layers = nn.ModuleList([GATConv(64, 64) for _ in range(3)]) def forward(self, user, item): user_emb = self.embedding(user) item_emb = self.embedding(item) path_embs = [] for path in find_paths(user, item): emb = torch.mean(torch.stack([self.embedding(n) for n in path])) path_embs.append(emb) return torch.mean(torch.stack(path_embs)) In the code implementation above, the embedding layer maps users and items to 64-dimensional vectors; `gat_layers` contains a 3-layer graph attention network, with each layer calculating the attention weights of neighboring nodes using `GATConv`; the `find_paths` function retrieves the connection paths between users and items; the `path_embs` list stores the mean embedding of each path, obtained by averaging the node embeddings along the path; finally, the recommendation score combines the base prediction from matrix factorization and the augmented signals from the knowledge graph. Here, `num_entities` specifies the entity size of the knowledge graph, 64 is the embedding dimension, and the 3-layer GAT design balances model depth and overfitting risk, achieving personalized recommendations with knowledge enhancement.

[0067] In summary, the knowledge graph-based reasoning method in this application employs a dynamically decaying multi-hop reasoning mechanism and combines recursive path search with semantic similarity calculation to achieve deep mining of knowledge associations. The knowledge-aware hybrid recommendation model collaboratively utilizes collaborative filtering and graph attention networks to jointly model user behavior data with structured knowledge, enabling intelligent question answering and personalized recommendation services that are accurate, interpretable, and timely, significantly improving cognitive reasoning capabilities in complex scenarios.

[0068] This application also provides an apparatus for constructing a knowledge graph. For example... Figure 3 As shown, the knowledge graph construction apparatus 300 of this application embodiment may specifically include: a first acquisition module 301, an extraction module 302, a fusion module 303, and a construction module 304. Wherein: The first acquisition module 301 is used to acquire multimodal knowledge.

[0069] Extraction module 302 is used to extract multimodal features corresponding to multimodal knowledge.

[0070] The fusion module 303 is used to perform weighted fusion of multimodal features to generate structured knowledge triples.

[0071] Module 304 is used to construct a multimodal knowledge graph including spatiotemporal attributes based on knowledge triples.

[0072] In the embodiments of this application, the specific process by which each module and unit implements its function can be found in the relevant description of any of the above-mentioned knowledge graph construction method embodiments, and will not be repeated here.

[0073] The knowledge graph construction apparatus of this application automatically generates structured knowledge triples based on multimodal knowledge and constructs a multimodal knowledge graph including spatiotemporal attributes. This enables dynamic updating of the multimodal knowledge graph, improving the comprehensiveness and timeliness of knowledge base coverage. High-value data sources can be selected based on the relevance, content quality, and timeliness of each data source. Accurate extraction of multimodal knowledge can be achieved by combining adaptive extraction rules generated by a large model with BERT semantic understanding capabilities. By introducing a rule verification feedback loop and a cross-modal unified indexing system, it possesses continuous evolution capabilities. While ensuring legality and compliance, it effectively solves the problems of blind target selection, rigid content parsing, and fragmented multimodal data inherent in traditional crawling methods, providing a high-quality, structured multimodal data foundation for subsequent knowledge processing. Context-aware adaptive text segmentation is achieved by quantifying the semantic association strength between sentences. A cross-modal attention fusion mechanism is used to achieve dynamic fusion of multimodal features, ultimately outputting structured knowledge triples with rich contextual information. This provides high-quality structured data input for subsequent knowledge graph construction, significantly improving the accuracy of subsequent knowledge reasoning. The evolutionary modeling of dynamic knowledge was achieved through a spatiotemporal graph neural network (ST-GNN). By using the InfoNCE loss framework, the distance between relevant modal samples was shortened in the feature space, while irrelevant samples were pushed away. This enabled the alignment and fusion of heterogeneous modalities such as text and images in a unified semantic space, providing a consistent multimodal entity representation foundation for subsequent graph reasoning.

[0074] This application also provides a reasoning device based on a knowledge graph. For example... Figure 4 As shown, the knowledge graph-based reasoning device 400 of this application embodiment may specifically include: a second acquisition module 401, a third acquisition module 402, and a generation module 403. Wherein: The second acquisition module 401 is used to acquire user queries.

[0075] The third acquisition module 402 is used to acquire a pre-built knowledge graph, which is constructed by the knowledge graph construction device of any of the above embodiments.

[0076] The generation module 403 is used to generate corresponding reasoning results based on user queries and knowledge graphs.

[0077] In the embodiments of this application, the specific process by which each module and unit implements its function can be found in the relevant description of any of the above-mentioned knowledge graph-based reasoning method embodiments, and will not be repeated here.

[0078] The knowledge graph-based reasoning device in this application employs a dynamically decaying multi-hop reasoning mechanism. Through recursive path search combined with semantic similarity calculation, it can achieve deep mining of knowledge associations. The knowledge-aware hybrid recommendation model collaboratively utilizes collaborative filtering and graph attention networks to jointly model user behavior data with structured knowledge, enabling intelligent question answering and personalized recommendation services that are accurate, interpretable, and timely, significantly improving cognitive reasoning capabilities in complex scenarios.

[0079] This application also provides an electronic device. For example... Figure 5 As shown, the electronic device 500 can vary significantly due to differences in configuration or performance. It may include one or more processors 501 and memory 502, with memory 502 storing one or more programs or instructions. Memory 502 can be temporary or permanent storage. The program stored in memory 502 may include one or more modules (not shown), each module including a series of computer-executable instructions for the electronic device 500. Furthermore, processor 501 may be configured to communicate with memory 502, executing the series of programs or computer-executable instructions stored in memory 502 on the electronic device 500. The electronic device 500 may also include one or more power supplies 503, one or more wired or wireless network interfaces 504, one or more input / output interfaces 505, and one or more keyboards 506.

[0080] Specifically, in the embodiments of this application, the electronic device includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of any of the above-described knowledge graph construction method embodiments, or implement the steps of any of the above-described knowledge graph-based reasoning method embodiments.

[0081] In the electronic device of this application embodiment, structured knowledge triples are automatically generated based on multimodal knowledge, and a multimodal knowledge graph including spatiotemporal attributes is constructed. This enables dynamic updating of the multimodal knowledge graph, improving the comprehensiveness and timeliness of knowledge base coverage. A dynamically decaying multi-hop reasoning mechanism, combined with recursive path search and semantic similarity calculation, enables deep mining of knowledge associations. A knowledge-aware hybrid recommendation model collaboratively utilizes collaborative filtering and graph attention networks to jointly model user behavior data with structured knowledge, achieving intelligent question answering and personalized recommendation services that are accurate, interpretable, and timely, significantly improving cognitive reasoning capabilities in complex scenarios.

[0082] This application also proposes a readable storage medium storing one or more computer programs or instructions that, when executed by a processor in an electronic device, enable the processor in the electronic device to perform the steps of any of the above-described knowledge graph construction method embodiments, or to perform the steps of any of the above-described knowledge graph-based reasoning method embodiments.

[0083] In the readable storage medium of this application embodiment, structured knowledge triples are automatically generated based on multimodal knowledge, and a multimodal knowledge graph including spatiotemporal attributes is constructed. This enables dynamic updating of the multimodal knowledge graph, improving the comprehensiveness and timeliness of knowledge coverage in the knowledge base. A dynamically decaying multi-hop reasoning mechanism, combined with recursive path search and semantic similarity calculation, enables deep mining of knowledge associations. A knowledge-aware hybrid recommendation model collaboratively utilizes collaborative filtering and graph attention networks to jointly model user behavior data and structured knowledge, achieving intelligent question answering and personalized recommendation services that are accurate, interpretable, and timely, significantly improving cognitive reasoning capabilities in complex scenarios.

[0084] It should be understood that the training and prediction processes of the AI ​​models involved in the various embodiments of this specification all adhere to multiple legal and compliant principles, including legal data sources, compliant data content, compliant data governance, compliant training objectives and schemes, compliant training processes, compliant training environments and tools, and compliant ethical verification of training results, and comply with the requirements of Article 5 of the Patent Law. Among them: Data source legitimacy: All datasets used for AI model training were obtained through legal means, covering three categories: publicly authorized data, data authorized by partners, and self-collected compliant data. Publicly authorized data comes from compliant data sources following open-source licenses such as Apache 2.0, with complete copyright attribution and authorization scope clearly marked, and no unauthorized open-source code or data reuse. Data authorized by partners has been subject to formal data usage agreements, clearly defining the scope, duration, and confidentiality obligations, and possessing a complete authorization chain. For self-collected data involving personal information, strict informed consent procedures have been followed, and anonymization processes (including but not limited to field masking, feature anonymization, and differential privacy technology applications) have been implemented to remove personally identifiable information, fully complying with the requirements of relevant laws and regulations such as the "Interim Measures for the Administration of Generative Artificial Intelligence Services" and the "Personal Information Protection Law."

[0085] Data content compliance: The AI ​​model's dataset undergoes multiple screenings and cleaning processes to remove all content that may violate social morality or harm public interests. It contains no obscene, pornographic, violent, discriminatory, or information that endangers national or public safety, nor does it involve the illegal acquisition or use of genetic resources. For data in sensitive fields (such as healthcare and finance), an additional privacy-preserving computation module (including federated learning and secure multi-party computation technologies) ensures that the data is "usable but not visible," avoiding compliance risks during the original data transmission process and ensuring that the data application scenarios and uses comply with public order and good morals and industry regulatory requirements.

[0086] Data governance norms: A complete data traceability system is established during the AI ​​model training process to automatically record the source, collection time, annotation process, cleaning rules, and permission allocation of training data, generating traceable compliance reports to ensure that the data is verifiable throughout its entire lifecycle. The dataset annotation process for AI models is completed by a professional human R&D team, clearly defining the proportion of human creative contributions and avoiding reliance on AI-generated data that has not undergone substantial human modification, thus meeting the examination requirements for "human main contributions" in AI patent applications.

[0087] Training objectives and plans are compliant: The AI ​​model training objective focuses on rule extraction and generation. The training scheme and final output results do not violate any mandatory provisions of laws and administrative regulations, do not harm the public interest or the legitimate rights and interests of others, and do not pose any potential risk of being used for illegal activities, privacy infringement, or public safety disruption. The model strictly adheres to the ethical principle of "intelligent for good".

[0088] Training process compliance: A closed-loop training framework is adopted to ensure compliance and controllability of the training process. The specific process is as follows: First, training samples are obtained through compliant data sources. After the aforementioned data cleaning and desensitization, they are input into the neural network model to generate preliminary training results. Second, an expert system is introduced to verify the preliminary results. Based on preset rules and human expert experience, the feasibility of the results is evaluated, and outputs that may pose ethical risks or compliance hazards are corrected (such as removing decision-making logic that violates public order and good morals, and adjusting model parameters that do not comply with safety regulations). Finally, the loss function weights are dynamically optimized based on expert system feedback to strengthen the model's learning of compliant results, avoid overfitting errors or non-compliant labels, and form a closed-loop control of "data input - model training - expert verification - parameter optimization - result feedback" to ensure that the entire training process complies with A5 ethical review requirements.

[0089] Training environment and tool compliance: AI model training is implemented using nationally licensed chips and a compliant training platform. All open-source frameworks and components used in the training process have obtained their corresponding licenses, and copyright statements and patent citation information are fully retained, with no instances of infringement or reuse. The training environment is built using virtual devices (containers / virtual machines) with fixed random seeds and initial parameter configurations to ensure the reproducibility of the training process. Furthermore, through access control and operation log recording, risks such as data leakage and parameter tampering during training are prevented, ensuring the security and compliance of the training process.

[0090] Training results ethical verification compliance: After the model is trained, it undergoes an additional third-party ethical compliance assessment and algorithm filing review to verify that the model output does not violate social morality or harm public interests. For potentially sensitive scenarios (such as public services and intelligent decision-making), a dedicated result verification mechanism is established to ensure that the model always complies with Article 5 of the Patent Law and relevant laws and regulations in practical applications.

[0091] In summary, the data and training process used in the AI ​​model of this specification strictly comply with the relevant provisions of Article 5 of the Patent Law and the Patent Examination Guidelines (2023 Edition), and there are no violations of laws, social ethics, public interests, or illegal use of genetic resources. It fully meets the compliance requirements for patent authorization.

[0092] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0093] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0094] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0095] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0096] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0097] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0098] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0099] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0100] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0101] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0102] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0103] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for constructing a knowledge graph, characterized in that, include: Acquire multimodal knowledge; Extract the multimodal features corresponding to the multimodal knowledge; The multimodal features are weighted and fused to generate structured knowledge triples; Construct a multimodal knowledge graph including spatiotemporal attributes based on the knowledge triples.

2. The method according to claim 1, characterized in that, The acquisition of multimodal knowledge includes: Target data sources are selected based on their relevance to demand, content quality, and timeliness. The multimodal knowledge is obtained from the target data source according to the extraction rules.

3. The method according to claim 2, characterized in that, Also includes: Calculate the accuracy of the multimodal knowledge acquisition; When the accuracy rate is lower than a preset accuracy threshold, the extraction rules are reconstructed.

4. The method according to claim 1, characterized in that, The extraction of multimodal features corresponding to the multimodal knowledge includes: The multimodal knowledge is divided into blocks to obtain multiple knowledge units; Extract the multimodal features corresponding to each knowledge unit.

5. The method according to claim 1, characterized in that, The weighted fusion of the multimodal features to generate structured knowledge triples includes: The fusion weights of each modality feature are determined based on the semantic importance and contextual relevance of each modality feature. The modal features are weighted and fused according to the fusion weights to generate the knowledge triples.

6. The method according to claim 1, characterized in that, The construction of a multimodal knowledge graph including spatiotemporal attributes based on the knowledge triples includes: Semantic modeling is performed on the knowledge triples to obtain a semantic model; A time encoding function is used to convert the time intervals between nodes in the semantic model into periodic position encoding vectors; Based on the periodic position encoding vector, a time-aware message vector is generated by a multilayer perceptron; The multimodal knowledge graph is constructed based on the message vector.

7. A reasoning method based on knowledge graphs, characterized in that, include: Get user queries; Obtain a pre-constructed knowledge graph, wherein the knowledge graph is constructed according to the method described in any one of claims 1-6; Generate corresponding reasoning results based on the user query and the knowledge graph.

8. The method according to claim 7, characterized in that, The step of generating corresponding reasoning results based on the user query and the knowledge graph includes: Based on the user query, a recursive multi-hop reasoning mechanism is used to search for the optimal path in the knowledge graph; The inference result is generated based on the optimal path.

9. A knowledge graph construction apparatus, characterized in that, include: The first acquisition module is used to acquire multimodal knowledge; The extraction module is used to extract the multimodal features corresponding to the multimodal knowledge; The fusion module is used to perform weighted fusion of the multimodal features to generate structured knowledge triples; The construction module is used to construct a multimodal knowledge graph including spatiotemporal attributes based on the knowledge triples.

10. A reasoning device based on a knowledge graph, characterized in that, include: The second acquisition module is used to acquire user queries; The third acquisition module is used to acquire a pre-built knowledge graph, wherein the knowledge graph is constructed by the device as described in claim 9; The generation module is used to generate corresponding reasoning results based on the user query and the knowledge graph.