Dynamic knowledge fusion image question and answer method based on path correlation
By classifying the relationship types and dynamically fusing knowledge in the external knowledge base, generating candidate knowledge subgraphs and pruning and optimizing them, the problems of low-confidence knowledge interference and knowledge overload in image question answering are solved, and the accuracy and efficiency of the model are improved.
Patent Information
- Application Number
- CN202510944325.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing image question answering technology suffers from low-confidence knowledge interference and knowledge overload problems, cannot effectively distinguish between high-quality and low-quality knowledge, and lacks the ability to dynamically adjust the degree of knowledge fusion, resulting in low model reasoning efficiency.
By classifying the relationship types of the external knowledge base, retaining the lightweight knowledge graph, extracting image and question text entities, generating candidate knowledge subgraphs, and realizing dynamic knowledge fusion and adjusting the degree of knowledge embedding through greedy pruning strategy and knowledge weight coefficient.
By effectively eliminating low-relevance knowledge, the accuracy and computational efficiency of the image question answering model on complex problems are improved, especially in scenarios that require common sense reasoning.
Smart Images

Figure CN120804379A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of data processing, and in particular to a dynamic knowledge fusion image question answering method based on path correlation. BACKGROUND
[0002] Image question answering based on external knowledge is a complex task across the fields of computer vision and natural language processing, aiming to enable computers to answer natural language questions about image content under the condition of embedding external knowledge. Existing image question answering technologies have the problems of interference of low-confidence knowledge and knowledge overload caused by full knowledge fusion. In the aspect of external knowledge selection, existing models often simply embed all retrieved knowledge triples into the reasoning process, failing to effectively distinguish high and low quality knowledge. When low-confidence or irrelevant knowledge is contained in the external knowledge base, these noise knowledge will be embedded into the model, misleading the reasoning path. In addition, traditional knowledge enhancement methods usually adopt a full knowledge fusion strategy, embedding all possible relevant knowledge into the model at once. This approach not only increases the computational complexity, but also reduces the reasoning efficiency of the model due to the interference of a large amount of redundant knowledge. At the same time, the use of a predefined static fusion strategy lacks the ability to dynamically adjust the knowledge fusion degree according to different problem texts. For complex problems that require stronger knowledge support, the fixed fusion strategy cannot meet the differentiated needs. Therefore, it is of great significance to study a technology that can dynamically filter high-quality knowledge, flexibly adjust the knowledge fusion strength, and effectively handle multi-source knowledge, in order to improve the performance of image question answering systems. SUMMARY
[0003] The purpose of the application is to provide a dynamic knowledge fusion image question answering method based on path correlation, and the specific technical solutions are as follows:
[0004] The dynamic knowledge fusion image question answering method based on path correlation comprises the following steps: S1, classifying the relationship types of the external knowledge base, retaining the corresponding relationship types according to the characteristics of the image question answering task, and obtaining a lightweight knowledge graph; S2, extracting entities from the image and the problem text respectively, mapping the entities to the nodes of the lightweight knowledge graph in S1, and obtaining a seed node set; S3, starting from the seed node set in S2, generating a candidate knowledge subgraph through breadth-first search; S4, scoring the knowledge paths in the knowledge subgraph in S3 according to the text matching degree and visual consistency; S5, iteratively optimizing the knowledge subgraph through a greedy pruning strategy according to the knowledge path scores in S4; S6, calculating the knowledge weight coefficient, realizing the dynamic fusion of external knowledge, adaptively adjusting the knowledge embedding degree according to the problem text, and inputting the image question answering model to generate the final answer.
[0005] The relationship types in S1 include a physical attribute and function relationship, a spatial and composition relationship, a dynamic and causality relationship, a hierarchical relationship, a social interaction relationship, a language structure relationship, a logical abstraction relationship, and other relationships.
[0006] The relationship types reserved according to the image question answering task characteristics in S1 need to meet at least one of the following requirements: directly mapping to visual attributes; supporting physical or spatial reasoning; reflecting entity functions or social common sense.
[0007] In S2, the entities are extracted from the image and the question text respectively: the visual entities in the image are extracted through the target detection model Faster R-CNN, obtaining the visual entity set and the detection confidence of each entity attached; the text entities and attributes in the question text are extracted through the natural language processing tool SpaCy, obtaining the text entity set.
[0008] In S2, the extracted entities are mapped to the lightweight knowledge graph nodes in S1, and the visual entity set and the text entity set are mapped to the knowledge graph nodes through the entity and node embedding matching method, obtaining the seed node set.
[0009] When scoring the knowledge path relevance in S4, it includes: constructing a multi-modal scoring function at the path level, for any path:
[0010]
[0011] The correlation scoring function is calculated as:
[0012]
[0013] Wherein, S q (P,Q) is a text matching degree function, S I (P,I) is a visual consistency function, which evaluates the semantic similarity between the path entities and relationships and the keywords of the question, and the mapping strength of the path entities in the image, providing a quantitative basis for the importance of the path.
[0014] The text matching degree function is the matching degree calculation of the question text and the relationship and node in the path, which is used to completely capture the semantic association between the question and the path.
[0015] The visual consistency function is a joint confidence-based fusion mechanism, which introduces the entity existence probability of target detection and the semantic correlation weight between knowledge nodes, and constructs a multi-granularity visual-semantic joint verification function.
[0016] In S5, the greedy pruning strategy is used to iteratively optimize the knowledge subgraph, including: S5.1, calculating the average relevance score S origS5.2, arranging candidate pruning nodes in the subgraph in descending order of node degree; S5.3, performing tentative pruning on each candidate node in turn, temporarily removing the node and its associated edges; S5.4, recalculating the path average correlation score S of the subgraph after pruning pruned ; S5.5, if S pruned > S orig , confirming the pruning operation, updating the subgraph and score benchmark; otherwise, retaining the node; S5.6, repeating steps S5.3 to S5.5 until all candidate nodes are traversed or consecutive multiple pruning fails to improve the quality of the subgraph.
[0017] The calculation method of the knowledge weight coefficient in S6 includes: S6.1, considering the problem text, historical fusion state and current knowledge content, calculating the knowledge weight coefficient a k through a gating mechanism; S6.2, the calculation formula of the knowledge weight coefficient is: a k = sigma (W k · [f (j-1) ; q; k]), wherein f (j-1) is the previous stage fusion state, q is the problem text feature, and k is the knowledge feature; S6.3, when a k tends to 1, the model strengthens the knowledge guided reasoning path; when a k tends to 0, the model degenerates into a pure visual-text fusion mode, realizing dynamic adjustment of the knowledge fusion intensity.
[0018] The above technical scheme of the present application has the following beneficial technical effects: the present application provides a dynamic knowledge fusion image question answering method based on path correlation, which filters the relationship types of the external knowledge base through structuring, retains the core relationship related to visual or problem semantics, and designs a path correlation scoring function to comprehensively consider text matching degree and visual consistency, realizing high-quality knowledge retrieval. The method uses a greedy pruning strategy to iteratively optimize the knowledge subgraph, effectively eliminates low correlation nodes, and realizes dynamic fusion of external knowledge through a knowledge weight coefficient, so that the model can adaptively adjust the knowledge embedding degree according to different problem texts. The method effectively solves the problems of low confidence knowledge interfering with the reasoning path and knowledge overload affecting the efficiency of the model, significantly improves the accuracy on knowledge-intensive problems, especially in complex reasoning scenarios, while maintaining good computational efficiency, and has strong practical value. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart of the present application;
[0020] Figure 2 is a schematic diagram of knowledge entity extraction, seed node set acquisition and candidate subgraph generation in the present application;
[0021] Figure 3 is a process schematic diagram of iteratively optimizing knowledge sub-graphs by a greedy pruning strategy in the present application. DETAILED DESCRIPTION
[0022] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with specific embodiments and with reference to the drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present application. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present application.
[0023] A dynamic knowledge fusion image question answering method based on path relevance, comprising:
[0024] As shown in Figure 1 and 2 , S1, the relationship type classification is performed on the external knowledge base, and the corresponding relationship type is retained according to the characteristics of the image question answering task to obtain a lightweight knowledge graph. In order to strip the language level redundant association on the premise of retaining the common sense reasoning ability, form a lightweight knowledge sub-graph for the image question answering task, and perform a relationship type filtering operation on ConceptNet. The relationship types of ConceptNet are divided into eight categories, including physical property and function relationship, spatial and composition relationship, dynamic and causal relationship, hierarchical relationship, social interaction relationship, language structure relationship, logical abstraction, and other categories. In the image question answering task, the embedded external knowledge needs to at least meet one of the following requirements: directly mapping to visual attributes to answer questions of color, material and the like; supporting physical / spatial reasoning to answer questions of position, causality type; reflecting entity function or social common sense. Through this structured screening, the key knowledge required for visual reasoning is retained, and the language structure and abstract logic irrelevant to visual semantics are eliminated.
[0025] S2, entities are extracted from the image and the question text respectively, and the entities are mapped to the node of the lightweight knowledge graph in S1 to obtain a seed node set. As shown in Figure 1 , the present embodiment includes three parts of knowledge entity extraction, seed node set acquisition and candidate sub-graph generation. First, the entities of the image and the question text are extracted. For visual image entities, the visual entity set of the image is extracted by using the target detection model Faster R-CNN to constitute D={v1,v2,...,v m}, and the detection confidence f vis (v|I) of each entity is obtained, which will be used for subsequent path relevance calculation. For text entity analysis, the text entities and attributes of the question text are extracted by using the natural language processing tool SpaCy to obtain K={k1,k2,...,km}。
[0026] After obtaining the image and the extracted entity set D and K, the entity set is mapped to the knowledge graph node by the way of entity and node embedding matching. Compared with the way of string matching and Levenshtein distance, embedding matching can realize accurate mapping from image and question entities to nodes with different naming forms in the knowledge graph. The entity set finally mapped in the knowledge graph is called the seed set
[0027] S3, starting from the seed node set in S2, a candidate knowledge subgraph is generated by breadth-first search. Starting from the seed node set , a candidate subgraph is generated by the way of breadth-first search (BFS) along the relationship edges of the knowledge graph, and then the union set is taken to merge the repeated nodes and edges:
[0028]
[0029] Wherein, d max is the maximum expansion depth to control the size of the subgraph. The way of BFS can ensure the coverage of the multi-hop reasoning path of the core relationship, while avoiding the local over-expansion caused by depth-first search (DFS).
[0030] As shown in Figure 3 , S4, according to the text matching degree and visual consistency, the relevance score of the knowledge path in the knowledge subgraph in S3 is calculated. The core of the path relevance score calculation is to quantify the semantic relevance of each path in the knowledge subgraph with the question text and image content, as shown in Figure 2 . Specifically, for the initial subgraph generated by the multi-modal entity alignment, first, a multi-modal scoring function at the path level is constructed. For any path The calculation formula of its relevance score function is:
[0031]
[0032] Wherein, S q (P,Q) is the text matching degree function, and S I (P,I) is the visual consistency function. The function quantifies the semantic similarity between the path entity and relationship and the question keywords, and the mapping strength of the path entity in the image, providing a quantitative basis for the importance of the path.
[0033] The text matching degree function includes the matching degree calculation of the question text and the relationship and node in the path, in order to completely capture the semantic association between the question and the path. Specifically, for a given question Q and a candidate path P = {r1, n1, r2, n2, …, r m ,nm} and text matching degree S q (Q, P) is calculated in two levels, relationship matching degree S rel and node matching degree S node . For relationship matching, first extract all relationships R P ={r1, r2, …, r m} in the path, encode each relationship r i into a vector through the pre-trained language model BERT. Input the question embedding vector Q at the same time, calculate the overall matching score of the relationship set
[0034]
[0035] Through average pooling processing of path relationships of different lengths, avoid long path because of relationship quantity is high and score is high. For node matching degree, first extract the entity keyword set K in the question, encode k, and then calculate the maximum similarity according to:
[0036]
[0037] and path node N P ={n1, n2, …, n m}.
[0038] Wherein is the embedding vector of the graph node. The node matching degree adopts the maximum value sum method, aiming to require the nodes in the path to be highly related to at least one question entity. Then fuse the two to form the final text matching degree S q (Q, P) = S rel (Q, R P ) + S node (K, N P ).
[0039] In terms of visual consistency, the embodiment proposes a fusion mechanism based on joint confidence, constructs a multi-granularity visual-semantic joint verification function by introducing the entity existence probability of target detection and the semantic correlation weight between knowledge nodes. Specifically, the visual consistency function S I (I, P) is defined as the product of the visual existence probability and the semantic weight of all nodes in the path, and the calculation formula is:
[0040]
[0041] Wherein, n i represents the node in the knowledge path P, v represents the visual entity set instance output by the Faster RCNN target detection model; f vis (n i|I) represents the node n i The existence probability in image I. The entity coverage granularity of the target detection model is usually limited to predefined categories, which is difficult to cover the fine-grained attributes in the knowledge graph. Therefore, different calculation strategies are adopted for different situations: if the node n i is directly identified by the target detection model, f vis (n i |I) is the confidence score of the detection box; if n i is not detected, the confidence score is calculated by:
[0042]
[0043]
[0044] where D is the set of all detected entities, cos(n i ,v) is the semantic similarity between the knowledge node n i and the detected entity v, which is calculated by the cosine similarity of the pre-trained word vector. By assigning a reasonable confidence to the undetected node through semantic similarity, the final result is calculated by jointly mapping the detection result and the semantic similarity, rather than directly detecting the output, avoiding discarding the correct path due to the defects of the detection model. In addition, the max operation is used to ensure that only the most relevant detection entity is relied on, avoiding the influence of multiple entities. S5, according to the knowledge path score in S4, the knowledge subgraph is iteratively optimized by the greedy pruning strategy;
[0045] The subgraph pruning method proposed in this embodiment is to iteratively optimize by the greedy strategy, and then gradually eliminate low-relevance nodes by the iterative pruning strategy based on the score of the path relevance, and finally retain the most concise and high-confidence reasoning path. The overall process is shown in Figure 3 . Specifically, first, the weighted average value of the scores of all paths in the original subgraph is calculated.
[0046]
[0047]
[0048] where N is the total number of paths. By adding up the scores of all paths in the original subgraph and then averaging, the overall quality of the current subgraph is represented as the basis for subsequent pruning decisions. The candidate nodes are processed in descending order of node degree, forming a candidate node list, and the nodes with dense connections are preferentially pruned. Subsequently, for each candidate node n, a tentative pruning is performed, i.e. temporarily removing the node n and its associated edges, regenerating the pruned subgraph and calculating the updated path score average S pruned .If S pruned >S orig , it is determined that the node n contributes negatively to the overall reasoning, and the node and the associated edges will be pruned.
[0049] In addition, the termination condition is set to terminate the pruning operation when the post-exploration pruning does not result in an improved score. The motivation for this design is that when local optimization cannot further improve the subgraph quality, it is inclined to consider that the current subgraph is close to the optimal state, and further pruning may result in the loss of critical paths due to oversimplification.
[0050] S6, calculate the knowledge weight coefficient, realize the dynamic fusion of external knowledge, adaptively adjust the knowledge embedding degree according to the problem text, input the image question answer model to generate the final answer. For the model input problem text feature and external knowledge The task aims to generate a knowledge weight coefficient α k ∈[0,1], to guide the fusion strength of external knowledge and other modal features.
[0051] The design of the knowledge weight coefficient calculation formula is:
[0052]
[0053] Since the dependence on knowledge varies significantly for different problem texts, such as "whether there is" type questions that require weak knowledge dependence, and "historical background" type that requires strong knowledge retrieval, α k as the fusion weight can adapt to different degrees of external knowledge demand. When α k →1, the model strengthens the knowledge-guided reasoning path; when α k →0, it degenerates into a pure visual-text fusion mode.
[0054] The specific tests of the embodiment are, for example, as follows:
[0055] Training corpus:
[0056] Training example 1: the image content is a blue bird standing on a branch, and the question is "What color is the bird?"
[0057] Training example 2: the image content is a coffee shop scene with a hand-pouring pot and filter paper, and the question is "What is the purpose of the equipment on the counter?"
[0058] Test corpus:
[0059] Test example 1: the image content is a kitchen scene with various cooking tools, and the question is "What is the typical use of the knife shown in the image?"
[0060] Test Case 2: The image content is a pigeon pecking an olive branch, and the question is "What does this scene symbolize?"
[0061] In Test Case 1, the input image is a kitchen scene, and the question asks about the typical use of a knife. This example first extracts entities such as "knife" and "kitchen" from the image, and keywords such as "knife" and "use" from the question. Through breadth-first search, an initial candidate subgraph is generated, containing the following paths:
[0062] Path 1: Knife → UsedFor → Cutting → AtLocation → Kitchen
[0063] Path 2: Knife → UsedFor → Carving
[0064] Path 3: Knife → MadeOf → Metal
[0065] Path 4: Knife → SymbolOf → Danger
[0066] Applying the path relevance score calculation, the text matching degree of Path 1 is 0.91 (the question "the typical use" completely matches the "UsedFor" relationship), the visual consistency is 0.88 (the image indeed detects a knife and a kitchen scene), and the total score is 1.79; the total score of Path 2 is 1.42; the total score of Path 3 is 0.65; the total score of Path 4 is 0.51.
[0067] After the greedy pruning strategy, Path 3 and Path 4 are pruned because the material and symbolic meaning are irrelevant to the current question, and Path 1 and Path 2 are retained. The final generated knowledge subgraph is converted into natural language knowledge: "Knife is used for cutting", "Cutting is usually done in the kitchen", "Knife is used for carving". Through the dynamic fusion mechanism, the knowledge weight coefficient α k = 0.75, which belongs to a high level, because it is a question about the functional use. The model finally outputs the answer "cutting", successfully utilizing external knowledge to complete the reasoning.
[0068] In Test Case 2, the input image is a pigeon pecking an olive branch, and the question asks about the symbolic meaning of the scene. This example first extracts entities such as "pigeon" and "olive branch" from the image, and keywords such as "symbolize" and "scene" from the question. Through breadth-first search, an initial candidate subgraph is generated, containing the following paths:
[0069] Path 1: Pigeon → SymbolOf → Peace
[0070] Path 2: Olive Branch → SymbolOf → Peace
[0071] Path 3: Pigeon -> IsA -> Bird -> CapableOf -> Fly Path 4: Olive Branch -> PartOf -> Olive Tree -> AtLocation -> Mediterranean Sea
[0072] Applying the path relevance score calculation, the text matching degree of path 1 is 0.94 (the "symbolize" in the question is completely matched with the "SymbolOf" relationship), the visual consistency is 0.85, and the total score is 1.79; the total score of path 2 is 1.82; the total score of path 3 is 0.56; and the total score of path 4 is 0.48.
[0073] After the greedy pruning strategy, paths 3 and 4 are pruned (physical properties are irrelevant to symbolic meanings), and paths 1 and 2 are retained. The finally generated knowledge subgraph is converted into natural language knowledge: "pigeon symbolizes peace" and "olive branch symbolizes peace".
[0074] Through the dynamic fusion mechanism, the knowledge weight coefficient a is calculated k = 0.89, which belongs to a higher level, because it is a symbolic meaning problem that requires cultural common sense. The model finally outputs the answer "peace", successfully using external knowledge to complete the reasoning.
[0075] As can be seen from the above test cases, through relationship type filtering, knowledge retrieval and pruning based on path relevance, and dynamic fusion mechanism, the embodiment can effectively extract knowledge related to the current question and image from the external knowledge base, and dynamically adjust the knowledge fusion strength according to the question text, so as to accurately answer the image question answering problem that requires common sense reasoning.
[0076] The dynamic knowledge fusion image question answering model based on path relevance of the embodiment can be applied to various types of image understanding scenarios, such as medical diagnosis, education assistance, intelligent customer service, content review, etc. It can process image question answering tasks that require common sense reasoning and professional knowledge support in these fields. The model has also been applied in various application scenarios in the industry, such as: in the medical field, intelligent interpretation of visual data such as medical images and pathological sections, combined with medical knowledge graph to provide preliminary diagnosis suggestions, thereby providing auxiliary decision support for doctors; in the education field, the platform can use knowledge-enhanced image question answering technology to intelligently analyze textbook illustrations and experimental phenomena, and provide personalized learning guidance for students; in the e-commerce and product display field, the system can use this technology to deeply understand product pictures and answer consumers' questions about product functions, materials, and usage methods, thereby improving the shopping experience; for example, it can be used in car sales platforms to analyze vehicle pictures and answer potential buyers' professional questions about vehicle features and configuration functions, thereby improving conversion rates.
[0077] The above is further detailed description of the present application in combination with specific preferred embodiments, and cannot be deemed as limitation of the specific implementation of the present application to these descriptions. For those skilled in the art to which the present application belongs, without departing from the concept of the present application, a number of equivalent substitutions or obvious variations can be made, and the performance or use is the same, which should be deemed as falling within the protection scope of the present application.
Claims
1. A dynamic knowledge fusion image question answering method based on path correlation, characterized by: include: S1. Classify the relationship types of the external knowledge base, retain the corresponding relationship types according to the characteristics of the image question answering task, and obtain a lightweight knowledge graph; S2. Extract entities from the image and question text respectively, and map the entities to the lightweight knowledge graph nodes in S1 to obtain a set of seed nodes; S3, starting from the seed node set in S2, generating candidate knowledge subgraphs through breadth-first search; S4, scoring the relevance of the knowledge paths in the knowledge subgraph in S3 according to text matching and visual consistency; S5. According to the knowledge path scores in S4, the knowledge subgraph is iteratively optimized through a greedy pruning strategy; S6. Calculate the knowledge weight coefficient to achieve dynamic integration of external knowledge, adaptively adjust the degree of knowledge embedding according to the question text, and input the image question answering model to generate the final answer.
2. The dynamic knowledge fusion image question answering method based on path correlation according to claim 1, characterized in that: In S1, the relationship types of the external knowledge base are classified, where the relationship types include: physical attribute and function relationship, space and composition relationship, dynamic and causal relationship, hierarchical relationship, social interaction relationship, language structure relationship, logical abstract relationship and other relationships.
3. The dynamic knowledge fusion image question answering method based on path correlation as claimed in claim 2, characterized in that: The relationship types retained in S1 based on the characteristics of the image question answering task need to meet at least one of the following requirements: Direct mapping to visual attributes; Support physical or spatial reasoning; Reflect entity functions or social common sense.
4. The dynamic knowledge fusion image question answering method based on path correlation according to claim 1, characterized in that: When extracting entities from images and question texts in S2 respectively: The Faster R-CNN object detection model is used to extract visual entities from the image, obtaining a set of visual entities and the detection confidence associated with each entity. The natural language processing tool SpaCy is used to extract text entities and attributes from the question text to obtain a text entity set.
5. The dynamic knowledge fusion image question answering method based on path correlation as claimed in claim 4, characterized in that: When the extracted entities are mapped to the lightweight knowledge graph nodes in S1 in S2, the visual entity set and the text entity set are mapped to the knowledge graph nodes by matching the entity and node embedding to obtain a seed node set.
6. The dynamic knowledge fusion image question answering method based on path correlation according to claim 1, characterized in that: When scoring the relevance of the knowledge paths in S4, it includes: Construct a path-level multimodal scoring function. For any path: The calculation formula of its relevance scoring function is: Among them, S q (P,Q) is the text matching function, S I (P,I) is a visual consistency function that provides a quantitative basis for path importance by bidirectionally evaluating the semantic similarity between path entities and relations and question keywords, as well as the mapping strength of path entities in the image.
7. The dynamic knowledge fusion image question answering method based on path correlation according to claim 6, characterized in that: The text matching function calculates the matching degree between the question text and the relationships and nodes in the path, and is used to fully capture the semantic association between the question and the path.
8. The dynamic knowledge fusion image question answering method based on path correlation according to claim 6, characterized in that: The visual consistency function is a fusion mechanism based on joint confidence. By introducing the entity existence probability of target detection and the semantic relevance weight between knowledge nodes, a multi-granularity visual-semantic joint verification function is constructed.
9. The dynamic knowledge fusion image question answering method based on path correlation according to claim 1, characterized in that: The iterative optimization of the knowledge subgraph by the greedy pruning strategy in S5 includes: S5.
1. Calculate the average relevance score S of all paths in the original candidate knowledge subgraph orig ; S5.
2. Arrange the candidate pruning nodes in the subgraph in descending order of node degree; S5.
3. Perform tentative pruning on each candidate node in turn, temporarily removing the node and its associated edges; S5.
4. Recalculate the average path correlation score S of the pruned subgraph pruned ; S5.5, if S pruned >S orig , then confirm the pruning operation and update the subgraph and score benchmark; otherwise, keep the node; S5.
6. Repeat steps S5.3 to S5.5 until all candidate nodes are traversed or multiple consecutive pruning attempts fail to improve the subgraph quality.
10. The dynamic knowledge fusion image question answering method based on path correlation according to claim 1, characterized in that: The calculation method of the knowledge weight coefficient in S6 includes: S6.1: Comprehensively consider the question text, historical fusion status and current knowledge content, and calculate the knowledge weight coefficient α through the gating mechanism k ; S6.2: The calculation formula of knowledge weight coefficient is: α k =σ(W k ·[f (j-1) ;q;k]), where f (j-1) is the fusion state of the previous stage, q is the question text feature, and k is the knowledge feature; S6.3: When α k When it approaches 1, the model strengthens the reasoning path guided by knowledge; when α k When it approaches 0, the model degenerates into a pure vision-text fusion mode, achieving dynamic adjustment of the knowledge fusion intensity.
Citation Information
Cited By
Answer generation method and device based on semiconductor knowledge base and medium
CN121256004A