A classification system completion method and system based on visual injection

CN122310236BActive Publication Date: 2026-09-15NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610685466.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-09-15
Estimated Expiration
2046-05-19

AI Technical Summary

Technical Problem

[0004]真实图像稀缺与抽象概念不可直接检索:在医学、法律、工业等领域,许多概念缺乏可直接检索的高质量图像;对于动作/过程/抽象概念等,真实视觉数据往往更难获得,导致传统多模态方法难以覆盖全体系

Benefits of technology

本发明通过弥补仅文本方法的“感知鸿沟”:通过为概念生成合成图像并进行视觉注入,可利用视觉差异与视觉分组线索辅助结构推断,降低由词汇重叠导致的误判,显著提升分类体系补全的插入位置准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122310236B_ABST
    Figure CN122310236B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electric digital processing, and particularly relates to a classification system completion method and system based on visual injection. The method comprises the following steps: obtaining an existing classification system structure and a query concept to be inserted, generating a synthetic image according to the node and the query concept; converting the synthetic image into an injection input sequence through a structure-perception visual mapping module; performing deep injection fusion on the query concept and the injection input sequence to obtain a final representation; calculating a matching score between a query representation of the final representation and a position representation of the final representation, determining an optimal insertion position of the query concept in the classification system structure according to the matching score, and inserting the query concept into the optimal insertion position. The present application can improve the structure discrimination ability and robustness of the classification system completion by making up the "sensory gap" problem, and is particularly suitable for the classification system completion scene of abstract concepts or professional field concepts lacking real images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital processing technology, and in particular to a classification system completion method and system based on visual injection. Background Technology

[0002] Classification systems are used to organize concepts into hierarchical structures, serving as a crucial foundation for knowledge-driven applications such as search, recommendation, knowledge graph construction, medical subject heading organization, and domain terminology management. As new terms, products, and disease names constantly emerge, existing classification systems often lag behind, necessitating the automatic addition of new concepts—a task known as classification system completion. This task typically involves determining the hierarchical relationship between new and existing concepts and selecting appropriate parent-child positions for insertion.

[0003] In existing technologies, classification system completion methods mostly rely on textual definitions or names as the primary source of information, but in practice, the following problems exist: The problem of the "sensory gap" that relies solely on text: When there is a high degree of lexical overlap in the definitions of different concepts (such as similar products or similar terms), textual methods are prone to misjudging lexical similarity as hierarchical dependence, thus producing structural errors; while visual cues (differences in appearance, differences in construction, key attributes) can provide important signals to distinguish concepts at the same level.

[0004] The scarcity of real images and the inability to directly retrieve abstract concepts: In fields such as medicine, law, and industry, many concepts lack high-quality images that can be directly retrieved; for actions / processes / abstract concepts, real visual data is often even more difficult to obtain, making it difficult for traditional multimodal methods to cover the entire system.

[0005] The quality dilemma and training bias of introducing generated images: Generative models can supplement visual signals to some extent, but generated images may contain illusory noise; at the same time, visual similarity between concepts of the same level may form valuable difficult negative samples. Without a differentiation mechanism, conventional filtering strategies may treat "noise" and "difficult negative samples" as anomalies, leading to the wrong deletion of valid samples or the neglect of key discriminative signals, thus affecting the contrastive learning training effect.

[0006] The problem of unbalanced label density in multimodal fusion: When fusing a small number of visual labels with long text definitions, the visual signal is easily "overwhelmed" by a large number of text labels; while simply increasing the visual weight may amplify the noise brought by low-quality generated images, causing fusion gating conflicts and instability.

[0007] Therefore, a new classification system completion scheme is urgently needed: it can introduce usable visual signals when real visual data is scarce, achieve deep cross-modal interaction in structural reasoning, and have training strategies to suppress illusory noise and mine visually difficult negative samples, while solving the fusion failure problem caused by the imbalance between visual and text label densities. Summary of the Invention

[0008] This invention aims to at least solve one of the technical problems existing in related technologies. To this end, this invention provides a classification system completion method and system based on visual injection, which realizes that by generating synthetic images for concepts and performing visual injection, visual differences and visual grouping cues can be used to assist in structural inference, reduce misjudgments caused by word overlap, and significantly improve the accuracy of insertion position in classification system completion.

[0009] This invention provides a classification system completion method based on visual injection, comprising: S1: Obtain the existing classification system architecture and the query concept to be inserted. The classification system architecture includes a set of nodes, where each node in the set is natural language text, and the query concept is natural language text. S2: Generate a composite image based on the node and the query concept; S3: The synthesized image is converted into an injection input sequence through a structure-aware visual mapping module; S4: Deeply inject and fuse the query concept and the injected input sequence to construct a unified input sequence. Perform cross-modal interactive encoding through a text encoder to obtain the query representation and the location representation. S5: The query representation and the location representation are fused using an adaptive residual fusion mechanism to obtain the final representation; S6: Calculate the matching score between the query concept in the final representation and the candidate position in the final representation, determine the optimal insertion position of the query concept in the classification system based on the matching score, and insert the query concept.

[0010] According to the classification system completion method based on visual injection provided by the present invention, the process of step S2 is as follows: the natural language text of the node is structurally concatenated to construct prompt words, the prompt words and query concepts are input into an image generative model, a synthetic image is generated, and the detailed description text corresponding to the generated image is stored. The detailed description text is composed of the corrected prompt words returned by the generative model.

[0011] According to the classification system completion method based on visual injection provided by the present invention, the process of step S3 is as follows: S31: Construct the input sequence and the target sequence, wherein the format of the input sequence is "A photo of " For query concept The pseudo-label sequence, wherein the target sequence is a detailed description text; S32: Construct a query template, the query template format being "Query Node:Def: Image: The input sequence is replaced with a pseudo-label sequence to represent the query concept name. Query Node is the query node, Def is a text-defined marker, and Image is a marker for image placeholders. Placeholders for the image to be filled; S33: Construct a triplet template, wherein the triplet template format is: Parent Node: Def: Image: ; Child Node:Def: Image: ; Sibling Node: Def: Image: ; In this system, Parent Node represents the parent node, Child Node represents the child node, and Sibling Node represents the sibling node. Define the text for the parent node. Define the text for child nodes. Define the text for sibling nodes. This is a placeholder image for the parent node to be filled. These are placeholders for the images to be filled in the child nodes. Placeholders for the images to be filled in for sibling nodes; S34: Extract the plain text names of the parent node, child node, and sibling node of the synthesized image and fill them into the triplet template to generate a text target sequence. Extract the visual images of the parent node, child node, and sibling node respectively, process them into pseudo-label sequences through a mapping network, and fill them into the name positions corresponding to the triplet template to generate an injection input sequence.

[0012] According to the classification system completion method based on visual injection provided by the present invention, the calculation expression of the pseudo-label sequence in step S31 is as follows: in, For query concept pseudo-labeled sequences, For mapping networks, The visual embedding extracted from the frozen visual encoder.

[0013] According to the classification system completion method based on visual injection provided by the present invention, the process of step S4 is as follows: S41: Construct a unified input sequence for query concepts, replacing the placeholders for query concept names in the query template with a pseudo-tag sequence; S42: Construct a unified input sequence for candidate positions, and replace the name placeholders of the parent node, child node and sibling node in the triplet template with the corresponding pseudo-tag sequences respectively; S43: Input the unified input sequence into the text encoder. The pseudo-label sequence participates in the global self-attention calculation, enabling the text context to dynamically pay attention to relevant visual cues. The calculation results are obtained by mean pooling to obtain query representation and location representation.

[0014] According to the classification system completion method based on visual injection provided by the present invention, the calculation formula for step S5 is as follows: in, For the final representation, This is a query representation obtained by standard pooling the complete sequence. This is a position representation obtained by pooling only the visual pseudo-label sequence. For element-wise multiplication, For learnable scalars, Context-aware gating is used to verify semantic relevance. It is the Sigmoid activation function. Represents the learnable weight matrix; Represents a learnable bias vector; Indicates to and Perform the splicing operation.

[0015] According to the classification system completion method based on visual injection provided by the present invention, the method for calculating the matching score in step S6 is as follows: in, For query concept With candidate positions Match score, The cosine similarity function is used. The query concept that is ultimately represented. These are the candidate positions for the final characterization.

[0016] According to the classification system completion method based on visual injection provided by the present invention, a multimodal guided adaptive reweighting strategy is used to perform comparative learning training on the model. The steps are as follows: S71: Calculate text scores separately using the frozen encoder. and visual score : in, For query concept Textual representation, Candidate positions Textual representation, For query concept Visual representation, For the sample The visual representation, max The operation indicates that the query is visually aligned with any node in the candidate positions. Let p be the sample, c be the sample of the parent node, and s be the sample of the child node. S72: When and Samples rejected by both modalities are marked as noise and subjected to the first soft weighting. This preserves effective long-tail concepts that are ambiguous in a single modality but supported by another modality. S73: If a negative sample is marked as highly similar by any modality, it is considered a hard negative sample and the second soft weighting is increased. The model is forced to distinguish between symmetrical visual similarity and asymmetrical hierarchical implication relationships. S74: Optimizing the weighted pairwise comparison marginal loss: in, For loss function, Cosine distance Marginal value, for operate, For the actual location, For the negative sample set, These are negative samples in the set of negative samples.

[0017] This invention also provides a classification system completion system based on visual injection, comprising: Visual acquisition module: acquires the existing classification system architecture and the query concept to be inserted, wherein the classification system architecture includes a set of nodes, each node in the set of nodes is natural language text, and the query concept is natural language text; generates a synthetic image based on the nodes and the query concept; Visual mapping and pre-training module: The synthetic image is converted into an injection input sequence through a structure-aware visual mapping module; the query concept and the injection input sequence are fused by deep injection to construct a unified input sequence, and cross-modal interactive encoding is performed through a text encoder to obtain query representation and location representation; Deep injection and encoding module: used to perform multimodal fusion of the query representation and the location representation through an adaptive residual fusion mechanism to obtain the final representation; Adaptive residual fusion module: calculates the matching score between the query concept in the final representation and the candidate position in the final representation, determines the optimal insertion position of the query concept in the classification system based on the matching score, and inserts the query concept.

[0018] The above-described one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: This invention bridges the "perceptual gap" of text-only methods: by generating synthetic images for concepts and performing visual injection, visual differences and visual grouping cues can be used to assist in structural inference, reducing misjudgments caused by word overlap and significantly improving the accuracy of insertion positions in classification system completion.

[0019] This invention does not rely on real images, enhancing its feasibility: it constructs usable visual representations for all nodes and query concepts through a generative model, and, in conjunction with a hierarchical fallback strategy, perfectly adapts to long-tail concepts, professional fields, and classification system completion scenarios lacking publicly available real image resources.

[0020] The structure-aware visual mapping of this invention improves the quality of pseudo-tags: through three types of proxy tasks—richness distillation, concept alignment, and structural composability—pseudo-tags can retain fine-grained visual information and adapt to structural roles such as parent, child, and sibling, thereby supporting complex hierarchical topological reasoning.

[0021] The deep injection and adaptive residual fusion of this invention improve cross-modal interaction and robustness: pseudo-tags directly participate in the global self-attention calculation of the text encoder, realizing deep interaction between text and vision; the adaptive residual fusion mechanism decouples "amplification magnitude" and "relevance verification", which not only avoids the visual signal being diluted by long text, but also effectively suppresses irrelevant visual noise interference.

[0022] The multimodal guided reweighting of this invention improves training stability and discriminativeness: by performing noise soft weighting (mutual rescue mechanism) and hard negative sample reinforcement (complementary mining mechanism) through cross-modal consensus, it can still stably optimize the contrastive learning target when there is uncertainty in the synthesized image, and greatly enhance the model's ability to distinguish similar concepts and structural associations.

[0023] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating a classification system completion method based on visual injection provided by the present invention.

[0026] Figure 2 This is a schematic diagram of the structure of a classification system completion system based on visual injection provided by the present invention.

[0027] Figure label: 101. Visual acquisition module; 102. Visual mapping and pre-training module; 103. Depth injection and encoding module; 104. Adaptive residual fusion module. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. The following embodiments are used to illustrate this invention but cannot be used to limit the scope of this invention.

[0029] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0030] The following is combined with Figures 1 to 2 This invention is described.

[0031] Example like Figure 1 As shown, Figure 1 A flowchart illustrating a classification system completion method based on visual injection provided by the present invention includes: S1: Obtain the existing classification system architecture and the query concept to be inserted. The classification system architecture includes a set of nodes and a set of edges. Each node in the set of nodes is natural language text, and the query concept is natural language text. S2: Generate a composite image based on the node and the query concept; S3: The synthesized image is converted into an injection input sequence through a structure-aware visual mapping module; S4: Deeply inject and fuse the query concept and the injected input sequence to construct a unified input sequence. Perform cross-modal interactive encoding through a text encoder to obtain the query representation and the location representation. S5: The query representation and the location representation are fused using an adaptive residual fusion mechanism to obtain the final representation; S6: Calculate the matching score between the query concept in the final representation and the candidate position in the final representation, determine the optimal insertion position of the query concept in the classification system based on the matching score, and insert the query concept.

[0032] Specifically, in step S1 of this embodiment, each node in the node set corresponds to a natural language text and has a text definition, which can be a concept definition, an encyclopedia definition, a professional term definition, etc.; the edge set represents the hierarchical relationship between concepts.

[0033] In this embodiment, the classification system can be Represented as a directed acyclic graph or a tree-like hierarchical structure ,in For a set of concept nodes, Let this be the set of upper and lower edges. The query concept is denoted as... .

[0034] To facilitate the construction of candidate positions, in this embodiment, candidate positions are represented as triplets containing "parent node, child node, and sibling node": in, Let p be the parent node sample, c be the child node sample, and s be the sibling node sample, used to provide clues for discrimination within the same layer. The method for constructing the position of this triplet can be referenced as follows: for each parent-child edge... Generate candidate locations and from Randomly or according to a strategy, select one node from the child nodes as... (e.g., randomly selected) (a child node).

[0035] Specifically, in this embodiment, step S2 involves: structurally concatenating the natural language text of the node to construct prompt words; inputting the prompt words and query concepts into an image generative model to output a synthesized image; storing the detailed description text corresponding to the generated image, wherein the detailed description text is composed of corrected prompt words returned by the generative model, including: S21: Standard generation uses the text definition of the node (or query concept) as a prompt, and generates the corresponding visual image through a generative model. To alleviate ambiguity and polysemy, both the "concept name + text definition" can be used as the prompt. An example prompt format could be: "Please generate a representative image of the concept '[Term]', which is defined as '[Definition]'." S22: Layered rollback strategy. For nodes that cannot be generated or whose generation is limited, a layered strategy is adopted: First, try rewriting the prompt words using "scientific illustration style / abstract description" and then regenerating them; If that still fails, use text rendering to render the concept name or brief definition into the image as a placeholder image to ensure that each concept has visual input.

[0036] S23: Stores detailed description text for the generated image. This detailed description text can be composed of correction prompts returned by the generative model, such as internal model prompts for more detailed descriptions / rewrites. (Note: The last part, "S23," is a typo and can be left as is.) This is used for the pre-training task of subsequent structure-aware visual mapping.

[0037] Specifically, a hierarchical and progressive strategy for constructing prompts and generating images can be employed to ensure both coverage and security. This includes the following steps: First level: The entity name and text definition of the node are structurally concatenated using a preset template. For example, the preset prompt template is: "Please generate the image of concept '[Term]'. Its definition is [Definition]'." By explicitly including the node's text definition in the input, the ambiguity of the node entity name (Term) can be effectively resolved, thereby generating an image that accurately matches the semantics.

[0038] Level Two: If the above standard generation triggers the model's safety policy restrictions, the input model's prompt word template is replaced with an abstract description in the style of a scientific illustration, such as: "An abstract, educational diagram representing '[Definition]' (Concept: '[Term]'). Minimalist line art, safe for work." This abstract style significantly reduces the probability of triggering safety restrictions while retaining the core structured semantic features of the node concept.

[0039] The third level: In extreme cases where images still cannot be generated after the above two attempts, or for artificially assisted structural nodes in the classification system (such as pseudo-root nodes and pseudo-leaf nodes), the system uses text rendering as placeholders. Specifically, the term names of the nodes are rendered as high-contrast text (such as white text) on a solid color background (such as a black background) in a programmatic manner. This setting ensures that the subsequent visual encoder can still obtain valid input by extracting the visual features of the text (such as using optical character recognition (OCR) mechanisms), avoiding abnormal model calculations due to empty input vectors.

[0040] Regarding the format and post-processing of the output data, in this embodiment, the synthetic visual image output by the generative model is preferably 1024x1024 pixels. Simultaneously, to further enrich semantic alignment, when the system calls the generative model (such as the DALL-E3 API), it synchronously extracts and stores the revised prompt returned by the model as a detailed description of the node. Because this revised prompt contains the refinement logic of the image within the generative model, it typically contains richer visual details than the original input, thus more accurately reflecting the actual content of the synthetic image and being used for subsequent comparative learning and alignment training.

[0041] Specifically, the process of step S3 is as follows: S31: Construct the input sequence and the target sequence, wherein the format of the input sequence is "A photo of " For query concept The pseudo-label sequence, wherein the target sequence is a detailed description text; S32: Construct a query template, the query template format being "Query Node:Def: Image: The input sequence is replaced with a pseudo-label sequence to represent the query concept name. Query Node is the query node, Def is a text-defined marker, and Image is a marker for image placeholders. Placeholders for the image to be filled; S33: Construct a triplet template, wherein the triplet template format is: Parent Node: Def: Image: ; Child Node:Def: Image: ; Sibling Node: Def: Image: ; In this system, Parent Node represents the parent node, Child Node represents the child node, and Sibling Node represents the sibling node. Define the text for the parent node. Define the text for child nodes. Define the text for sibling nodes. This is a placeholder image for the parent node to be filled. These are placeholders for the images to be filled in the child nodes. Placeholders for the images to be filled in for sibling nodes; S34: Extract the plain text names of the parent node, child node, and sibling node of the synthesized image and fill them into the triplet template to generate a text target sequence. Extract the visual images of the parent node, child node, and sibling node respectively, process them into pseudo-label sequences through a mapping network, and fill them into the name positions corresponding to the triplet template to generate an injection input sequence.

[0042] Specifically, to prevent the mapping network from ignoring visual signals when the definition text is explicitly present, a two-level Dropout is applied to the definition field, with Dropout probability... Discard the entire defined field; Define internal random masking for discrete word tags, for example, using masked lexical probabilities. Masking terms. This forces the mapping network to learn the "semantic compensation ability of visual tags for missing text," making the model more willing to "consult" pseudo-tags rather than treat them as redundant noise during subsequent deep injection. In this embodiment, stage one only trains the mapping network (freezing the visual encoder and text encoder), trains for 60 epochs, has a batch size of 128, and a learning rate of 1e-4; It is 0.3. The number of pseudo-labels is 0.1. A value of 4 can be chosen to achieve a compromise between information richness and label density.

[0043] Specifically, the calculation expression for the pseudo-label sequence in step S31 is: in, For query concept pseudo-labeled sequences, For mapping networks, The visual embedding extracted from the frozen visual encoder.

[0044] Specifically, the process of step S4 is as follows: S41: Construct a unified input sequence for query concepts, replacing the concept name placeholders in the query template with a pseudo-tag sequence; S42: Construct a unified input sequence for candidate positions, and replace the name placeholders of the parent node, child node and sibling node in the triplet template with the corresponding pseudo-tag sequences respectively; S43: Input the unified input sequence into the text encoder. The pseudo-label sequence participates in the global self-attention calculation, enabling the text context to dynamically pay attention to relevant visual cues. The calculation results are obtained by mean pooling to obtain query representation and location representation.

[0045] Specifically, step S5 uses an adaptive residual fusion mechanism to perform multimodal fusion of query representation and location representation. This mechanism includes a visual enhancement module and a relevance verification module to address the density-selectivity contradiction of "few pseudo-labels and long text definitions causing visual contributions to be submerged in mean pooling": it aims to enhance the visibility of visual signals while suppressing noise from low-quality generated images.

[0046] The calculation formula for step S5 is: in, For the final representation, This is a query representation obtained by standard pooling the complete sequence. This is a position representation obtained by pooling only the visual pseudo-label sequence. For element-wise multiplication, For learnable scalars, Context-aware gating is used to verify semantic relevance. It is the Sigmoid activation function. Represents the learnable weight matrix; Represents a learnable bias vector; Indicates to and Perform the splicing operation.

[0047] Specifically, the matching score is calculated in step S6 as follows: in, For the concept With candidate positions Match score, The cosine similarity function is used. The query concept that is ultimately represented. These are the candidate positions for the final characterization.

[0048] Specifically, the classification system completion method employs a multimodal-guided adaptive reweighting strategy to train the model through comparative learning. The steps are as follows: S71: Calculate text scores separately using the frozen encoder. and visual score : in, For query concept Textual representation, Candidate positions Textual representation, For query concept Visual representation, For the sample The visual representation, max The operation indicates that the query is visually aligned with any node in the candidate positions. Let p be the sample, c be the sample of the parent node, and s be the sample of the child node. S72: When and Both modalities simultaneously reject the sample, i.e. hour, and Labeled as noise, and the first soft weighting is applied. This preserves effective long-tail concepts that are ambiguous in one modality but supported by another; among them, Noise filtering threshold This indicates a logical AND operation.

[0049] S73: If a negative sample is labeled as highly similar by any modality, i.e. If the sample is a difficult negative sample, then the second soft weighting is increased. The model is forced to distinguish between symmetrical visual similarity and asymmetrical hierarchical implication relationships; among which, For a high similarity threshold, This indicates a logical OR.

[0050] S74: Optimizing the weighted pairwise comparison marginal loss: in, For loss function, Cosine distance Marginal value, for operate, For the actual location, For the negative sample set, These are negative samples in the set of negative samples.

[0051] To verify the effectiveness of the method of the present invention, in a specific embodiment, the above-mentioned classification system completion method based on visual injection was experimentally tested according to the full candidate evaluation settings of the disclosed benchmark, and compared with the existing technical methods.

[0052] This embodiment uses the following metrics to evaluate the results of "sorting by insertion position of query concept": 1. Macro Mean Rank (MR): The macro mean of the rank of the actual insertion position among all candidate positions. The smaller the value, the better (marked with "↓" in the table).

[0053] 2. Mean Reciprocal Rank (MRR): The average of the last two rank values ​​of the actual insertion position; the higher the value, the better.

[0054] 3. Recall (R@k, k=1 / 5 / 10): The degree of coverage of the actual insertion position set in the Top-k predicted position set. The higher the value, the better (expressed as a percentage in the table).

[0055] 4. Hit rate (H@k, k=1 / 5 / 10): The proportion of Top-k predicted positions that hit at least one true insertion position. The higher the value, the better (expressed as a percentage in the table).

[0056] The comparison method selected existing representation learning-based classification system completion (TC) methods: TMN, TaxoEnrich, QEN, TaxoComplete, and CoSTC; at the same time, the classification system expansion (TE) methods TaxoExpan and Arborist were adapted as classification system completion tasks as controls.

[0057] Table 1 Comparison of SemEval-Food dataset results

[0058] Table 2 Comparison of MeSH Dataset Results

[0059] Table 3 Comparison of WordNet-Verb 3 dataset results

[0060] As shown in Tables 1 to 3, the test data provided by this invention achieves superior completion results on three different domain datasets: First, in the SemEval-Food dataset, the overall H@1 score of this invention reaches 60.1%, a significant improvement compared to the control method, indicating that introducing visual cues can effectively distinguish concepts with highly similar text definitions. Second, in the WordNet-Verb and MeSH datasets for specialized medical concepts, where the level of abstraction is high, the overall H@1 score of this invention increases to 31.3%, indicating that the visual injection and multimodal guidance mechanism of this invention has good applicability and noise resistance in abstract concept scenarios. Third, testing verifies that the multimodal guided filtering, adaptive residual enhancement, and visual alignment mechanism modules in this invention all play a positive role in improving the accuracy of the insertion position ranking in classification system completion.

[0061] like Figure 2As shown below, a classification system completion system based on visual injection provided by the present invention will be described. The classification system completion system based on visual injection described below and the classification system completion method based on visual injection described above can be referred to and correspond to each other.

[0062] Visual acquisition module 101: acquires an existing classification system structure and a query concept to be inserted, wherein the classification system structure includes a set of nodes, each node in the set of nodes is natural language text, and the query concept is natural language text; generates a synthetic image based on the nodes and the query concept; Visual mapping and pre-training module 102: The synthesized image is converted into an injected input sequence through a structure-aware visual mapping module; the query concept and the injected input sequence are fused by deep injection to construct a unified input sequence; cross-modal interactive encoding is performed through a text encoder to obtain query representation and location representation; Deep injection and encoding module 103: used to perform multimodal fusion of the query representation and the location representation through an adaptive residual fusion mechanism to obtain the final representation; Adaptive residual fusion module 104: calculates the matching score between the query concept in the final representation and the candidate position in the final representation, determines the optimal insertion position of the query concept in the classification system based on the matching score, and inserts the query concept.

[0063] It should be noted that the embodiments of this disclosure can be implemented using hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a programmable memory or a data carrier such as an optical or electronic signal carrier.

[0064] Furthermore, although the operation of the methods of this disclosure is described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowcharts may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps. It should also be noted that the features and functions of two or more devices according to this disclosure may be embodied in one device. Conversely, the features and functions of one device described above may be further divided and embodied by multiple devices.

[0065] While this disclosure has been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the specific embodiments disclosed. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A method for completing a classification system based on visual injection, characterized in that, include: S1: Obtain the existing classification system architecture and the query concept to be inserted. The classification system architecture includes a set of nodes, where each node in the set is natural language text, and the query concept is natural language text. S2: Generate a composite image based on the node and the query concept; S3: The synthesized image is converted into an injection input sequence through a structure-aware visual mapping module; the process of step S3 is as follows: S31: Construct the input sequence and the target sequence, wherein the format of the input sequence is "A photo of " For query concept The pseudo-label sequence, wherein the target sequence is a detailed description text; S32: Construct a query template, the query template format being "Query Node:Def: Image: The input sequence is replaced with a pseudo-label sequence to represent the query concept name. Query Node is the query node, Def is a text-defined marker, and Image is a marker for image placeholders. Placeholders for the image to be filled; S33: Construct a triplet template, wherein the triplet template format is: Parent Node:Def: ,Image: ; Child Node:Def: ,Image: ; Sibling Node:Def: ,Image: ; In this system, Parent Node represents the parent node, Child Node represents the child node, and Sibling Node represents the sibling node. Define the text for the parent node. Define the text for child nodes. Define the text for sibling nodes. This is a placeholder image for the parent node to be filled. These are placeholders for the images to be filled in the child nodes. Placeholders for the images to be filled in for sibling nodes; S34: Extract the plain text names of the parent node, child node and sibling node of the synthesized image and fill them into the triplet template to generate a text target sequence. Extract the visual images of the parent node, child node and sibling node respectively, process them into pseudo-label sequences through the mapping network and fill them into the name positions corresponding to the triplet template to generate an injection input sequence. S4: Deeply inject and fuse the query concept and the injected input sequence to construct a unified input sequence. Perform cross-modal interactive encoding through a text encoder to obtain the query representation and the location representation. S5: The query representation and the location representation are fused using an adaptive residual fusion mechanism to obtain the final representation; S6: Calculate the matching score between the query concept in the final representation and the candidate position in the final representation, determine the optimal insertion position of the query concept in the classification system based on the matching score, and insert the query concept.

2. The classification system completion method based on visual injection according to claim 1, characterized in that, The process of step S2 is as follows: the natural language text of the node is structurally concatenated to construct prompt words, the prompt words and query concepts are input into the image generative model, a synthetic image is generated, and the detailed description text corresponding to the generated image is stored. The detailed description text is composed of the corrected prompt words returned by the generative model.

3. The classification system completion method based on visual injection according to claim 1, characterized in that, The calculation expression for the pseudo-label sequence in step S31 is: in, For mapping networks, The visual embedding extracted from the frozen visual encoder.

4. The classification system completion method based on visual injection according to claim 3, characterized in that, The process for step S4 is as follows: S41: Construct a unified input sequence for query concepts, replacing the placeholders for query concept names in the query template with a pseudo-tag sequence; S42: Construct a unified input sequence for candidate positions, and replace the name placeholders of the parent node, child node and sibling node in the triplet template with the corresponding pseudo-tag sequences respectively; S43: Input the unified input sequence into the text encoder. The pseudo-label sequence participates in the global self-attention calculation, enabling the text context to dynamically pay attention to relevant visual cues. The calculation results are obtained by mean pooling to obtain query representation and location representation.

5. The classification system completion method based on visual injection according to claim 1, characterized in that, The calculation formula for step S5 is: in, For the final representation, This is a query representation obtained by standard pooling the complete sequence. This is a position representation obtained by pooling only the visual pseudo-label sequence. For element-wise multiplication, For learnable scalars, Context-aware gating is used to verify semantic relevance. It is the Sigmoid activation function. Represents the learnable weight matrix; Represents a learnable bias vector; Indicates to and Perform the splicing operation.

6. The classification system completion method based on visual injection according to claim 1, characterized in that, The matching score is calculated in step S6 as follows: in, For query concept With candidate positions Match score, The cosine similarity function is used. The query concept that is ultimately represented. These are the candidate positions for the final characterization.

7. The classification system completion method based on visual injection according to claim 1, characterized in that, The classification system completion method employs a multimodal-guided adaptive reweighting strategy to comparatively train the model. The steps are as follows: S71: Calculate text scores separately using the frozen encoder. and visual score : in, For query concept Textual representation, Candidate positions Textual representation, For query concept Visual representation, For the sample The visual representation, max The operation indicates that the query should be visually aligned with any node in the candidate positions. Let p be the sample, c be the sample of the parent node, and s be the sample of the child node. S72: When and Samples rejected by both modalities are marked as noise and subjected to the first soft weighting. This preserves effective long-tail concepts that are ambiguous in one modality but supported by another; S73: If a negative sample is marked as highly similar by any modality, it is considered a hard negative sample and the second soft weighting is increased. The model is forced to distinguish between symmetrical visual similarity and asymmetrical hierarchical implication relationships. S74: Optimizing the weighted pairwise comparison marginal loss: in, For loss function, Cosine distance Marginal value, for operate, For the actual location, For the negative sample set, These are negative samples in the negative sample set.

8. A classification system completion system based on visual injection, used to execute a classification system completion method based on visual injection as described in any one of claims 1 to 7, characterized in that, include: Visual acquisition module: acquires the existing classification system architecture and the query concept to be inserted, wherein the classification system architecture includes a set of nodes, each node in the set of nodes is natural language text, and the query concept is natural language text; generates a synthetic image based on the nodes and the query concept; Visual mapping and pre-training module: The synthetic image is converted into an injection input sequence through a structure-aware visual mapping module; the query concept and the injection input sequence are fused by deep injection to construct a unified input sequence, and cross-modal interactive encoding is performed through a text encoder to obtain query representation and location representation; Deep injection and encoding module: used to perform multimodal fusion of the query representation and the location representation through an adaptive residual fusion mechanism to obtain the final representation; Adaptive residual fusion module: calculates the matching score between the query concept in the final representation and the candidate position in the final representation, determines the optimal insertion position of the query concept in the classification system based on the matching score, and inserts the query concept.

Citation Information

Patent Citations

  • Patent image few-sample classification method based on multi-modal representation fusion

    CN120375133A

  • Fine-grained continuous learning method based on residual concept guidance

    CN120851131A