Patent retrieval analysis system and method based on multi-modal fusion and dynamic weight adjustment
Patent Information
- Application Number
- CN202611186687.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-06
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本发明为了弥补现有技术的不足,提供了一种基于多模态融合与动态权重调整的专利检索分析系统和方法,旨在解决传统专利检索系统及现有多模态方法存在的信息孤岛、检索精准度低、跨模态融合不足、权重分配固定、忽视结构化信息与图像精细特征提取等问题,实现对专利多模态信息的深度利用、动态权重调整及差异分析可视化,提升专利检索的准确性和用户体验
[0009]本发明的有益效果在于:通过多模态数据处理模块实现专利文本结构化分割与图像组件级特征提取,解决了传统方法忽视结构化信息和图像精细处理的问题;联合表征构建模块强化了文本与图像的语义关联,打破信息孤岛;多维融合检索模块通过动态权重调整,使不同模态和特征的重要性适配检索需求,提升检索精准度;差异分析可视化模块通过对比矩阵和热力图,支持多维关联分析,提升用户体验,有效满足复杂专利检索及跨领域、多模态深度融合场景的需求。
Smart Images

Figure CN122838569A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, specifically to a patent retrieval and analysis system and method based on multimodal fusion and dynamic weight adjustment. Background Technology
[0002] Traditional patent search systems primarily rely on keyword matching or single-modal (such as text) similarity calculations, which suffer from information silos, inaccurate search results, and a lack of multi-dimensional correlation analysis. With the rapid growth of patent data, how to efficiently utilize multi-modal information such as text and images for in-depth retrieval and analysis has become an urgent problem to be solved.
[0003] In the prior art, some patent retrieval methods attempt to introduce multimodal data, but they generally suffer from the following problems: (1) the semantic association between text and images is not close enough, making it difficult to achieve effective cross-modal fusion; (2) the weight allocation is fixed during the retrieval process, and it is impossible to dynamically adjust the importance of different modalities or features; (3) traditional methods usually ignore the important role of structured information (as claimed, descriptions of figures, etc.) when processing patent texts, and at the same time lack the ability to extract features and model associations of technical components in images. These limitations make it difficult for existing systems to meet the needs of complex patent retrieval, especially in scenarios involving deep fusion of cross-domain or multimodal information. Therefore, there is an urgent need for a patent retrieval method that can make full use of multimodal data, achieve dynamic weight adjustment, and support the visualization of difference analysis, so as to improve the accuracy of retrieval and user experience. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this invention provides a patent retrieval and analysis system and method based on multimodal fusion and dynamic weight adjustment. It aims to solve the problems of information silos, low retrieval accuracy, insufficient cross-modal fusion, fixed weight allocation, and neglect of structured information and fine image feature extraction in traditional patent retrieval systems and existing multimodal methods. It achieves in-depth utilization of patent multimodal information, dynamic weight adjustment, and visualization of difference analysis, thereby improving the accuracy of patent retrieval and user experience.
[0005] The patent retrieval and analysis system and method based on multimodal fusion and dynamic weight adjustment of the present invention includes a system consisting of a multimodal data processing module, a joint characterization construction module, a multidimensional fusion retrieval module, and a difference analysis visualization module. The functions and specific implementations of each module are as follows: I. Multimodal Data Processing Module This module is used for text semantic parsing and image feature extraction of patents, specifically including structured segmentation of the patent text and component-level feature extraction of the patent image. The specific processing procedure is as follows: (A1) The patent text is structured and segmented into paragraph units such as title, abstract, claims, and description of drawings, and named entity recognition is performed through a patent entity recognition model; (A2) Perform grayscale conversion and component detection on the patent image, and use the YOLO model to locate component regions.
[0006] II. Joint Characterization Construction Module This module is used to extract textual and image features from patents and construct a heterogeneous knowledge graph, specifically including: (B1) A patent text encoder model finely tuned with patent domain knowledge is used to perform semantic encoding on each paragraph to generate text semantic vectors; (B2) The patent image encoder model (patent-ViT) fine-tuned with the patent image set is used to encode the preprocessed patent images to generate image feature vectors, and the patent-embedding is used to semantically encode the explanatory text of the patent images to generate image feature-text semantic vector pairs. (B3) Map text entities and image components as heterogeneous graph nodes, and generate semantic and visual association edges through co-occurrence relations, linguistic similarity, and visual similarity to construct a heterogeneous knowledge graph. (Where E represents an entity and R represents a relation edge).
[0007] III. Multidimensional Fusion Search Module This module is used to calculate and fuse multi-dimensional similarity, dynamically adjust the weights of user-input patent queries, and perform multi-dimensional fusion retrieval to complete the patent search ranking. Specifically, it includes: (C1) The patent document is structure-aware modeled using a structure-aware model, and the weights of different paragraphs of the patent, as well as multiple dimensions such as text semantics, image features, vector pair associations, and graph relationships are calculated. (C2) Set domain adaptation coefficients for different technical fields. When the search involves a specific technical field, automatically increase the initial weight ratio of high-frequency feature dimensions in that field. (C3) The semantic vector, patent image feature vector, and image feature-text semantic vector are mapped to the corresponding vector library for vector retrieval, and the retrieval similarity is multiplied by their respective weights to obtain a weighted similarity. The entity relationship similarity based on the heterogeneous knowledge graph is calculated and weighted to obtain a weighted entity relationship similarity. (C4) The text semantic similarity, image feature similarity, vector pair association similarity and graph relationship similarity are fused and a weighted sum is used to obtain a comprehensive similarity score. The comprehensive similarity score is multiplied by the domain adaptation coefficient to obtain the final multidimensional fusion retrieval similarity. The candidate patents are retrieved and ranked based on the multidimensional fusion retrieval similarity.
[0008] IV. Visualization Module for Difference Analysis This module is used to generate contrast matrices and heatmaps, constructing a contrast matrix that includes textual semantic differences and image structural differences, and visualizing the results. Specifically, it includes: (D1) The difference between the text semantic vectors of the query and the candidate patent is calculated by cosine similarity, and the difference between the spatial layout of the image components is calculated by structural difference index to generate a two-dimensional comparison matrix. The matrix element values represent the degree of difference. (D2) Convert the contrast matrix into a heatmap, where color depth is positively correlated with the difference value, and allow users to click on matrix elements to view specific difference segments through an interactive interface, while also providing a quantitative score for the difference.
[0009] The beneficial effects of this invention are as follows: the multimodal data processing module achieves structured segmentation of patent text and extraction of image component-level features, solving the problem of traditional methods neglecting structured information and fine image processing; the joint representation construction module strengthens the semantic association between text and image, breaking down information silos; the multidimensional fusion retrieval module, through dynamic weight adjustment, adapts the importance of different modalities and features to retrieval needs, improving retrieval accuracy; the difference analysis visualization module, through comparison matrices and heatmaps, supports multidimensional association analysis, improves user experience, and effectively meets the needs of complex patent retrieval and cross-domain, multimodal deep fusion scenarios. Attached Figure Description
[0010] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1 This is a flowchart of a patent retrieval method for multimodal fusion and dynamic weight adjustment according to an embodiment of the present invention;
[0012] Figure 2 This is a flowchart of a patent retrieval system with multimodal fusion and dynamic weight adjustment according to an embodiment of the present invention; Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the following detailed description of the patent retrieval and analysis system and method based on multimodal fusion and dynamic weight adjustment of this invention is provided in conjunction with specific embodiments.
[0014] The system described in this invention consists of a multimodal data processing module, a joint characterization construction module, a multidimensional fusion retrieval module, and a difference analysis visualization module. These modules work together to achieve multimodal in-depth retrieval and analysis of patents. The following example, using patent retrieval in the field of "new energy vehicle battery structure," details the specific implementation process of each module.
[0015] I. Multimodal Data Processing Module This module is used for structured processing and feature extraction of patent text and images. The specific steps are as follows:
[0016] (1) Structured segmentation and entity recognition of patent text Taking a patent titled "A High-Safety Power Battery Housing Structure" as an example, its text is structurally segmented into paragraphs including title, abstract, claims (8 in total), and description of drawings (4 figures). Using a BERT-based patent entity recognition model, technical entities such as "power battery housing," "explosion-proof valve," and "heat dissipation channel," as well as "claim 3," "appendix," etc., are identified. Figure 2 Document structure entities such as "".
[0017] (2) Patent image component-level feature extraction Appendix to this patent Figure 2 (Exploded view of battery casing) First, grayscale conversion is performed to eliminate color interference. Then, the three core areas of "casing body", "explosion-proof valve assembly" and "heat dissipation fins" are identified by the finely tuned YOLOv10 model positioning technology component. The bounding box coordinates and confidence scores (both ≥0.9) are output.
[0018] II. Joint Characterization Construction Module This module is used to generate text-image joint vectors and heterogeneous knowledge graphs, and the specific implementation is as follows:
[0019] (1) Text and image feature encoding A patent-embedding model fine-tuned in the patent field is used to semantically encode each paragraph. For example, the paragraph "the pressure trigger threshold of the explosion-proof valve is 0.3 MPa" in claim 3 generates a 512-dimensional semantic vector.
[0020] The patent-ViT model, fine-tuned using a patent image set, is applied to the attached... Figure 2 Encode the data to generate a 512-dimensional image feature vector; simultaneously, the "attached" part of the accompanying description... Figure 2The assembly relationship between the explosion-proof valve and the shell was demonstrated and semantically encoded to form an "image feature-text semantic" vector pair.
[0021] (2) Construction of heterogeneous knowledge graph Text entities (such as "power battery housing") and image components (such as "explosion-proof valve assembly") are mapped as heterogeneous graph nodes, and associated edges are generated through co-occurrence relationships: "claim 3" and "explosion-proof valve assembly" form a semantic associated edge due to the technical limitation relationship (weight 0.85); "explosion-proof valve assembly" and "heat dissipation fins" form a visual associated edge due to the spatial layout relationship (weight 0.78).
[0022] When the knowledge graph is dynamically updated, if a new patent involving "solid-state battery casing" is added to the database, the "solid-state battery" is identified as a variant of "power battery" through a bidirectional LSTM-CRF model. Related edges are added based on the same family citation relationship, and the historical association weights are adjusted according to the time decay factor (set to 0.9).
[0023] III. Multidimensional Fusion Search Module This module achieves multi-dimensional retrieval through dynamic weight adjustment, and the steps are as follows:
[0024] (1) Model training and dynamic weight adjustment Taking a user query for "power battery casing with explosion-proof structure" as an example, the feature vector of the query patent is first generated through a multimodal data processing module. When training the Transformer-based multi-head attention network, a dataset containing 5000 pairs of "query-candidate patents" is used, introducing "claim 3 and appendix..." Figure 2 The mapping relationship is used as a strong supervision signal.
[0025] For searches in the new energy field, a field adaptation coefficient of 1.2 is set to increase the initial weight of high-frequency features such as "explosion-proof performance" and "heat dissipation efficiency".
[0026] (2) Multi-dimensional similarity calculation and ranking Based on a structure-aware model, multi-dimensional weights were calculated: text semantics, image features, vector pairs, and graph relationships were weighted at 0.3, 0.2, 0.25, and 0.25, respectively. The multi-dimensional similarities between the query patent and the candidate patent were calculated as follows: text semantic similarity 0.82, image feature similarity 0.76, vector pair association similarity 0.85, and graph relationship similarity 0.8. The weighted sum yielded a comprehensive similarity of 0.81, placing the candidate patent second in the search results.
[0027] IV. Visualization Module for Difference Analysis This module generates a comparison matrix and a heatmap, as detailed below:
[0028] A difference analysis was performed on the searched patent and the top-ranked candidate patent (related to "liquid-cooled battery casing"). The textual semantic difference was calculated to be 0.35 (low semantic overlap) using cosine similarity, and the image component layout difference was calculated to be 0.6 using structural difference index. A two-dimensional comparison matrix was constructed, where the textual difference element value was 0.35 and the image difference element value was 0.6.
[0029] The matrix is converted into a heatmap: text differences are displayed in dark red, and image differences are displayed in orange. Users can click on the dark red area to view specific differences (such as "explosion-proof valve" vs. "liquid-cooled pipe"). The system simultaneously outputs a quantitative score of 62 points (out of 100) to visually demonstrate the technical differences.
[0030] This embodiment solves the problems of weak cross-modal correlation and fixed weights in traditional retrieval by using multimodal fusion and dynamic weight adjustment, thereby improving the accuracy and depth of patent retrieval and making it suitable for complex patent retrieval scenarios in fields such as new energy and machinery.
Claims
1. A patent search and analysis method based on multimodal fusion and dynamic weight adjustment, characterized in that: Includes the following steps: S1: Processing the patent text and images, specifically including: (A1) The patent text is structured and segmented into paragraph units such as title, abstract, claims, and description of drawings, and named entity recognition is performed through a patent entity recognition model; (A2) Perform grayscale conversion and component detection on the patent image, and use the YOLO model to locate component regions; S2: Constructing a heterogeneous knowledge graph, specifically including: (B1) A patent text encoder model finely tuned with patent domain knowledge is used to perform semantic encoding on each paragraph to generate text semantic vectors; (B2) The patent image encoder model (patent-ViT) fine-tuned with the patent image set is used to encode the preprocessed patent images to generate image feature vectors; and the patent-embedding is used to semantically encode the explanatory text of the patent images to generate image feature-text semantic vector pairs. (B3) Map text entities and image components as heterogeneous graph nodes, and generate semantic and visual association edges through co-occurrence relations, linguistic similarity, and visual similarity to construct a heterogeneous knowledge graph. Where E represents an entity and R represents a relation edge; S3: Perform dynamic weight adjustment and retrieval, specifically including: (C1) The patent document is structure-aware modeled using a structure-aware model, and the weights of different paragraphs of the patent, as well as multiple dimensions such as text semantics, image features, vector pair associations, and graph relationships are calculated. (C2) Set domain adaptation coefficients for different technical fields. When the search involves a specific technical field, automatically increase the initial weight ratio of high-frequency feature dimensions in that field. (C3) The semantic vector, patent image feature vector, and image feature-text semantic vector are mapped to the corresponding vector library for vector retrieval, and the retrieval similarity is multiplied by their respective weights to obtain a weighted similarity. The entity relationship similarity based on the heterogeneous knowledge graph is calculated and weighted to obtain a weighted entity relationship similarity. (C4) The text semantic similarity, image feature similarity, vector pair association similarity and graph relationship similarity are fused and a weighted sum is used to obtain a comprehensive similarity score. The comprehensive similarity score is multiplied by the domain adaptation coefficient to obtain the final multidimensional fusion retrieval similarity. The candidate patents are retrieved and ranked based on the multidimensional fusion retrieval similarity.
2. The patent retrieval and analysis method based on multimodal fusion and dynamic weight adjustment according to claim 1, characterized in that: The patent search and analysis methods also include S4 for constructing a comparison matrix and visualization, specifically including: (D1) The difference between the text semantic vectors of the query and the candidate patent is calculated by cosine similarity, and the difference between the spatial layout of the image components is calculated by structural difference index to generate a two-dimensional comparison matrix. The matrix element values represent the degree of difference. (D2) Convert the contrast matrix into a heatmap, where color depth is positively correlated with the difference value, and allow users to click on matrix elements to view specific difference segments through an interactive interface, while also providing a quantitative score for the difference.
3. The patent retrieval and analysis method based on multimodal fusion and dynamic weight adjustment according to claim 2, characterized in that: The training process of the structure-aware model described in step (C1) includes: (E1) Construct a multimodal heterogeneous training dataset, which includes text paragraph annotations, image component association labels, and cross-modal semantic matching scores for query-candidate patent pairs. The patent-specific "claim-figure" mapping relationship is introduced as a strong supervision signal. (E2) Design a two-stage dynamic loss function: The first stage adopts cross-modal contrastive loss, which distinguishes the feature distribution of text-dominant, image-dominant and mixed-type patents through triple sampling strategy, so that the text-image joint vectors of patents with the same topic form clusters in the feature space; The second stage introduces relation constraint loss, which uses entity association edges in heterogeneous knowledge graphs to construct hard negative samples, and strengthens the model's ability to distinguish subtle semantic differences. (E3) Adopts a domain-adaptive meta-learning strategy. For different technical fields such as mechanics, electronics, and chemistry, it dynamically adjusts the weight allocation mechanism of the attention head through meta-parameters and uses a small number of domain-specific samples for rapid fine-tuning to achieve the generalization ability of cross-domain retrieval. (E4) An adversarial noise training module is embedded, which randomly injects patent term synonym substitution noise into the text semantic vector and adds component occlusion noise into the image feature vector. The robustness of the model to noisy data is improved through game training of generative adversarial network (GAN) while retaining the representation ability of core technical features. (E5) During training, the Normalized Discounted Cumulative Gain (NDCG) metric of retrieval ranking is monitored in real time. When there is no improvement for N consecutive epochs, the feature reweighting mechanism is triggered to automatically strengthen the learning intensity of the model for low-contribution paragraph / component features and avoid getting trapped in local optima.
4. The patent search and analysis method based on multimodal fusion and dynamic weight adjustment according to claim 1, characterized in that: The dynamic updating of the heterogeneous knowledge graph includes: (F1) Construct an entity recognition framework for patent lifecycle awareness. Combine the examination opinion notices of newly entered patents, the citation relationship of family patents and priority information, identify technical term variations and legal entities through a bidirectional LSTM-CRF model, and map them to a predefined entity type system of the knowledge graph. Among them, the limiting terms appearing in the claims are given entity recognition priority. (F2) Design a dynamic calculation model for cross-modal correlation strength: (F2.1) The basic association weights are initially assigned values based on text co-occurrence frequency, visual similarity of image components, and existing relational paths in the knowledge graph; (F2.2) Introduce time decay factor and technology evolution coefficient, apply exponential decay weight to historical correlation edges that are more than M years old, and increase evolution gain weight for entity pairs that are in the same IPC subclass and have high frequency co-occurrence in the past two years. (F2.3) Enhance the reliability of association through text-image mutual verification mechanism: When the "technical feature A corresponds to component B in the attached figure" recorded in the patent text is consistent with the image detection result, the confidence of the AB association edge is increased; if there is a contradiction, the manual verification interface is triggered, and the verification result is used as a model fine-tuning sample. (F3) Construct a predictive incremental update module: (F3.1) Graph neural networks (GNNs) are used to predict the evolution trend of the topology of knowledge graphs, identify potential technology association gaps, and generate a candidate set of association predictions; (F3.2) Combining the "conflicting application" determination rules in the patent examination history, the predicted associations are filtered for compliance, and potential associations that meet the inventiveness requirements of the patent law are retained; (F3.3) Incremental updates are achieved by adopting a hierarchical storage architecture: the core layer stores the claim entities and strongly related edges, and the edge layer stores the attached components and weakly related edges. Intelligent migration of data between the two layers is achieved through a dynamic threshold mechanism, while recording the source of each edge.
5. A patent retrieval and analysis system based on multimodal fusion and dynamic weight adjustment, characterized in that, The patent search and analysis method based on multimodal fusion and dynamic weight adjustment as described in any one of claims 1-4 includes: The multimodal data processing module performs structured segmentation of patent text and component-level feature extraction of patent images; the joint representation construction module extracts textual and image features from patents and constructs a heterogeneous knowledge graph; the multidimensional fusion retrieval module dynamically adjusts the weights of user-inputted query patents and performs multidimensional fusion retrieval to complete the patent search ranking; the difference analysis visualization module constructs a comparison matrix containing textual semantic differences and image structural differences and presents it visually.