Stamp retrieval method and system based on knowledge graph, electronic equipment and medium
By constructing multi-level labels and stamp knowledge graphs, the problem of the existing stamp search methods being highly dependent on state is solved, and high accuracy retrieval on damaged, wrinkled, and faded stamps are achieved, which improves the efficiency and accuracy of stamp search.
Patent Information
- Application Number
- CN202510856792.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing stamp retrieval methods are highly dependent on the stamp status, resulting in a decrease in the recognition accuracy rate when stamps are seriously damaged, wrinkled, faded, etc.
By constructing a stamp search method based on knowledge graph, the entity information and image information in the stamp image are extracted, multi-level labels are constructed and weights are assigned, KV databases and stamp knowledge graphs are constructed, candidate labels are obtained and tag expansion is performed, and incremental stamp sets are finally constructed for sorting search.
It improves the accuracy of stamp retrieval, especially when the stamp is in poor condition, which enhances the speed and accuracy of data query.
Smart Images

Figure CN120386887A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of stamp retrieval, and in particular, to a stamp retrieval method, system, electronic device, and medium based on a knowledge graph. Background Art
[0002] Existing stamp retrieval methods are responsible for capturing images containing stamps to be detected through an image acquisition unit, then the stamp positioning unit is responsible for positioning and extracting the stamps, and then the stamp retrieval unit is responsible for comparing the stamps with those in the database to quickly determine the basic information of the stamps; finally, the result display unit is responsible for outputting and displaying the basic information of the stamps such as name, issue number, printing type, etc.
[0003] The above existing stamp retrieval methods are highly dependent on the state of the stamps. When the stamps are severely damaged, wrinkled, faded, etc., the recognition accuracy of the existing stamp retrieval methods may be greatly affected. Summary of the Invention
[0004] This application aims to propose a stamp retrieval method, system, electronic device, and medium based on a knowledge graph, which can improve the accuracy of stamp retrieval.
[0005] In a first aspect, an embodiment of this application provides a stamp retrieval method based on a knowledge graph, and the method includes: Extracting entity information and image information from a stamp image, where the entity information is text information related to the stamp, and the image information is image information on the stamp image; Constructing multi-level tags according to the entity information and the image information, and assigning a first weight to tags at different levels; Based on the stamp image, the multi-level tags, and the first weight, constructing a KV database, and constructing a stamp knowledge graph according to the stamp entity data and tag entity data in the KV database; Extracting multiple target tags from the text to be queried, and obtaining a candidate tag set by obtaining multiple candidate tags from the stamp knowledge graph according to the multiple target tags; If the candidate tag set contains only one tag, performing tag extension on the tag in the candidate tag set to obtain an extended tag set; Merging the candidate tag set and the extended tag set to obtain a final tag set; Based on the final tag set, constructing an incremental stamp set including a first stamp and a stamp weight; Sorting the first stamps in the incremental stamp set according to the stamp weight to obtain a stamp sorting result, so as to retrieve the target stamp according to the stamp sorting result.
[0006] Compared with the prior art, the first aspect of the present application has the following beneficial effects: In this method, entity information and image information in the stamp image are extracted. The entity information is text information related to the stamp, and the image information is the image information on the stamp image. According to the entity information and the image information, multi-level tags are constructed, and first weights are assigned to tags at different levels. Based on the stamp image, the multi-level tags, and the first weights, a KV database is constructed, and according to the stamp entity data and tag entity data in the KV database, a stamp knowledge graph is constructed. Multiple target tags in the text to be queried are extracted, and multiple candidate tags are obtained from the stamp knowledge graph according to the multiple target tags to obtain a candidate tag set. If the candidate tag set contains only one tag, the tag in the candidate tag set is expanded to obtain an expanded tag set. The candidate tag set and the expanded tag set are merged to obtain a final tag set. Based on the final tag set, an incremental stamp set including the first stamp and the stamp weight is constructed. According to the stamp weight, the first stamps in the incremental stamp set are sorted to obtain a stamp sorting result, so as to retrieve the target stamp according to the stamp sorting result. In this way, first, a KV database is constructed through the constructed multi-level tags, and then a stamp knowledge graph is constructed based on the data in the KV database, which can improve the data query speed, and constructing a knowledge graph with stamp characteristics lays a data foundation for accurately finding stamps in the later stage. Then, multiple candidate tags related to the multiple target tags in the text to be queried are obtained through the stamp knowledge graph, and in the case of only one tag, the tag is expanded to obtain the final tag set, and then the target stamp is retrieved according to the final tag set, which can retrieve stamps according to the text to be queried input by the user when the stamps are severely damaged, wrinkled, faded, etc., and can improve the accuracy of stamp retrieval.
[0007] In some embodiments, constructing the multi-level tags according to the entity information and the image information includes: Constructing the entity information as first-level tags; Constructing the image information as second-level tags; Combining the first-level tags and the second-level tags to construct query information; Retrieving information fragments related to the query information from an external knowledge source to generate third-level tags, where the external knowledge source is a channel for obtaining information data from the outside.
[0008] In some embodiments, obtaining the multiple candidate tags from the stamp knowledge graph according to the multiple target tags to obtain a candidate tag set includes: Matching the multiple target tags with the tags in the KV database to obtain multiple first matching results; Obtain multiple first candidate labels in the stamp knowledge graph from the multiple first matching results to obtain an initial label set; Perform similar semantic expansion on each unmatched target label among the multiple target labels to obtain multiple similar semantic expansion words corresponding to each unmatched target label; Match the multiple similar semantic expansion words with the labels in the KV database to obtain multiple second matching results; Use the similar semantic expansion word with the largest semantic similarity in the multiple second matching results as the second candidate label corresponding to each unmatched target label to obtain a second label set; Merge the initial label set and the second label set to obtain a candidate label set.
[0009] In some embodiments, the performing similar semantic expansion on each unmatched target label among the multiple target labels to obtain multiple similar semantic expansion words corresponding to each unmatched target label includes: Generate a set of candidate synonyms corresponding to each unmatched target label among the multiple target labels; Obtain the first synset set of each unmatched target label, and obtain the second synset set of each candidate synonym in the set of candidate synonyms; Calculate the similarity between each synset in the first synset set and all synsets in the second synset set to obtain a first similarity result; According to the first similarity result, calculate the global similarity between each unmatched target label and its corresponding candidate synonym to obtain a second similarity result; Correct the second similarity result to obtain a third similarity result; Select multiple candidate synonyms as similar semantic expansions according to the third similarity result to obtain multiple similar semantic expansion words corresponding to each unmatched target label.
[0010] In some embodiments, the calculating the global similarity between each unmatched target label and its corresponding candidate synonym according to the first similarity result to obtain a second similarity result includes: ; Wherein, represents the second similarity result, represents the unmatched target label of the first synset set, represents the th synset in the first synset set, represents the synset weight, Indicates the th semantic primitive in the second semantic primitive set, represents the second semantic primitive set of candidate near-synonyms , represents the first similarity result, indicating the search for the maximum similarity.
[0011] In some embodiments, constructing an incremental stamp set including a first stamp and a stamp weight based on the final tag set includes: Obtaining the tag nodes of each tag in the final tag set ; Determining the corresponding tag entity node or stamp entity node of each tag node in the knowledge graph; Constructing an adjacency list according to the tag entity node and the stamp entity node; Obtaining the first stamp associated with each tag in the final tag set according to the adjacency list; Calculating the tag weight between each tag in the final tag set and its associated first stamp; Accumulating the tag weights corresponding to the first stamps to obtain the stamp weight; Constructing an incremental stamp set according to the first stamp and the stamp weight.
[0012] In some embodiments, calculating the tag weight between each tag in the final tag set and its associated first stamp includes: ; wherein, represents the th tag in the final tag set and its associated first stamp between the tag weights, represents the th tag similarity weight, represents the tag and the first stamp in the edge relationship two-dimensional vector between the weights, represents the logarithmic function, represents the tag and the first stamp in the edge relationship two-dimensional vector between the tag levels.
[0013] Second aspect, the embodiment of the present application also provides a stamp retrieval system based on a knowledge graph, the system includes: A data extraction unit for extracting entity information and image information from a stamp image, where the entity information is text information related to the stamp, and the image information is the image information on the stamp image; A first construction unit for constructing multi-level tags according to the entity information and the image information, and assigning first weights to tags at different levels; A second construction unit for constructing a KV database based on the stamp image, the multi-level tags, and the first weights, and constructing a stamp knowledge graph according to the stamp entity data and tag entity data in the KV database; A data acquisition unit for extracting multiple target tags from the text to be queried, and obtaining a candidate tag set by acquiring multiple candidate tags from the stamp knowledge graph according to the multiple target tags; A tag extension unit for performing tag extension on the tag in the candidate tag set if the candidate tag set contains only one tag to obtain an extended tag set; A data merging unit for merging the candidate tag set and the extended tag set to obtain a final tag set; A third construction unit for constructing an incremental stamp set including first stamps and stamp weights based on the final tag set; A stamp retrieval unit for sorting the first stamps in the incremental stamp set according to the stamp weights to obtain a stamp sorting result, so as to retrieve target stamps according to the stamp sorting result.
[0014] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute a stamp retrieval method based on a knowledge graph as described above.
[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute a stamp retrieval method based on a knowledge graph as described above.
[0016] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For the relevant descriptions, please refer to the above first aspect and will not be repeated here. Description of the Drawings
[0017] The above and / or additional aspects and advantages of the present application will become apparent and readily understood from the following description of embodiments in conjunction with the accompanying drawings, where: Figure 1 is a schematic flowchart of an embodiment of the stamp retrieval method based on a knowledge graph provided by the present application; Figure 2 is a schematic overall method flowchart in the best embodiment of the stamp retrieval method based on a knowledge graph provided by the present application; Figure 3 is a schematic diagram of stamp knowledge graph construction in the best embodiment of the stamp retrieval method based on a knowledge graph provided by the present application; Figure 4 is a schematic diagram of the acquisition and storage of multi-level tags in the best embodiment of the stamp retrieval method based on a knowledge graph provided by the present application; Figure 5 is a schematic diagram of stamp knowledge graph retrieval in the best embodiment of the stamp retrieval method based on a knowledge graph provided by the present application; Figure 6 is a schematic structural diagram of an embodiment of the stamp retrieval system based on a knowledge graph provided by the present application; Figure 7 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. Detailed Description of the Embodiment
[0018] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.
[0019] In the description of the present application, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0020] In the description of the present application, it should be understood that with respect to the orientation description, such as up, down, etc., the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present application.
[0021] In the description of this application, it should be noted that unless otherwise clearly defined, terms such as "set", "install", and "connect" should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in this application in combination with the specific content of the technical solution.
[0022] First, parse several nouns involved in this application: BM3D open-source algorithm: It is an open-source image denoising algorithm with excellent effects, especially having a strong denoising effect in retaining textures and edges.
[0023] Ordered set: A special set data structure that not only ensures the uniqueness of elements in the set (no duplicates) but also additionally maintains the sorting of all elements according to an associated score. Its core features are unique elements, automatic sorting of elements by score (usually from small to large), and support for efficient queries based on score ranges (such as obtaining rankings and interval elements).
[0024] HowNet-based similarity calculation: It is a method that uses the concept sememes and semantic relationships in the HowNet knowledge base to measure the similarity between words or sentences. This method calculates the distance between them in the semantic space by analyzing the sememe hierarchical structure and semantic associations of words, thereby evaluating their similarity.
[0025] Sememe: The most basic semantic unit in the HowNet knowledge base, used to describe the core concept features of words. Sememes construct a semantic network through a hierarchical organization method, which can express the abstract attributes and specific meanings of words, provide fine-grained knowledge representation for semantic computing, and support word similarity calculation and semantic reasoning.
[0026] Lowest Common Subsumer (LCS): Refers to the deepest-level ancestor node shared by two nodes in a tree-like or hierarchical structure. It can be used to measure the similarity of words in the semantic hierarchy and calculate the semantic distance.
[0027] HowNet semantic relationship: Refers to the association types between words, such as hyponymy, meronymy, and attribute relationships, etc. These relationships form a rich semantic network, enabling the computer to understand the logical connections between words and supporting applications such as semantic reasoning and knowledge graph construction.
[0028] Database key-value location: It is an efficient data retrieval method based on key-value pairs that quickly locates the corresponding value through a unique key.
[0029] Text cleaning technology: Refers to the preprocessing of original text to remove noise, correct errors, and unify formats. It includes operations such as removing special symbols, correcting spelling mistakes, and standardizing date formats, aiming to improve the quality of text data and make it more suitable for subsequent natural language processing tasks.
[0030] Inverted index: A data structure used for quickly retrieving documents. By recording the list of documents where each term appears, it realizes the reverse mapping from terms to documents. It can significantly improve the efficiency of keyword queries and support fast matching of massive texts.
[0031] Adjacency list: A data storage method for representing graph structures. By maintaining a list of adjacent nodes for each node, it describes the connection relationships of the graph. In graph computing and network analysis, the adjacency list can efficiently store sparse graph data and support operations such as node traversal and path finding.
[0032] Parallel computing framework: A technical architecture that uses multi-processors or multi-machine clusters to execute computing tasks simultaneously, which can accelerate data processing and improve the efficiency of model training and inference.
[0033] Topic level enhancement: A technique for enhancing topic distinctiveness in text modeling. By adjusting the parameters of the topic model or introducing external knowledge, the generated topics are made more representative and interpretable.
[0034] Bucket sorting: An algorithm that distributes data into different "buckets" according to specific rules and then sorts the data within each bucket. It can be used to efficiently handle large-scale vocabulary statistics or feature selection tasks and reduce the time complexity of sorting calculations.
[0035] Due to the strong dependence of existing stamp retrieval methods on the status of stamps, when stamps are severely damaged, wrinkled, faded, etc., the recognition accuracy of existing stamp retrieval methods may be greatly affected.
[0036] To solve the problem of relatively low recognition accuracy of existing stamp retrieval methods, this application proposes a stamp retrieval method, system, electronic device, and medium based on a knowledge graph.
[0037] Refer to Figure 1 , the flowchart of the stamp retrieval method based on a knowledge graph provided by an embodiment of this application. The stamp retrieval method based on a knowledge graph is applied to an electronic device, which can be a server or a mobile terminal, etc. As Figure 1 shown, the stamp retrieval method based on a knowledge graph may include the following steps: Step S100, extract the entity information and image information in the stamp image. The entity information is the text information related to the stamp, and the image information is the image information on the stamp image; Step S200: Construct multi-level tags based on entity information and image information, and assign first weights to tags at different levels; Step S300: Construct a KV database based on the stamp image, multi-level tags, and first weights, and construct a stamp knowledge graph according to the stamp entity data and tag entity data in the KV database; Step S400: Extract multiple target tags from the text to be queried, and obtain a candidate tag set by retrieving multiple candidate tags from the stamp knowledge graph according to the multiple target tags; Step S500: If the candidate tag set contains only one tag, perform tag expansion on the tag in the candidate tag set to obtain an expanded tag set; Step S600: Merge the candidate tag set and the expanded tag set to obtain a final tag set; Step S700: Based on the final tag set, construct an incremental stamp set containing the first stamp and stamp weights; Step S800: Sort the first stamps in the incremental stamp set according to the stamp weights to obtain a stamp sorting result, so as to retrieve the target stamp according to the stamp sorting result.
[0038] In this embodiment, by extracting the entity information and image information in the stamp image, the entity information is the text information related to the stamp, and the image information is the image information on the stamp image; according to the entity information and image information, a multi-level label is constructed, and a first weight is assigned to labels at different levels; based on the stamp image, the multi-level label, and the first weight, a KV database is constructed, and according to the stamp entity data and label entity data in the KV database, a stamp knowledge graph is constructed; multiple target labels in the text to be queried are extracted, and multiple candidate labels are obtained from the stamp knowledge graph according to the multiple target labels to obtain a candidate label set; if the candidate label set contains only one label, the label in the candidate label set is expanded to obtain an expanded label set; the candidate label set and the expanded label set are merged to obtain a final label set; based on the final label set, an incremental stamp set including the first stamp and the stamp weight is constructed; according to the stamp weight, the first stamps in the incremental stamp set are sorted to obtain a stamp sorting result, so as to retrieve the target stamp according to the stamp sorting result. In this way, first, a KV database is constructed through the constructed multi-level label, and then a stamp knowledge graph is constructed based on the data in the KV database, which can improve the data query speed, and constructing a knowledge graph with stamp characteristics lays a data foundation for accurately finding stamps in the later stage; then, multiple candidate labels related to the multiple target labels in the text to be queried are obtained through the stamp knowledge graph, and label expansion is performed when there is only one label to obtain the final label set, and then the target stamp is retrieved according to the final label set, which can retrieve stamps according to the text to be queried input by the user when the stamps are severely damaged, wrinkled, faded, etc., and can improve the accuracy of stamp retrieval.
[0039] The above extraction of the entity information and image information in the stamp image can be to use the same large model to extract the entity information and image information in the stamp image, such as Qwen2.5-vl, and this large model is open source.
[0040] The above extraction of multiple target labels in the text to be queried can be to use the word segmentation technology and natural language processing technology (NLP) to extract multiple target labels in the text to be queried.
[0041] In some embodiments, according to the entity information and image information, constructing a multi-level label includes: Construct the entity information into a first-level label; Construct the image information into a second-level label; Combine the first-level label and the second-level label to construct query information; Retrieve information fragments related to the query information from an external knowledge source to generate a third-level label, and the external knowledge source is a channel for obtaining information data from the outside.
[0042] In this embodiment, entity information is constructed as first-level tags; image information is constructed as second-level tags; by combining the first-level tags and the second-level tags, query information is constructed; information fragments related to the query information are retrieved from an external knowledge source to generate third-level tags, where the external knowledge source is a channel for obtaining information data from the outside. In this way, by constructing multi-level tags, a good data foundation is laid for constructing a knowledge graph in the later stage, solving the problems of single tags and lack of semantic associations in traditional retrieval methods.
[0043] In some embodiments, multiple candidate tags are obtained from the stamp knowledge graph according to multiple target tags to obtain a candidate tag set, including: Match the multiple target tags with the tags in the KV database to obtain multiple first matching results; Obtain multiple first candidate tags in the stamp knowledge graph through the multiple first matching results to obtain an initial tag set; Perform similar semantic expansion on each unmatched target tag among the multiple target tags to obtain multiple similar semantic expansion words corresponding to each unmatched target tag; Match the multiple similar semantic expansion words with the tags in the KV database to obtain multiple second matching results; Use the similar semantic expansion word with the largest semantic similarity among the multiple second matching results as the second candidate tag corresponding to each unmatched target tag to obtain a second tag set; Merge the initial tag set and the second tag set to obtain a candidate tag set.
[0044] In this embodiment, the multiple target tags are matched with the tags in the KV database to obtain multiple first matching results; multiple first candidate tags in the stamp knowledge graph are obtained through the multiple first matching results to obtain an initial tag set; similar semantic expansion is performed on each unmatched target tag among the multiple target tags to obtain multiple similar semantic expansion words corresponding to each unmatched target tag; the multiple similar semantic expansion words are matched with the tags in the KV database to obtain multiple second matching results; the similar semantic expansion word with the largest semantic similarity among the multiple second matching results is used as the second candidate tag corresponding to each unmatched target tag to obtain a second tag set; the initial tag set and the second tag set are merged to obtain a candidate tag set. In this way, first, the multiple target tags are matched with the tags in the KV database, and then the relevant positions in the stamp knowledge graph are located according to the matching results, which can avoid large-scale traversal of the knowledge graph, thereby improving the data query speed. Then, by performing similar semantic expansion on each unmatched target tag among the multiple target tags, the user's query intention (i.e., multiple target tags) is mapped to the tag nodes in the knowledge graph, and appropriate semantic tags can be matched for all the multiple target tags, thus providing a good data foundation for subsequent stamp retrieval.
[0045] In some embodiments, for each unmatched target label among a plurality of target labels, similar semantic expansion is performed to obtain a plurality of similar semantic expansion words corresponding to each unmatched target label, including: Generating a set of candidate near-synonyms corresponding to each unmatched target label among the plurality of target labels; Obtaining a first set of original meanings of each unmatched target label, and obtaining a second set of original meanings of each candidate near-synonym in the set of candidate near-synonyms; Calculating the similarity between each original meaning in the first set of original meanings and all original meanings in the second set of original meanings to obtain a first similarity result; According to the first similarity result, calculating the global similarity between each unmatched target label and its corresponding candidate near-synonym to obtain a second similarity result; Correcting the second similarity result to obtain a third similarity result; Selecting a plurality of candidate near-synonyms as similar semantic expansions according to the third similarity result to obtain a plurality of similar semantic expansion words corresponding to each unmatched target label.
[0046] In this embodiment, through semantic similarity calculation, similar semantic expansion is performed for each unmatched target label among a plurality of target labels, so as to match a suitable semantic label for each unmatched target label, thereby providing a good data basis for subsequent stamp retrieval.
[0047] In some embodiments, according to the first similarity result, calculating the global similarity between each unmatched target label and its corresponding candidate near-synonym to obtain a second similarity result, including: ; Wherein, represents the second similarity result, represents an unmatched target label of the first set of original meanings, represents the th original meaning in the first set of original meanings, is a positive integer, represents the weight of the original meaning of, represents the th original meaning in the second set of original meanings, represents the candidate near-synonym of the second set of original meanings, represents the first similarity result, represents finding the maximum similarity.
[0048] In some implementations, constructing an incremental stamp set comprising a first stamp and a stamp weight based on the final label set includes: Get the label node for each label in the final label set ; Determine the label node for each label in the knowledge graph The corresponding label entity node or stamp entity node; Construct an adjacency list based on the label entity node and the stamp entity node; According to the adjacency list, obtain the first stamp associated with each label in the final label set; Calculate the label weight between each label and its associated first stamp in the final label set; Accumulate the label weight corresponding to the first stamp to obtain the stamp weight; Construct an incremental stamp set based on the first stamp and stamp weight.
[0049] In this embodiment, by obtaining the label node of each label in the final label set ; Determine the label node for each label in the knowledge graph The corresponding label entity node or stamp entity node; construct an adjacency list based on the label entity node and the stamp entity node; obtain the first stamp associated with each label in the final label set based on the adjacency list; calculate the label weight between each label in the final label set and its associated first stamp; accumulate the label weight corresponding to the first stamp to obtain the stamp weight; construct an incremental stamp set based on the first stamp and the stamp weight. In this way, the label node of each label is determined in the knowledge graph. The corresponding tag entity nodes or stamp entity nodes can map the semantic tags in the final tag set to specific stamp instances through the association relationships in the knowledge graph to form a stamp set (i.e., an incremental stamp set). By efficiently traversing the stamp-tag association relationships in the knowledge graph, the accuracy and efficiency of subsequent stamp retrieval can be improved.
[0050] In some implementations, calculating a tag weight between each tag in the final tag set and its associated first stamp includes: ; in, Indicates the final label set Tags The first stamp associated with it The label weights between Indicates The first stamp, is a positive integer, Indicates The similarity weight of the tags, Representation tag The weight value in the edge relationship two-dimensional vector with the first stamp Denote the logarithmic function Representation tag The label level in the edge relationship two-dimensional vector with the first stamp
[0051] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments Since the existing stamp retrieval method is highly dependent on the stamp status, when the stamp is severely damaged, wrinkled, faded, etc., the recognition accuracy of the existing stamp retrieval method may be greatly affected. At the same time, the database needs to be continuously maintained and updated, with weak dynamic adjustment ability, and does not combine well with the issuance characteristics and practical significance of stamps
[0052] To solve the problems existing in the existing stamp retrieval method, referring to Figure 2 , this embodiment proposes a stamp retrieval scheme based on knowledge graph and similarity calculation. This scheme extracts the content of the stamp through a large model, and combines with the RAG technology to inductively generate connotation tags highly relevant to the stamp content and assign relevant weights. According to the relationship between the stamp and the tag, a corresponding KV database and a stamp knowledge graph are constructed. Finally, based on similarity calculation, stamps highly relevant to the user input are retrieved. Compared with the traditional stamp method, this embodiment starts from the stamp and its connotation, and can not only quickly and accurately locate the stamps of interest to the user, but also deeply explore the cultural, historical and other connotation information behind the stamps, providing richer and more comprehensive retrieval results for the user
[0053] The technical solution of this embodiment specifically includes the following content 1. A stamp knowledge graph construction scheme based on multi-level tags
[0054] In order to comprehensively represent stamp information and explore its potential connotation, this embodiment innovatively proposes a stamp knowledge graph construction scheme based on multi-level tags. Its core is the large model and retrieval augmented generation (RAG) technology. Through the large model technology, entity information is extracted, stamp image data is parsed, theme classification and image recognition are realized, and then multi-level tags are generated in combination with the RAG technology. Then, through the weight assignment algorithm combined with the KV database, the stamp data and multi-level tags are further processed to construct a KV database about stamps and tags, and a corresponding stamp knowledge graph is constructed based on this KV database. This scheme upgrades the traditional stamp retrieval to "vision-semantics-knowledge" triple drive through a hierarchical technology stack, with advantages such as high efficiency, high precision and interpretability. The schematic diagram of the stamp knowledge graph construction is as Figure 3 shown
[0055] 1. Acquisition and storage of multi - level tags.
[0056] The generation of multi - level tags is the first step of this solution. Through the layer - by - layer parsing from entity information to pixels and then to semantics, the stamp theme and visual content are combined and extended, and finally structured knowledge units are formed. This system constructs a three - layer tag architecture of "core theme - local elements - derived knowledge", providing a data foundation for the subsequent construction of a knowledge graph and solving the problems of single tags and lack of semantic associations in traditional retrieval methods. The schematic diagram of the acquisition and storage of multi - level tags is as Figure 4 shown. The specific process is as follows: (1) Data input: Receive stamp image data.
[0057] (2) Data pre - processing: Use the BM3D open - source algorithm to eliminate the noise and ink stain interference generated when scanning stamp images, improving the clarity of stamp image data for subsequent processing and recognition by large models.
[0058] (3) Acquisition of first - level and second - level tags: First, classify the entity information of stamps through large - model technology. The entity information of stamps includes stamp names, stamp themes, issue years, etc. These core theme information are identified as first - level tags. Then, use the large - model to detect the stamp image and extract the local elements (such as bamboo, airplane, rice ears, etc.) that appear in the image. These local elements are identified as second - level tags.
[0059] (4) Generation of third - level tags: The generation of third - level tags is based on the first - level and second - level tags. Through the RAG technology, combine the first - level and second - level tags to construct relevant queries (i.e., query information). This query aims to obtain information such as the symbolic meaning, spiritual connotation, and abstract concepts of the tags. Then, retrieve relevant database information fragments from external knowledge sources. Finally, use the large - model to reason and generalize these information fragments and tags, and output derivative tags highly relevant to the stamp theme (for example, bamboo can be extended to "tenacity"). These tags are identified as third - level tags. External knowledge sources can be various channels and resources for obtaining information, data, and knowledge from the outside, such as public publications and digital resources, etc.
[0060] (5) Assignment of tag weights: According to the importance of the content represented by tags of different levels, assign relevant weights to tags of different levels. The specific formula is as follows: 1) Assume that the sum of tag weights for each level is , and each level has tags.
[0061] 2) The sum of weights satisfies .
[0062] 3) The weight distributions of each level of tags are as follows: , , . Among them, , , are the weights of the th tags of the first-level tag, second-level tag, and third-level tag respectively.
[0063] (6) Construction of the KV (such as Redis) database: 1) Data preparation, organizing the relationships between the stamps and tags at all levels and the corresponding weight relationships obtained in the above steps. 2) Define the data types used in the database, which are three types: Hash, Sorted Set, and Set. Among them, Hash is used to store the stamp metadata (i.e., the stamp image and its title name). Sorted Set is used to store weights, first-level tags, second-level tags, and third-level tags. Sorted Set is divided into three levels, corresponding to the first-level tag, second-level tag, and third-level tag respectively. Such a setting facilitates the establishment of an inverted index from the tag name to the node. The specific formula is: ; Among them, is the inverted index from the tag name to the tag node, is the text name of the tag (such as "bamboo", "tough", etc.), is the unique identifier of the corresponding tag node in this database. After that, the tag node ID can be quickly obtained through hash lookup. Then there is Set, whose function is to store the stamp ID.
[0064] Principle of rapid positioning of the KV database: The storage form of the KV database is key-value mapping type, ensuring that each key is unique in the database, and the value is corresponding to the key one by one. This enables the direct positioning of the data location by applying the hash function with the input key and returning the corresponding value. At the same time, multiple tags of stamps can be efficiently stored through hash tables and ordered sets, and the tag weights can be automatically sorted. When retrieving, directly obtain according to the score range, avoiding the steps of "traversing the index" and "full table scan" in the traditional query mode.
[0065] 2. Construction of the stamp knowledge graph.
[0066] A knowledge graph is a knowledge base that organizes data in a graph structure. It describes things and their connections in the real world through three elements: entities (nodes), relationships (edges), and attributes. Based on the generated multi-level tags and stamp image data, this embodiment can construct a stamp knowledge graph. The specific steps include: (1) Data reception: Receive the stamp entity data and multi-level tag data (i.e., tag entity data) from the KV database.
[0067] (2) Definition of entity nodes: Two types of knowledge graph nodes are established. One is the stamp entity node, and the other is the label entity node, including primary labels, secondary labels, and tertiary labels.
[0068] (3) Definition of entity node attributes: Based on the data information received from the KV database, structured attribute information such as stamp name and label hierarchy is defined for the stamp entity node and the label entity node.
[0069] (4) Definition of relationship edges: Based on the entity node types and their semantic associations, the two-dimensional vector of the edge relationship is used to define the pointing logic and weight attributes of the relationship edges, where represents the label level, represent the primary label, secondary label, and tertiary label respectively, represents the weight size. There are two types of relationship types in this knowledge graph, namely stamp-label and label-label. The final result obtained is, for example: Stamp of "Chinese Bamboo Culture" -> Bamboo -> Perseverance.
[0070] II. Stamp knowledge graph retrieval scheme based on similarity calculation.
[0071] In order to achieve efficient and accurate stamp retrieval, this embodiment innovatively proposes a stamp knowledge graph retrieval scheme based on similarity calculation. This scheme includes strategies such as label set extraction, similarity calculation, label association expansion, and instantiation retrieval. Its core lies in using similarity calculation, NLP technology, and database key-value positioning to map the user's query intention to the label nodes in the knowledge graph, and finally realizing instantiation stamp retrieval based on the relationship characteristics of stamp-label, improving the efficiency and accuracy of stamp retrieval. This scheme not only significantly improves the accuracy and robustness of stamp retrieval in complex semantic environments, but also has high flexibility and scalability, suitable for different levels of retrieval requirements. The schematic diagram of stamp knowledge graph retrieval is as Figure 5 shown.
[0072] 1. Introduction to label set extraction.
[0073] This is the first step of this scheme. Its core idea is to use the calculation based on HowNet similarity combined with database key-value positioning to convert the user input query (sentence or label) into a set of semantic labels with weights, providing a data basis for subsequent stamp retrieval. The specific process is as follows: (1) Query input: Receive the user's natural language query text (i.e., the text to be queried).
[0074] (2)Query processing: Use text cleaning technology to remove special symbols such as "@", "#" in the text content and filter out some stop words such as "not as good as", "compared with", etc., to standardize the text data. Then use word segmentation technology and natural language processing technology (NLP) to extract and classify the core words (i.e., target tags) in the text data.
[0075] (3)Initial tag set formation: Perform efficient key-value matching between the extracted target tags and the constructed KV database, and store the candidate tags that can be exactly matched with the tag nodes in the knowledge graph into the initial tag set , where, represents the candidate tag to be stored, is the similarity weight, here takes the value of 1.0, indicating a perfect match.
[0076] First, perform an exact match between the extracted tags and the tags stored in the KV database. For example, "tough" can only match "tough" and cannot match "strong". The tags matched in the KV database can efficiently locate the relevant tag positions in the knowledge graph. This can avoid large-scale traversal of the knowledge graph. Because the entire stamp knowledge graph may have hundreds of thousands of entity nodes and hundreds of thousands of edges, simply inputting a tag such as "tough" into the stamp knowledge graph cannot directly locate the position of the "tough" tag entity node, but rather has to traverse the entire stamp knowledge graph bit by bit, and this process has to be carried out for each tag search, resulting in an excessive amount of computation. By combining the KV database, the relevant positions in the stamp knowledge graph can be efficiently located without traversing the entire stamp knowledge graph bit by bit, improving the efficiency of data query.
[0077] (4)Similarity semantic expansion: Use a large language model to generate a set of candidate synonyms for the unmatched target tag t, and calculate the semantic similarity between each synonym and the target tag based on the semantic similarity of HowNet. The specific method is as follows: 1)Sememe definition: HowNet describes the smallest semantic unit of a word through sememes. Each sememe has a hierarchical relationship. Let the sememe set of the word be , where is a sememe. Define the path of the sememe as the hierarchical path from this sememe to the root node ROOT , represents other sememes between the sememe and the root node, and the of the sememe is the length of the path . For example, the word "strong" may contain sememes For a mental state with a path of: [mental state, mental characteristic, abstract concept, ROOT], its depth is .
[0078] 2) Sememe extraction: Obtain the unmatched target labels from HowNet and the sememe sets of each near-synonym are respectively represented as and .
[0079] 3) Sememe similarity calculation: The similarity between sememes depends on their lowest common subsumer (LCS) and path difference. For each sememe of the target label , calculate its maximum similarity (i.e., the first similarity result) with all sememes of . Its formula is: ; where is the lowest common subsumer of sememe and sememe , represents the path length difference of sememes, that is , is the decay coefficient (default value is 0.2). The exponential term penalizes the path difference to prevent the similarity from being distorted due to too long paths of deep sememes.
[0080] 4) Comprehensive global similarity of words: Through the weighted matching of their respective sememe similarities, the global similarity of word pairs (i.e., the second similarity result) can be calculated. Its formula is: ; where is the weight of sememe , which can be obtained by normalizing the sememe frequency in HowNet, refers to finding the maximum similarity of each sememe of word in the sememe set of word , is the sememe set (i.e., the first sememe set) of word (i.e., each unmatched target label), is the sememe set (i.e., the second sememe set) of word (i.e., the candidate near-synonym).
[0081] 5) Semantic relation correction: There are explicit semantic relations (such as synonymy, antonymy, whole - part) between sememes in HowNet, which can be further used to correct the similarity to obtain the third similarity result. Its semantic relation correction formula is: ; Among them, is a correction factor, which can be adjusted according to semantic relationships: if it is a synonym relationship, ; if it is an antonym relationship, ; if it is a whole - part relationship, .
[0082] 6) Result screening: Set a customizable similarity threshold , and retain near - synonyms.
[0083] (6) Formation of the overall label set: Continue to perform efficient key - value queries in the KV database for the near - synonyms corresponding to the target labels that cannot be precisely matched. If there are near - synonyms that can be precisely matched, then take the near - synonym with the largest and store it in the label set , where . The overall label set is .
[0084] (7) Label node expansion: When the label set contains only one label, there may be inaccurate stamp retrieval due to lack of semantics. To solve this problem, a label expansion strategy can be adopted. The specific method is: Select 2 to 3 labels with the highest (the value of the two - dimensional vector representing the node relationship in the knowledge graph) associated with this label and store them in the set , where , is the similarity weight of the target label.
[0085] (8) Formation of the final label set: Normalize all label weights to form the final label set .
[0086] 2. Introduction to stamp instantiation retrieval.
[0087] Stamp instantiation retrieval is the second step of this solution. Its core idea is to map the semantic labels in the final label set to specific stamp instances through the association relationships in the knowledge graph to form a stamp set, and return several most relevant stamps based on weight calculation. This module efficiently traverses the stamp - label association relationships in the knowledge graph and combines semantic similarity weights to ensure that the stamp retrieval results not only meet the user's query intent but also make full use of the semantic relevance of the stamp knowledge graph, thereby improving the accuracy, efficiency, and coverage of stamp retrieval.
[0088] 2.1 Obtaining the stamp set.
[0089] Stamp set acquisition is the first step of the stamp instantiation retrieval scheme. Its core idea is to use the association relationships in the knowledge graph to map the semantic tags in the final tag set to specific stamp instances, and construct a stamp set containing weight information to provide a data basis for subsequent Top-K sorting. The specific process is as follows: (1)Data input: The system receives the output result of the final tag set from the tag set acquisition module: , where represents the th tag in the output result of the final tag set, represents the th similarity weight of the tag, represents the total number of tags.
[0090] (2)Inverted index for accelerated positioning: According to the inverted index constructed in the KV database, the tag node can be quickly obtained through the hash table, and the tag node set is obtained, denoted as .
[0091] (3)Batch edge traversal: A pre-constructed adjacency list is used, where represents the tag entity node obtained according to the tag node , is the stamp entity node obtained according to the tag node , and are the and values of the two-dimensional vector in the knowledge graph edge relationship respectively. According to this adjacency list, the stamps associated with the tag (i.e., the first stamps) can be obtained in batch .
[0092] (4)Weight calculation optimization: A parallel computing framework is adopted to calculate the tag weights by strengthening the topic level: ; where, represents the tag weight between the tag and the stamp (i.e., the first stamp) , represents the value in the two-dimensional vector of the edge relationship between the tag and the stamp , represents the value in the two-dimensional vector of the edge relationship between the tag and the stamp . After that, the cumulative influence of all associated tags on the stamp can be aggregated by calculating multiple tags 。
[0093] (5) Incremental stamp set construction: Initialize an empty set and for each , adopt the following processing strategy: 1) If , 。
[0094] 2) If , , where is the set of associated tags of stamp , is the total aggregation weight of stamp (i.e., the stamp weight).
[0095] 1) The specific representation of the stamp is: 。
[0096] 2) Adjacency list structure: The adjacency list is a storage method of the graph data structure, which maintains a list of adjacent nodes for each node. In this embodiment, it is a pre-constructed adjacency relationship table of label node → [associated stamp] , where each associated record contains a triple ([[]] , topic level and weight ).
[0097] 3) Topic level enhancement: Topic level enhancement is a key adjustment factor in the weight calculation model, which is specifically used to amplify the contribution of the core tags (high-topic-level tags) in the stamp. Its mathematical expression is: ; where, represents the enhancement factor, and the logarithmic function ensures that the enhancement amplitude decays smoothly as increases.
[0098] 2.2 Optimization of stamp retrieval sorting.
[0099] This module uses the Bucket Sort algorithm to achieve efficient Top-K stamp retrieval. Its core idea is to pre-classify stamps into buckets in different intervals according to the weight distribution characteristics, and sort them with different algorithms, significantly reducing the sorting calculation amount. Compared with the traditional full-scale sorting algorithm, while ensuring the accuracy of the results, it can achieve a 5-8 times performance improvement. The specific process is as follows: (1) Introduction to bucket sorting: Bucket sorting is a non-comparative sorting algorithm, and its core idea follows the "divide and conquer - aggregation" paradigm. The specific stages are as follows: 1) Bucketing stage: Divide the data distribution range into several ordered buckets.
[0100] 2) Bucketing stage: Map the elements to the corresponding buckets according to their key values.
[0101] 3) In-bucket sorting: Perform local sorting on each non-empty bucket.
[0102] 4) Result merging: Concatenate all elements in the order of the buckets.
[0103] In the stamp weight scenario, the key value is normalized (with a value range of [0, 1]), which makes bucket sorting an ideal choice.
[0104] (2) Input data: Receive the incremental stamp set from the stamp set acquisition module.
[0105] (3) Weight distribution analysis: Generate a weight distribution histogram based on the input data. Using the weight as the horizontal axis, divide the interval [0, 1] into 100 equal-width bins, count the number of stamps falling into each bin, identify the high-density areas (e.g., 60% of the data is concentrated in the interval from 0.3 to 0.5) and sparse areas (e.g., high-weight stamps greater than 0.8 only account for 2%). Design non-uniform buckets based on the analysis results of the weight distribution histogram. For example, use narrow buckets (width of 0.05) in high-density areas to improve sorting accuracy, and wide buckets (width of 0.2) in sparse areas to reduce the probability of empty buckets.
[0106] (4) Dynamic bucketing: According to the weight distribution, in this embodiment, the stamps are divided into different buckets, which can be divided into four buckets: 1) High-value bucket: Store high-weight stamps (such as rare stamps), with a small quantity but the highest priority.
[0107] 2) Medium-value bucket: Store medium-weight stamps (such as popular theme stamps), with a moderate amount of data.
[0108] 3) Low-value bucket: Store low-weight stamps (such as ordinary stamps), with a large amount of data but the lowest priority.
[0109] 4) Discarded bucket: Directly filter and do not participate in sorting.
[0110] (5) Hierarchical sorting and result aggregation: In this embodiment, each bucket is processed according to the priority. The specific method is as follows: 1) High-value bucket: Since the amount of data is extremely small, insertion sorting can be used to ensure complete order and retain all results in full.
[0111] 2) Medium-value bucket: Use an optimized version of quicksort and only recursively process the partitions that may contain the Top-K items to reduce the computational amount.
[0112] 3) Low-value buckets: Only processed when the results of the first two buckets are less than K. Use partial sorting (such as heap selection) to extract the top K items.
[0113] When merging the results, this embodiment maintains a global Top-K list in real time. Once the cumulative quantity is greater than or equal to K, immediately check the maximum possible weight of the remaining buckets. If it cannot exceed the current Kth place, terminate the processing in advance, significantly reducing the computational amount.
[0114] After sorting is completed, output the formatted results, including stamp ID, name, picture, normalized weight, and core associated tags.
[0115] Among them: 1) Insertion sort: The insertion sort algorithm divides the data into two parts: sorted and unsorted. It inserts the unsorted elements into the correct position in the sorted part one by one, similar to the sorting method when arranging playing cards, and is suitable for efficient sorting of small sample sizes.
[0116] 2) Quick sort: Quick sort is an efficient divide-and-conquer sorting algorithm. Its core idea is to divide the sequence to be sorted into two parts through recursive partitioning: select a pivot value, move the elements smaller than the pivot value to its left, and those larger than the pivot value to its right. After forming two sub-sequences, repeat this process for the sub-sequences until they are ordered. Quick sort is especially suitable for processing large-scale random data due to its in-place sorting and cache friendliness.
[0117] 3) Heap sort: Heap sort is a comparison-based sorting algorithm based on the binary heap data structure, using the characteristics of the maximum heap or minimum heap of the heap for sorting. Heap sort is suitable for scenarios with large-scale data and no requirement for stability (such as priority queues, Top-K problems). Its core steps include: (1) Build the heap: Adjust the unordered array into a heap structure (usually start adjusting from the last non-leaf node to ensure that the parent node is greater than (or less than) the child nodes).
[0118] (2) Sorting: Repeatedly swap the top (maximum or minimum value) of the heap with the last element of the heap, shrink the heap range and readjust the heap until the entire array is ordered.
[0119] Compared with the prior art, the method of this embodiment has the following advantages: Through technologies such as large models, RAG, and KV databases, generate multi-level tags based on the characteristics of stamps to construct a knowledge graph rich in stamp characteristics, providing basic data for the rapid and accurate search of stamp information. Map the user's query intention to the tag nodes in the knowledge graph, and based on the multi-dimensional relationship characteristics of stamps - tags in the knowledge graph, achieve accurate and flexible instantiated stamp retrieval, thereby obtaining the target stamps and improving the accuracy of stamp retrieval.
[0120] This embodiment successfully solves the problems of the traditional solution, such as weak flexibility and high requirements for the stamp dataset. This embodiment not only improves the flexibility of stamp retrieval, but also successfully mines the connotative information of stamp data. In addition, the representation structure of the knowledge graph improves the interpretability of the model, making the content of the stamps and their label data more transparent and intuitive, facilitating understanding and analysis, and having broad application prospects.
[0121] Referring to Figure 6 , the embodiment of the present application also provides a stamp retrieval system based on a knowledge graph. The system includes a data extraction unit 100, a first construction unit 200, a second construction unit 300, a data acquisition unit 400, a label extension unit 500, a data merging unit 600, a third construction unit 700, and a stamp retrieval unit 800, where: The data extraction unit 100 is used to extract entity information and image information from the stamp image. The entity information is text information related to the stamp, and the image information is the image information on the stamp image; The first construction unit 200 is used to construct multi-level labels according to the entity information and image information, and assign first weights to labels at different levels; The second construction unit 300 is used to construct a KV database based on the stamp image, multi-level labels, and first weights, and construct a stamp knowledge graph according to the stamp entity data and label entity data in the KV database; The data acquisition unit 400 is used to extract multiple target labels from the text to be queried, and obtain a candidate label set by obtaining multiple candidate labels from the stamp knowledge graph according to the multiple target labels; The label extension unit 500 is used to perform label extension on the label in the candidate label set if the candidate label set contains only one label, to obtain an extended label set; The data merging unit 600 is used to merge the candidate label set and the extended label set to obtain a final label set; The third construction unit 700 is used to construct an incremental stamp set including the first stamp and the stamp weight based on the final label set; The stamp retrieval unit 800 is used to sort the first stamps in the incremental stamp set according to the stamp weight to obtain a stamp sorting result, so as to retrieve the target stamp according to the stamp sorting result.
[0122] It should be noted that since a stamp retrieval system based on a knowledge graph in this embodiment and the above-mentioned stamp retrieval method based on a knowledge graph are based on the same inventive concept, the corresponding content in the method embodiment also applies to this system embodiment and will not be elaborated here.
[0123] Referring to Figure 7, embodiments of the present application further provide an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the above-mentioned method for retrieving stamps based on a knowledge graph in the present disclosure.
[0124] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0125] The electronic device in the embodiments of the present application will be introduced in detail below.
[0126] The processor 1600 can be implemented in ways such as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure; The memory 1700 can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700, and the processor 1600 is called to execute the method for retrieving stamps based on a knowledge graph in the embodiments of the present disclosure.
[0127] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.
[0128] An embodiment of the present disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned stamp retrieval method based on a knowledge graph.
[0129] As a non-transitory computer-readable storage medium, a memory can be used to store a non-transitory software program and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0130] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0131] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0133] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0134] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0135] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0136] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0137] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0139] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs. The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present application.
[0140] The embodiments of the present application have been described in detail above with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present application.
Claims
1. A stamp retrieval method based on a knowledge graph, characterized in that, The method includes: Extracting entity information and image information from the stamp image, where the entity information is text information related to the stamp, and the image information is the image information on the stamp image; Constructing multi-level tags according to the entity information and the image information, and assigning a first weight to tags at different levels; Based on the stamp image, the multi-level tags, and the first weight, constructing a KV database, and constructing a stamp knowledge graph according to the stamp entity data and tag entity data in the KV database; Extracting multiple target tags from the text to be queried, and obtaining a candidate tag set by obtaining multiple candidate tags from the stamp knowledge graph according to the multiple target tags; If the candidate tag set contains only one tag, performing tag expansion on the tag in the candidate tag set to obtain an expanded tag set; Merging the candidate tag set and the expanded tag set to obtain a final tag set; Based on the final tag set, constructing an incremental stamp set including the first stamp and stamp weights; Sorting the first stamps in the incremental stamp set according to the stamp weights to obtain a stamp sorting result, so as to retrieve the target stamp according to the stamp sorting result.
2. The stamp retrieval method based on a knowledge graph according to claim 1, wherein The constructing multi-level tags according to the entity information and the image information includes: Constructing the entity information as first-level tags; Constructing the image information as second-level tags; Combining the first-level tags and the second-level tags to construct query information; Retrieving information fragments related to the query information from an external knowledge source to generate third-level tags, where the external knowledge source is a channel for obtaining information data from the outside.
3. The method for retrieving stamps based on a knowledge graph according to claim 1, wherein The obtaining a candidate tag set by obtaining multiple candidate tags from the stamp knowledge graph according to the multiple target tags includes: Matching the multiple target tags with the tags in the KV database to obtain multiple first matching results; Obtaining multiple first candidate tags in the stamp knowledge graph through the multiple first matching results to obtain an initial tag set; Performing similar semantic expansion on each unmatched target tag among the multiple target tags to obtain multiple similar semantic expansion words corresponding to each unmatched target tag; Matching the multiple similar semantic expansion words with the tags in the KV database to obtain multiple second matching results; Taking the similar semantic expansion word with the largest semantic similarity in the multiple second matching results as the second candidate tag corresponding to each unmatched target tag to obtain a second tag set; Merging the initial tag set and the second tag set to obtain a candidate tag set.
4. The stamp retrieval method based on a knowledge graph according to claim 3, characterized in that The performing similar semantic expansion on each unmatched target tag among the multiple target tags to obtain multiple similar semantic expansion words corresponding to each unmatched target tag includes: Generating a candidate synonym set corresponding to each unmatched target tag among the multiple target tags; Obtaining a first sememe set of each unmatched target tag, and obtaining a second sememe set of each candidate synonym in the candidate synonym set; Calculate the similarity between each sememe in the first sememe original set and all sememes in the second sememe set to obtain a first similarity result; According to the first similarity result, calculate the global similarity between each unmatched target label and its corresponding candidate near-synonym to obtain a second similarity result; Correct the second similarity result to obtain a third similarity result; According to the third similarity result, select multiple candidate near-synonyms as semantic similarity expansions to obtain multiple semantic similarity expansion words corresponding to each unmatched target label.
5. The method for retrieving stamps based on a knowledge graph according to claim 4, wherein The step of calculating the global similarity between each unmatched target label and its corresponding candidate near-synonym according to the first similarity result to obtain a second similarity result includes: ; Among them, represents the second similarity result, represents the set of first semantic primitives of the unmatched target label and represents the th semantic primitive in the set of first semantic primitives, represents the semantic primitive weight, represents the th semantic primitive in the set of second semantic primitives, represents the candidate near-synonym set of second semantic primitives, represents the first similarity result, represents finding the maximum similarity.
6. The stamp retrieval method based on a knowledge graph according to claim 1, characterized in that Based on the final label set, construct an incremental stamp set including first stamps and stamp weights, including: Obtain the label nodes of each label in the obtained final label set ; Determine the label nodes of each of the labels in the knowledge graph The corresponding label entity node or stamp entity node; Construct an adjacency list according to the label entity node and the stamp entity node; According to the adjacency list, obtain the first stamps associated with each label in the final label set; Calculate the label weights between each label in the final label set and its associated first stamp; Accumulate the label weights corresponding to the first stamps to obtain the stamp weights; Construct an incremental stamp set according to the first stamps and the stamp weights.
7. The method for retrieving stamps based on a knowledge graph according to claim 6, wherein, The step of calculating the label weights between each label in the final label set and its associated first stamp includes: ; Among them, represents the th label in the final label set and the weight of the label between it and the first stamp represents the similarity weight of the th label, represents the label and the weight value in the two-dimensional vector of the edge relationship between the first stamp, represents the logarithmic function, represents the label and the label level in the two-dimensional vector of the edge relationship between the first stamp.
8. A stamp retrieval system based on a knowledge graph, characterized in that, The system includes: A data extraction unit for extracting entity information and image information in the stamp image, where the entity information is text information related to the stamp, and the image information is image information on the stamp image; A first construction unit for constructing multi-level labels according to the entity information and the image information and assigning first weights to labels at different levels; A second construction unit for constructing a KV database based on the stamp image, the multi-level labels, and the first weights, and constructing a stamp knowledge graph according to the stamp entity data and label entity data in the KV database; A data acquisition unit for extracting multiple target labels in the text to be queried, and obtaining a candidate label set by obtaining multiple candidate labels from the stamp knowledge graph according to the multiple target labels; A label expansion unit for expanding the label in the candidate label set to obtain an expanded label set if the candidate label set contains only one label; A data merging unit for merging the candidate label set and the expanded label set to obtain a final label set; A third construction unit for constructing an incremental stamp set including first stamps and stamp weights based on the final label set; A stamp retrieval unit for sorting the first stamps in the incremental stamp set according to the stamp weights to obtain a stamp sorting result, so as to retrieve the target stamp according to the stamp sorting result.
9. An electronic device, characterized in that, Comprising at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the knowledge graph-based stamp retrieval method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the knowledge graph-based stamp retrieval method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge reasoning-based big data service label extension method and system
CN111737400A
Method and system for keyword search over a knowledge graph
CN113535890A
Quick stamp positioning and retrieval algorithm based on 2D image
CN116311339A
Label storage method and device for database, terminal equipment and medium
CN116541393A
Multi-intention recognition method and device, storage medium and equipment
CN117744667A