Metadata-driven book searching and recommending method based on virtual knowledge graph
By using a metadata-driven approach based on virtual knowledge graphs to generate temporary virtual nodes and perform semantic associations, the problems of frequent information updates and insufficient semantic understanding in library recommendation systems are solved, enabling efficient and personalized book search and recommendation.
Patent Information
- Application Number
- CN202511167147.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Existing library recommendation systems rely on user behavior data, making it difficult to achieve deep associations. Frequent updates to the knowledge graph lead to low search accuracy and insufficient semantic understanding capabilities.
The metadata-driven approach based on virtual knowledge graphs generates temporary virtual nodes by receiving query requests, uses semantic vectors to calculate and establish connections between nodes, labels relationship weights, constructs triples, and performs precise cross-source entity association. The recommendation engine generates recommendations based on multi-hop path reasoning.
Significantly reduces memory usage, improves response speed and recommendation accuracy, reduces server load, and enables personalized recommendations.
Smart Images

Figure CN120994712A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a metadata-driven book search and recommendation method based on virtual knowledge graphs. Background Technology
[0002] The rapid development of artificial intelligence technology has injected strong momentum into innovation across various industries. However, current library recommendation systems still primarily rely on user behavior data, with book-related information such as authors, publishers, and sales figures scattered across various databases, making deep correlation difficult. While some libraries have begun constructing knowledge graphs, the frequent updates to library book information make it difficult for existing knowledge graphs to adapt to changes in dynamic metadata. Furthermore, insufficient semantic understanding capabilities also lead to low search accuracy; for example, searching for "Introduction to Quantum Physics" fails to effectively correlate with related content such as "Fundamentals of Quantum Mechanics." Summary of the Invention
[0003] This invention addresses the technical problems existing in the prior art by providing a metadata-driven book search and recommendation method based on virtual knowledge graphs.
[0004] The technical solution of this invention to solve the above-mentioned technical problems is as follows: a metadata-driven book search and recommendation method based on virtual knowledge graphs, comprising the following steps: S101. Receive the query request, trigger the aggregation of multi-source metadata, generate temporary virtual nodes, and establish the relationship edges between nodes by calculating semantic vectors and label the relationship weights. S102. Virtual nodes are automatically released after a query response. During a second query, cached metadata is reused. Virtual node attributes are classified according to stability and marked with expiration time. S103. Parse the query into triples, construct a citation chain using the simple knowledge organization system vocabulary and citation data, and achieve accurate cross-source entity association. S104. The search engine returns matching results through semantic parsing, subgraph generation, and association strength, while the recommendation engine generates recommendations based on multi-hop path reasoning.
[0005] In a preferred embodiment, in step S101, a query request input by the user is received and limited to ISBN, book title, and keywords. The parsing unit uses a differentiated processing mechanism to parse the query request. For ISBN type, its standard format is matched using regular expressions. For book title type, a string fuzzy matching algorithm is used to handle input errors. For keyword type, core search identifiers are extracted using topic term extraction technology. All parsing results are converted into standardized search identifiers. Based on the standardized search identifiers, a distributed metadata collector is triggered to perform real-time aggregation. The data sources are divided into three categories, including basic data sources, dynamic data sources, and related data sources. The metadata collector connects to each data source through a standardized API interface to ensure data format uniformity. A temporary virtual node is created for each target book. The virtual node is stored in memory as a key-value pair data structure. The aggregated metadata is encapsulated into a temporary structured node and is not written to the physical storage device. The virtual node is generated separately for each target book and includes three types of attributes: basic attributes, semantic attributes, and derived attributes. The basic attributes cover ISBN, book title, author, publisher, and publication time. The basic attributes are used as the core identifier of the node. The semantic attributes include subject terms, abstract keywords, and subject classification to support semantic association. The derived attributes include user rating, price fluctuation value, and related data recommendation degree to reflect the dynamic characteristics and association potential of the book. Dynamic links are established by automatically constructing relational edges between virtual nodes through semantic vector computation. These dynamic links are implemented based on two types of operations: entity association and relation annotation. Entity association quantifies the semantic association strength between the book node and other entity nodes through semantic vector computation. The specific calculation formula is as follows: , in, Cosine similarity is used to determine whether an entity association exists. The semantic vector representing the book node. The semantic vector representing other entity nodes, Representing vectors The i-th component, Representing vectors The i-th component, where k represents the vector dimension. It represents the dot product of two vectors, automatically links book nodes with other entity nodes based on semantic association strength, explicitly marks the relationship type on the associated edges, and dynamically assigns relationship weights based on co-occurrence frequency, and quantifies the association strength based on the relationship weights.
[0006] In a preferred embodiment, in step S102, the virtual node is used as a temporary computing entity to support word query processing and result return. After the query response is completed, the node is automatically released. The automatic release of the node is triggered by the following conditions: after the user obtains the search and recommendation results, the automatic release process is started after a five-second delay, clearing all attribute information of the virtual node in memory without any physical storage. For secondary query scenarios, a caching mechanism is used to avoid repeated collection of metadata. During the first query, the metadata on which the virtual node depends is temporarily stored in the memory cache. When the secondary query is triggered, the metadata in the cache is directly called to regenerate the virtual node without having to send a collection request to the multi-source data source again. Virtual node attributes are categorized by stability and labeled with expiration dates. Virtual node attributes are divided into two categories based on stability: static attributes and dynamic attributes. The expiration dates of the attributes are labeled accordingly. Static attributes refer to metadata that remains unchanged for a long time, including author, ISBN, publisher, and publication time. These are marked as long-term valid and are only generated during the initial metadata collection. Dynamic attributes refer to metadata that may change in a short time, including price, user rating, and inventory status. These are marked as short-term valid. The expiration date labels are embedded in the attribute fields of the virtual node in the form of key-value pairs. When the remaining validity period of a dynamic attribute is insufficient, a new round of metadata collection request is automatically initiated from the corresponding data source to obtain the latest attribute value. After collection, if the newly obtained attribute value differs from the original attribute value, the corresponding attribute of the virtual node is immediately updated and the validity period is re-marked. If there is no significant difference, the original attribute value is maintained but the validity period mark is refreshed.
[0007] In a preferred embodiment, in S103, the user query request is parsed into a triplet of entity, attribute, and relation. The triplet is constructed using a semantic parsing model and expands to synonyms and near-synonyms based on a simple knowledge organization system lexicon. The expansion logic is as follows: when the subject term of a book node matches a concept in the simple knowledge organization system lexicon, it automatically associates all the subject terms corresponding to the synonyms and near-synonyms of that concept. Through the citation data of the academic database, a cited-reference relationship chain is constructed. The relationship chain is based on the citation records of the academic database. The citation records contain the citation information of the book, the citation position, and the citation purpose. The association logic is to construct a cited association edge for each book citation record and mark the citation strength.
[0008] In a preferred embodiment, in S104, the triples generated by entity relationship identification are mapped to the node attributes of the virtual knowledge graph to establish a semantic association between the query intent and the graph structure. The mapping process must ensure semantic consistency. Field matching is achieved through a preset attribute mapping table. A subgraph containing the target book and its directly associated nodes is dynamically constructed. The generation range of the subgraph is defined by the direct association of the query entity. For a single book entity, the subgraph contains the book node and its directly associated entity nodes. For a category entity, the subgraph contains all book nodes under that category and their common associated high-frequency entity nodes. The virtual subgraph is traversed using a breadth-first search algorithm to filter out all book nodes that satisfy the triplet constraint. The execution logic of the breadth-first search algorithm is as follows: starting from the core entity node of the subgraph, the adjacent nodes are traversed layer by layer, and the attributes of each node are checked to determine whether they satisfy the attribute constraints of the triplet. If they are satisfied, they are marked as matching nodes. Matching nodes are sorted in descending order based on association strength. By quantifying the degree of association between nodes and queries, the most relevant books are presented first. The specific formula for calculating association strength is as follows: , Where S represents the association strength, M represents the attribute matching degree, W represents the relationship weight, and T represents the timeliness of dynamic attributes. Indicates the attribute matching weight. The weights that represent the relation weights The weights representing the timeliness of dynamic attributes are used to achieve personalized recommendations based on the association paths of the virtual graph. By mining multi-hop association paths in the virtual graph, candidate books that are indirectly related to the target book are discovered. Multi-hop traversal is performed based on the association edges of the virtual graph. Typical paths include author association paths, topic association paths, and citation association paths. Through multi-path cross-validation, candidate books with stable associations to the target book are selected, providing a candidate pool for personalized recommendations.
[0009] The beneficial effects of this invention are as follows: This invention achieves low-latency data collection through API interfaces, updates dynamic attributes in real time according to short-term effective rules, adopts a fast automatic release and reasonable cache cycle reuse mechanism for virtual nodes, significantly reducing memory consumption, and improves response speed for secondary queries by reusing cached metadata, effectively reducing server load and avoiding resource redundancy. It also improves the accuracy of recommendations by mining deep relationships through cross-source links and path reasoning. Attached Figure Description
[0010] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0012] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0013] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0014] like Figure 1 This embodiment provides a metadata-driven book search and recommendation method based on virtual knowledge graphs, including the following steps: S101. Receive the query request, trigger the aggregation of multi-source metadata, generate temporary virtual nodes, and establish the relationship edges between nodes by calculating semantic vectors and label the relationship weights. Furthermore, the system receives user-input query requests, limiting them to ISBN, book title, and keywords. The parsing unit employs a differentiated processing mechanism to parse the query requests. For ISBN types, it matches their standard format using regular expressions; for book title types, it uses a string fuzzy matching algorithm to handle input errors; and for keyword types, it extracts core search identifiers using topic term extraction technology. All parsing results are converted into standardized search identifiers. Based on these standardized search identifiers, a distributed metadata collector is triggered to perform real-time aggregation. The data sources are divided into three categories: basic data sources, dynamic data sources, and related data sources. The metadata collector connects to each data source through a standardized API interface to ensure data format uniformity. A temporary virtual node is created for each target book. The virtual node is stored in memory as a key-value pair data structure. The aggregated metadata is encapsulated into a temporary structured node and is not written to the physical storage device. The virtual node is generated separately for each target book and includes three types of attributes: basic attributes, semantic attributes, and derived attributes. The basic attributes cover ISBN, book title, author, publisher, and publication time. The basic attributes are used as the core identifier of the node. The semantic attributes include subject terms, abstract keywords, and subject classification to support semantic association. The derived attributes include user rating, price fluctuation value, and related data recommendation degree to reflect the dynamic characteristics and association potential of the book. Dynamic links are established by automatically constructing relational edges between virtual nodes through semantic vector computation. These dynamic links are implemented based on two types of operations: entity association and relation annotation. Entity association quantifies the semantic association strength between the book node and other entity nodes through semantic vector computation. The specific calculation formula is as follows: , in, Cosine similarity is used to determine whether an entity association exists. The semantic vector representing the book node. The semantic vector representing other entity nodes, Representing vectors The i-th component, Representing vectors The i-th component, where k represents the vector dimension. It represents the dot product of two vectors, automatically links book nodes with other entity nodes based on semantic association strength, explicitly marks the relationship type on the associated edges, and dynamically assigns relationship weights based on co-occurrence frequency, and quantifies the association strength based on the relationship weights.
[0015] It should be noted that the basic data sources include authoritative identification information such as ISBN, author, and publication date provided by the library's OPAC system; content summaries and tables of contents provided by the publisher database to ensure the accuracy of the basic attributes of the books; dynamic data sources include the current price and inventory status returned in real time by e-commerce platforms; user ratings and tag clouds provided by the Yudu community to reflect the real-time market feedback of the books; and related data sources provide citation relationships through academic databases, recent discussion popularity through social platforms, and click and purchase-related data from user behavior logs to support the correlation analysis of subsequent recommendations. The establishment of dynamic links enables the scattered virtual nodes to form a network-like relational structure, which not only preserves the semantic logic between entities but also adapts to changes in the relational relationship through dynamic adjustment of weights (such as the association between new books and classic books being strengthened as user behavior increases).
[0016] S102. Virtual nodes are automatically released after a query response. During a second query, cached metadata is reused. Virtual node attributes are classified according to stability and marked with expiration time. Furthermore, virtual nodes are used as temporary computing entities to support word search processing and result return. After the query response is completed, the nodes are automatically released. The automatic release of nodes is triggered by the following conditions: after the user obtains the search and recommendation results, the automatic release process is started after a five-second delay, clearing all attribute information (including basic attributes, semantic attributes, and derived attributes) of the virtual node in memory, without performing any physical storage. For secondary query scenarios (such as when the user performs pagination, filtering, or sorting operations), a caching mechanism is used to avoid repeated collection of metadata. During the first query, the metadata on which the virtual node depends is temporarily stored in the memory cache. When a secondary query is triggered, the metadata in the cache is directly called to regenerate the virtual node without having to send collection requests to multiple data sources again. Virtual node attributes are categorized by stability and labeled with expiration dates. Virtual node attributes are divided into two categories based on stability: static attributes and dynamic attributes. The expiration dates of the attributes are labeled accordingly. Static attributes refer to metadata that remains unchanged for a long time, including author, ISBN, publisher, and publication time. These are marked as long-term valid and are only generated during the initial metadata collection. Dynamic attributes refer to metadata that may change in a short time, including price, user rating, and inventory status. These are marked as short-term valid. The expiration date labels are embedded in the attribute fields of the virtual node in the form of key-value pairs. When the remaining validity period of a dynamic attribute is insufficient (e.g., less than ten seconds), a new round of metadata collection request is automatically initiated to the corresponding data source to obtain the latest attribute value (e.g., the current real-time price, the updated user rating). After collection, if the newly obtained attribute value differs from the original attribute value, the corresponding attribute of the virtual node is immediately updated and the validity period is re-marked. If there is no significant difference, the original attribute value is maintained but the validity period mark is refreshed.
[0017] It should be noted that the coordinated operation of the expiration mark and node release mechanism forms a closed-loop management of the entire lifecycle of virtual nodes. This ensures data validity and achieves optimal resource allocation. The long-term validity of static attributes is compatible with the node release mechanism. Since static attributes do not need to be updated frequently, their cache validity period (30 minutes) covers the secondary query window after node release, ensuring that accurate static information can still be reused when a second node is generated. The short-term validity of dynamic attributes is linked with the node release mechanism through the update mechanism. If a virtual node is released within the validity period of dynamic attributes, the dynamic metadata in the cache will be marked as pending update synchronously with the node release. When a secondary query triggers node reconstruction, the latest dynamic attributes will be collected first to avoid using expired cache.
[0018] S103. Parse the query into triples, construct a citation chain using the simple knowledge organization system vocabulary and citation data, and achieve accurate cross-source entity association. Furthermore, user query requests are parsed into entity, attribute, and relation triples. These triples are constructed using a semantic parsing model and expanded to include synonyms and near-synonyms based on a simple knowledge organization system lexicon. The expansion logic is as follows: when the subject term of a book node matches a concept in the simple knowledge organization system lexicon, it automatically associates all the subject terms corresponding to that concept's synonyms and near-synonyms. Through citation data from academic databases, a cited-reference relationship chain is constructed. The foundation of this relationship chain is the citation records in the academic database, which contain the book's citation information, citation location, and citation purpose. The association logic is to construct a cited association edge for each book citation record and label the citation strength.
[0019] It should be noted that the construction of the triples follows strict semantic logic. The entity refers to the core object involved in the query, the attribute refers to the characteristic dimension of the entity, and the relation refers to the way the entity and the attribute are associated. The role of the relation chain is to reveal the knowledge connection between books, so that when querying related books, the relation chain can recommend related books to be cited. This provides a complete resource network for academic research from basic to applied. The relation chain constructs book associations from three dimensions: author organization, topic semantics, and knowledge inheritance. Together with the triples of entity relation recognition, it forms a multi-dimensional virtual graph association network.
[0020] S104. The search engine returns matching results through semantic parsing, subgraph generation, and association strength, while the recommendation engine generates recommendations based on multi-hop path reasoning. Furthermore, the triples generated by entity relationship recognition are mapped to the node attributes of the virtual knowledge graph to establish a semantic association between query intent and graph structure. The mapping process must ensure semantic consistency. Precise field matching is achieved through a preset attribute mapping table. Subgraphs containing target books and directly related nodes are dynamically constructed. The generation scope of subgraphs is defined by the direct association of the query entity. For a single book entity, the subgraph contains the book node and its directly related entity nodes. For a category entity, the subgraph contains all book nodes under that category and their common high-frequency entity nodes. The virtual subgraph is traversed using a breadth-first search algorithm to filter out all book nodes that satisfy the triplet constraint. The execution logic of the breadth-first search algorithm is as follows: starting from the core entity node of the subgraph, the adjacent nodes are traversed layer by layer (the first layer is the directly related nodes, and the second layer is the related nodes of the related nodes). The attributes of each node are checked to determine whether they satisfy the attribute constraints of the triplet. If they satisfy both, they are marked as matching nodes. Matching nodes are sorted in descending order based on association strength. By quantifying the degree of association between nodes and queries, the most relevant books are presented first. The specific formula for calculating association strength is as follows: , Where S represents the association strength, M represents the attribute matching degree, W represents the relationship weight, and T represents the timeliness of dynamic attributes. Indicates the attribute matching weight. The weights that represent the relation weights The weights representing the timeliness of dynamic attributes are used to achieve personalized recommendations based on the association paths of the virtual graph. By mining multi-hop association paths in the virtual graph, candidate books that are indirectly related to the target book are discovered. Multi-hop traversal is performed based on the association edges of the virtual graph. Typical paths include author association paths, topic association paths, and citation association paths. Through multi-path cross-validation, candidate books with stable associations to the target book are selected, providing a candidate pool for personalized recommendations.
[0021] It should be noted that the matching process needs to balance completeness and efficiency: the traversal depth of the breadth-first search algorithm is limited to 2 layers (to avoid overexpansion leading to redundant calculations), and the adjacent nodes of the matched nodes are checked first (such as books by the same author of a matched computer textbook). The pruning strategy (such as skipping nodes that have been checked but do not meet the conditions) is used to reduce repeated calculations. The output of the subgraph matching is the set of all nodes that meet the constraints, which is the direct input for sorting the results.
[0022] It should be noted that the author association path is: Book A → same author → other works → Book B; the topic association path is: Book C → highly similar topics → related topic books → Book D; and the citation association path is: Book E → frequently cited → cited books → Book F.
[0023] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0024] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0025] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0026] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0027] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0028] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0029] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A metadata-driven book search and recommendation method based on virtual knowledge graphs, characterized in that, Includes the following steps: S101. Receive the query request, trigger the aggregation of multi-source metadata, generate temporary virtual nodes, and establish the relationship edges between nodes by calculating semantic vectors and label the relationship weights. S102. Virtual nodes are automatically released after a query response. During a second query, cached metadata is reused. Virtual node attributes are classified according to stability and marked with expiration time. S103. Parse the query into triples, construct a citation chain using the simple knowledge organization system vocabulary and citation data, and achieve accurate cross-source entity association. S104. The search engine returns matching results through semantic parsing, subgraph generation, and association strength, while the recommendation engine generates recommendations based on multi-hop path reasoning.
2. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 1, characterized in that, In S101, a query request is received from the user and limited to ISBN, book title, and keywords. The parsing unit uses a differentiated processing mechanism to parse the query request. For ISBN type, its standard format is matched using regular expressions. For book title type, a string fuzzy matching algorithm is used to handle input errors. For keyword type, core search identifiers are extracted using topic term extraction technology. All parsing results are converted into standardized search identifiers. Based on the standardized search identifiers, the distributed metadata collector is triggered to perform real-time aggregation. The data sources are divided into three categories, including basic data sources, dynamic data sources, and related data sources. The metadata collector connects with each data source through a standardized API interface to ensure data format uniformity. A temporary virtual node is created for each target book. The virtual node is stored in memory as a key-value pair data structure. The aggregated metadata is encapsulated into a temporary structured node and is not written to the physical storage device. The virtual node is generated separately for each target book and includes three types of attributes: basic attributes, semantic attributes, and derived attributes. The basic attributes cover ISBN, book title, author, publisher, and publication time. The basic attributes serve as the core identifier of the node. The semantic attributes include subject terms, abstract keywords, and subject classification to support semantic association. The derived attributes include user ratings, price fluctuation values, and related data recommendation scores to reflect the dynamic characteristics and association potential of the book.
3. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 2, characterized in that, Dynamic links are established by automatically constructing relational edges between virtual nodes through semantic vector computation. These dynamic links are implemented based on two types of operations: entity association and relation annotation. Entity association quantifies the semantic association strength between the book node and other entity nodes through semantic vector computation. The specific calculation formula is as follows: , in, Cosine similarity is used to determine whether an entity association exists. The semantic vector representing the book node. The semantic vector representing other entity nodes, Representing vectors The i-th component, Representing vectors The i-th component, where k represents the vector dimension. It represents the dot product of two vectors, automatically links book nodes with other entity nodes based on semantic association strength, explicitly marks the relationship type on the associated edges, and dynamically assigns relationship weights based on co-occurrence frequency, and quantifies the association strength based on the relationship weights.
4. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 1, characterized in that, In S102, virtual nodes are used as temporary computing entities to support word search processing and result return. After the query response is completed, the nodes are automatically released. The automatic release of nodes is triggered by the following conditions: after the user obtains the search and recommendation results, the automatic release process is started after a five-second delay, clearing all attribute information of the virtual node in memory without any physical storage. For secondary query scenarios, a caching mechanism is used to avoid repeated collection of metadata. During the first query, the metadata on which the virtual node depends is temporarily stored in the memory cache. When a secondary query is triggered, the cached metadata is directly called to regenerate the virtual node without having to send a collection request to the multi-source data source again.
5. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 4, characterized in that, Virtual node attributes are categorized by stability and labeled with expiration dates. Virtual node attributes are divided into two categories based on stability: static attributes and dynamic attributes. The expiration dates of the attributes are labeled accordingly. Static attributes refer to metadata that remains unchanged for a long time, including author, ISBN, publisher, and publication time. These are marked as long-term valid and are only generated during the initial metadata collection. Dynamic attributes refer to metadata that may change in a short time, including price, user rating, and inventory status. These are marked as short-term valid. The expiration date labels are embedded in the attribute fields of the virtual node in the form of key-value pairs. When the remaining validity period of a dynamic attribute is insufficient, a new round of metadata collection request is automatically initiated from the corresponding data source to obtain the latest attribute value. After collection, if the newly obtained attribute value differs from the original attribute value, the corresponding attribute of the virtual node is immediately updated and the validity period is re-marked. If there is no significant difference, the original attribute value is maintained but the validity period mark is refreshed.
6. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 1, characterized in that, In S103, user query requests are parsed into triples of entity, attribute, and relation. These triples are constructed using a semantic parsing model and expanded to include synonyms and near-synonyms based on a simple knowledge organization system lexicon. The expansion logic is as follows: when the subject term of a book node matches a concept in the simple knowledge organization system lexicon, it automatically associates all the subject terms corresponding to that concept's synonyms and near-synonyms. Through citation data from the academic database, a cited-reference relationship chain is constructed. The foundation of the relationship chain is the citation records in the academic database, which contain the book's citation information, citation location, and citation purpose. The association logic is to construct a cited association edge for each book citation record and mark the citation strength.
7. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 1, characterized in that, In S104, the triples generated by entity relationship recognition are mapped to the node attributes of the virtual knowledge graph to establish a semantic association between the query intent and the graph structure. The mapping process must ensure semantic consistency. Precise field matching is achieved through a preset attribute mapping table. Subgraphs containing the target book and its directly associated nodes are dynamically constructed. The generation scope of the subgraph is defined by the direct association of the query entity. For a single book entity, the subgraph contains the book node and its directly associated entity nodes. For a category entity, the subgraph contains all book nodes under that category and their common high-frequency entity nodes.
8. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 7, characterized in that, The virtual subgraph is traversed using a breadth-first search algorithm to filter out all book nodes that satisfy the triplet constraint. The execution logic of the breadth-first search algorithm is as follows: starting from the core entity node of the subgraph, the adjacent nodes are traversed layer by layer, and the attributes of each node are checked to determine whether they satisfy the attribute constraints of the triplet. If they are satisfied, they are marked as matching nodes.
9. The metadata-driven book search and recommendation method based on virtual knowledge graphs according to claim 7, characterized in that, Matching nodes are sorted in descending order based on association strength. By quantifying the degree of association between nodes and queries, the most relevant books are presented first. The specific formula for calculating association strength is as follows: , Where S represents the association strength, M represents the attribute matching degree, W represents the relationship weight, and T represents the timeliness of dynamic attributes. Indicates the attribute matching weight. The weights that represent the relation weights The weights representing the timeliness of dynamic attributes are used to achieve personalized recommendations based on the association paths of the virtual graph. By mining multi-hop association paths in the virtual graph, candidate books that are indirectly related to the target book are discovered. Multi-hop traversal is performed based on the association edges of the virtual graph. Typical paths include author association paths, topic association paths, and citation association paths. Through multi-path cross-validation, candidate books with stable associations to the target book are selected, providing a candidate pool for personalized recommendations.
Citation Information
Patent Citations
Virtual knowledge graph construction method and device
CN111475503A
Method and system for improving document file retrieval efficiency based on knowledge graph
CN113221562A
Document book semantic retrieval system based on knowledge graph
CN115563313A
Document retrieval method and system based on knowledge graph, terminal and storage medium
CN116881436A
Database management method and related device
CN119127836A