A knowledge graph-based publishing field knowledge service construction method

By performing multi-granular content structure analysis and semantic graph construction on publishing resources, the problem of loose semantic association and content structure mapping in digital publishing is solved, realizing the orderly transformation of knowledge service paths and improving the coherence of content generation.

CN122173661APending Publication Date: 2026-06-09TIANWEN DIGITAL MEDIA TECH HUNAN
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610647711.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies in digital publishing lack tight semantic association and content structure mapping, resulting in insufficient knowledge service path organization capabilities and difficulty in meeting the needs of complex service scenarios.

Method used

By performing multi-granularity content structure analysis on the original publishing resources, a set of publishing content structure nodes is generated. Combined with semantic information extraction, a publishing semantic graph is constructed, forming a multi-granularity anchor point dual graph. This enables the collaborative mapping of semantic relationships and content carrying locations, identifies user intent, generates a graph service execution plan, and optimizes service path generation.

Benefits of technology

It enhances the precision and service orientation of knowledge organization, improves the coherence and adaptability of knowledge service content generation, and realizes the orderly transformation of service needs into retrieval paths and organization methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173661A_ABST
    Figure CN122173661A_ABST
Patent Text Reader

Abstract

The application discloses a kind of publishing field knowledge service construction methods based on knowledge graph, it is related to intelligent digital publishing technical field, including, according to publishing content structure node set, constructs publishing content structure diagram, and in combination with publishing original resource set executes publishing field semantic information extraction and semantic correlation analysis, obtains publishing semantic diagram;From the structural distribution characteristics and content anchoring characteristics of semantic node in publishing resource in publishing semantic diagram are extracted, and the anchor point edge between publishing semantic diagram and publishing content structure diagram is established, and the multi-granularity anchor point double diagram is formed;Based on graph service execution plan, execute double diagram collaborative retrieval and service path generation, obtain knowledge service candidate set, and carry out service ordering and arrangement, generate publishing field knowledge service content.The application realizes the collaborative mapping of semantic relationship and content bearing position, enhances knowledge organization precision and service direction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent digital publishing technology, and in particular to a method for constructing knowledge services in the publishing field based on knowledge graphs. Background Technology

[0002] Against the backdrop of the deep integration of digital publishing and knowledge services, in order to meet the needs of knowledge organization and service for massive and heterogeneous publishing content, conventional methods usually adopt a technical approach that combines bibliographic metadata integration, full-text retrieval, subject heading indexing, classification organization, and knowledge graph modeling. The aim is to structurally represent elements such as books, chapters, authors, core concepts, and citation relationships, and construct a knowledge graph model with semantic relationship networks at its core. Through the knowledge graph model, basic services can be provided for scenarios such as retrieval, navigation, question answering, and related recommendations, realizing the evolution of publishing content from document management to knowledge association services.

[0003] Conventional methods still have limited ability to coordinate the construction of semantic associations and the structure of published content. There is a lack of close mapping between semantic nodes and chapter levels, chart carriers and content carrier positions, which affects knowledge positioning and the connection of structured services. Natural language request processing mostly stays at the level of general parsing and result retrieval, and there is insufficient coordination of service goals, service objects, constraints and output requirements, which limits the ability to organize paths in complex service scenarios. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a knowledge graph-based method for constructing knowledge services in the publishing field to address the problems of weak semantic association and content structure mapping in publishing, as well as insufficient organization of knowledge service paths driven by service intent.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] This invention provides a method for constructing knowledge services in the publishing field based on knowledge graphs, which includes: performing multi-granularity content structure analysis on the original set of publishing resources, identifying the hierarchical organization of content and content-carrying objects in the publishing resources, and generating a set of publishing content structure nodes;

[0008] Based on the set of structural nodes of the published content, a structural graph of the published content is constructed, and semantic information extraction and semantic association analysis in the publishing domain are performed in combination with the original set of publishing resources to obtain a semantic graph of the publishing.

[0009] Extract the structural distribution features and content anchoring features of semantic nodes in publishing resources from the publishing semantic graph, and establish anchor edges between the publishing semantic graph and the publishing content structure graph to form a multi-granularity anchor point bi-graph;

[0010] The system performs service intent recognition and slot extraction on user natural language requests, obtains the intent slot structure, maps it to a multi-granularity anchor point dual graph, and generates a graph service execution plan.

[0011] Based on the graph service execution plan, the system performs dual-graph collaborative retrieval and service path generation to obtain a knowledge service candidate set, and then performs service-oriented sorting and arrangement to generate knowledge service content in the publishing field.

[0012] As a preferred embodiment of the knowledge graph-based method for constructing knowledge services in the publishing field according to the present invention, the steps for generating the publishing content structure node set are as follows:

[0013] Extract directory hierarchy information, page layout information, and content boundary information that characterize the organization of publishing resources from the original publishing resources to generate a publishing content structure analysis framework;

[0014] Based on the content structure analysis framework, the publishing resources are hierarchically divided, the hierarchical and sequential adjacency relationships between the publishing content at each level are identified, and a hierarchical relationship chain of publishing content is generated.

[0015] Based on the hierarchical relationship chain of publishing content, objectify the publishing content at each level, determine the content carrier, location, subordinate path and sequence relationship of each level of publishing content, and generate a set of publishing content object identifiers;

[0016] Based on the object identifier set of published content and the hierarchical relationship chain of published content, structural node identifiers are assigned to each level of published content and its carrier object, and node attribute descriptions are established to generate a set of structural nodes for published content.

[0017] As a preferred embodiment of the knowledge graph-based knowledge service construction method for the publishing field described in this invention, the steps for constructing the publishing content structure graph are as follows:

[0018] Based on the location, subordinate path, and sequence relationship of each structural node in the publishing content structure node set, establish the structural mapping relationship between each structural node and generate the publishing content structure networking rules;

[0019] Based on the network rules for publishing content structure, the structural nodes in the publishing content structure node set are connected in a graph to generate a publishing content structure graph.

[0020] As a preferred embodiment of the knowledge graph-based method for constructing knowledge services in the publishing field according to the present invention, the steps for obtaining the publishing semantic graph are as follows:

[0021] Based on the original set of publishing resources and combined with the structural positioning information in the publishing content structure diagram, semantic information of the publishing domain is extracted from the publishing content at each level to obtain publishing semantic elements;

[0022] Using the structural positioning information in the publication content structure diagram, semantic association analysis is performed on the semantic associations between various publishing semantic elements to determine the semantic association direction, association type and association strength, and to generate a publishing semantic association chain.

[0023] The semantic elements of publishing are graphed into semantic nodes, and the semantic associations of publishing in the semantic association chain are graphed into semantic edges, thus generating a semantic graph of publishing.

[0024] As a preferred embodiment of the knowledge graph-based knowledge service construction method for the publishing field described in this invention, the steps for extracting the structural distribution features and content anchoring features of semantic nodes in publishing resources are as follows:

[0025] The semantic nodes in the publishing semantic graph are backtracked and matched with the corresponding content fragments in the original publishing resource set. The occurrence position, distribution frequency and hierarchical landing point of each semantic node in the publishing content at each level are counted to generate a semantic node structure distribution spectrum.

[0026] By performing alignment analysis between semantic nodes and corresponding structural nodes in the publication content structure diagram, the content attachment position and structural pointing position of each semantic node in the publication content structure diagram are determined, and a semantic node anchoring feature table is generated.

[0027] As a preferred embodiment of the knowledge graph-based knowledge service construction method for the publishing field described in this invention, the steps for forming a multi-granularity anchor point dual graph are as follows:

[0028] The semantic node structure distribution spectrum and the semantic node anchoring feature table are associated and matched to determine the target structure node and anchoring connection method corresponding to each semantic node, and generate a dual-graph anchor point mapping chain.

[0029] Based on the dual-graph anchor mapping chain, anchor edges are established between semantic nodes and target structural nodes, and the anchor edges, publishing semantic graph and publishing content structural graph are associated and fused to generate a multi-granularity anchor dual-graph.

[0030] As a preferred embodiment of the knowledge graph-based knowledge service construction method for the publishing field described in this invention, the steps for obtaining the intent slot structure are as follows:

[0031] Semantic parsing and element decomposition are performed on user natural language requests, and request semantic elements that represent service needs are extracted;

[0032] Semantic classification and slot mapping are performed on the semantic elements of the request to determine the service goals, service objects, constraints and output requirements in the user's natural language request, and to generate the intent slot structure.

[0033] As a preferred embodiment of the knowledge graph-based knowledge service construction method for the publishing field described in this invention, the steps for generating the knowledge graph service execution plan are as follows:

[0034] The intent slot structure is mapped and matched with the corresponding nodes, associated paths and anchor relationships in the multi-granularity anchor dual graph to determine the target node range and service path constraints that are adapted to the user's natural language request, and to generate a graph service mapping table.

[0035] Based on the graph service mapping table, the calling method, node retrieval order, path generation method, and service organization method of the multi-granularity anchor point dual graph are arranged to generate a graph service execution plan.

[0036] As a preferred embodiment of the knowledge graph-based knowledge service construction method in the publishing field described in this invention, the steps for obtaining the knowledge service candidate set are as follows:

[0037] The map service execution plan is converted into retrieval control parameters, and the initial retrieval entry point is determined in the multi-granularity anchor point dual map based on the retrieval control parameters;

[0038] Starting from the initial retrieval entry point, semantic nodes, structural nodes, and anchor relationships corresponding to the service target are extracted from the multi-granularity anchor point dual graph to generate a dual graph candidate node table.

[0039] Based on the candidate node table of the two graphs, path splicing and constraint filtering are performed along the anchor edge between the publishing semantic graph and the publishing content structure graph to determine the knowledge service path that meets the service requirements and generate a knowledge service candidate set.

[0040] As a preferred embodiment of the knowledge graph-based method for constructing knowledge services in the publishing field according to the present invention, the steps for generating knowledge service content in the publishing field are as follows:

[0041] According to the service path sequence of the knowledge service execution plan, the candidate contents in the knowledge service candidate set are arranged in order to generate a knowledge service ranking table;

[0042] According to the knowledge service ranking table, the candidate content is merged into service paths and integrated into services to generate knowledge service content in the publishing field.

[0043] The beneficial effects of this invention are as follows: by establishing anchor edges between the publishing semantic graph and the publishing content structure graph, a multi-granularity anchor point dual graph is formed, realizing the collaborative mapping of semantic relationships and content carrying positions, and enhancing the accuracy of knowledge organization and service orientation; by mapping the intent slot structure to the multi-granularity anchor point dual graph, a graph service execution plan is generated, realizing the orderly transformation of service requirements into retrieval paths and organization methods, and improving the coherence and adaptability of knowledge service content generation in the publishing field. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A flowchart illustrating a method for constructing knowledge services in the publishing field based on knowledge graphs.

[0046] Figure 2 A flowchart for obtaining the publishing semantic graph.

[0047] Figure 3 A flowchart for generating a bi-graph with multiple anchor points.

[0048] Figure 4 A flowchart for generating knowledge service content in the publishing field. Detailed Implementation

[0049] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0050] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0051] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0052] Reference Figures 1-4 This is one embodiment of the present invention, which provides a method for constructing knowledge services in the publishing field based on knowledge graphs, including the following steps:

[0053] S1. Perform multi-granularity content structure analysis on the original publishing resource set, identify the hierarchical organization of content and content carriers in the publishing resources, and generate a set of publishing content structure nodes.

[0054] S1.1: Extract directory hierarchy information, page layout information, and content boundary information that represent the organization of publishing resources from the original publishing resource collection, and generate a publishing content structure analysis framework.

[0055] Furthermore, the system identifies the numerical sequence of titles, font size, bolding attributes, and left indentation values ​​of paragraphs in the original published resource set to extract directory hierarchy information representing logical levels; it uses layout analysis algorithms to perform region detection and segmentation on page images or layout descriptions in the original published resource set to identify the coordinate range of regions such as header, footer, body text, and illustration areas, and extracts layout partition information; it scans document format tags, consecutive blank lines, and the start and end position marks of text blocks in the original published resource set to extract content boundary information defining the independent scope of paragraphs, lists, and tables.

[0056] It should be noted that the layout analysis algorithm reads the text block positions, font distribution, white space, wireframe boundaries, and graphic layout information from the page image or layout description, performs horizontal and vertical projection statistics on the page, identifies the segmented areas with obvious interval changes on the page, and combines connected component clustering to merge adjacent text blocks, image blocks, and annotation blocks, dividing the page into header area, footer area, body text area, illustration area, and annotation area. Finally, based on the coordinate range, size ratio, relative position, and adjacency relationship of each area, the algorithm completes the area boundary correction, realizing the area detection and segmentation of the page layout.

[0057] The hierarchical structure and content development order of the publishing resources are determined according to the directory hierarchy information. The content of each level is mapped to the specific carrying area on the page according to the page layout information. The start and end positions of each level of content in the corresponding carrying area are defined according to the content boundary information. The directory hierarchy information, page layout information and content boundary information are correlated and organized in a unified manner to generate a publishing content structure analysis framework.

[0058] S1.2: Based on the publishing content structure analysis framework, the publishing resources are hierarchically divided, the hierarchical relationship and sequential adjacency relationship between the publishing content at each level are identified, and a hierarchical relationship chain of publishing content is generated.

[0059] Furthermore, based on the directory hierarchy information, the numbering sequence and indentation level of the titles are analyzed to determine the tags and nesting depth of the published content at each level, thereby dividing the hierarchy. For example, titles with a main number and no indentation are identified as upper-level published content, while titles with sub-numbers or deeper indentation are identified as lower-level published content. Combining the page layout information, the header area, footer area, body text area, illustration area, annotation area, and index area in the published resources are mapped to the corresponding level according to coordinates. The content boundary information is used to define the start and end positions of each published content based on format marks, blank line separators, paragraph start and end positions, and figure caption boundaries.

[0060] Analyze the nested structure of headings in the directory hierarchy information, for example, by comparing the parent-child relationship of numbering (such as 1 and 1.1) or the difference in indentation levels to determine the hierarchical relationship of a higher-level heading containing a lower-level heading; analyze the coordinate order of published content within the same level in the page layout information, for example, by comparing horizontal and vertical positions (such as left to right, top to bottom) to determine the sequential adjacency relationship between published content; and connect the hierarchical relationship and the sequential adjacency relationship according to the hierarchical nesting and positional connection order to generate a hierarchical relationship chain of published content.

[0061] S1.3: Based on the hierarchical relationship chain of published content, objectify the published content at each level, determine the content carrier, location, subordinate path and sequence relationship of each level of published content, and generate a set of published content object identifiers.

[0062] Furthermore, unique identifiers are assigned to each level of published content in the hierarchical relationship chain to achieve object-oriented identification. Based on the start and end positions defined in the content boundary information, corresponding content fragments are extracted from the published resources to determine the content carrier objects. Among them, content located in the main text area and within the boundaries of consecutive paragraphs corresponds to the main text carrier object; content located in the illustration area and corresponding to the boundary of the figure or chart caption corresponds to the figure or chart carrier object; content located in the annotation area and limited by the annotation boundary corresponds to the annotation carrier object; and content located in the index area or reference area and limited by the entry boundary corresponds to the index carrier object or the reference carrier object, respectively.

[0063] The coordinate range of the corresponding area of ​​each level of published content is obtained from the page layout information to determine its location; based on the hierarchical relationship in the published content hierarchy, an identifier sequence is constructed from the top-level published content to the current published content to form a subordinate path; based on the sequential adjacency relationship in the published content hierarchy, the order relationship is determined by comparing the coordinate position and the arrangement order within the same level; the objectified identifier, the content carrier object, the location, the subordinate path, and the order relationship are recorded one by one to generate a set of published content object identifiers.

[0064] S1.4: Based on the set of object identifiers for published content and the hierarchical relationship chain of published content, assign structural node identifiers to each level of published content and its carrier object, establish node attribute descriptions, and generate a set of structural nodes for published content.

[0065] Furthermore, a unique structural node identifier is assigned to each level of published content and content-carrying object corresponding to each objectified identifier. The content-carrying object type, its location coordinates, subordinate path, and sequence relationship are extracted from the set of published content object identifiers. The hierarchical relationship and adjacency relationship are extracted from the published content hierarchical relationship chain and integrated into a node attribute description, including node type, location coordinates, level depth, superior node identifier, subordinate node identifier, and same-level sequence information. Each structural node identifier and its corresponding node attribute description are combined into a structural node record, and a set of published content structural nodes is generated.

[0066] S2. Based on the set of nodes in the published content structure, construct a published content structure graph, and combine it with the original published resource set to perform semantic information extraction and semantic association analysis in the publishing field to obtain a publishing semantic graph.

[0067] S2.1: Based on the location, subordinate path and sequence relationship of each structural node in the publishing content structure node set, establish the structural mapping relationship between each structural node and generate the publishing content structure networking rules.

[0068] Furthermore, the node identifier sequence in the subordinate path is extracted to determine the hierarchical nesting relationship between each structural node. For example, the upper-level node identifier and the lower-level node identifier are identified by the path separator identifier. According to the same-level sequence number in the sequence relationship, the order of arrangement of structural nodes in the same level is determined. Combined with the horizontal and vertical coordinate values ​​in the position coordinates, the position interval and arrangement direction between structural nodes are calculated to determine the position mapping relationship between structural nodes.

[0069] The expression for calculating the positional spacing and arrangement direction between structural nodes is:

[0070] ;

[0071] ;

[0072] in, Represents structural nodes With structural nodes The positional interval between them; Represents structural nodes The horizontal coordinate value in the coordinates of its location; Represents structural nodes The horizontal coordinate value in the coordinates of its location; Represents structural nodes The vertical coordinate value in the coordinate system of its location; Represents structural nodes The vertical coordinate value in the coordinate system of its location; Represents structural nodes Pointing to structure node The direction of arrangement.

[0073] Based on hierarchical nesting, same-level order, and position mapping relationships, we define the upper and lower level connection rules (such as edges pointing from upper-level nodes to lower-level nodes), same-level order rules (such as connecting same-level nodes by sequentially increasing sequence numbers), and position mapping rules (such as determining adjacency relationships based on coordinates) between structural nodes, thereby establishing the structural mapping relationships between each structural node; and integrate the upper and lower level connection rules, same-level order rules, and position mapping rules to generate the publishing content structure network rules.

[0074] S2.2: Connect the structural nodes in the publishing content structure node set according to the publishing content structure networking rules to generate a publishing content structure graph.

[0075] Furthermore, each structural node in the set of publishing content structural nodes is treated as a graph node. According to the hierarchical connection rules, hierarchical connection edges are established between structural nodes with hierarchical nesting relationships. According to the same-level order rules, sequential connection edges are established between structural nodes at the same level and with a sequential arrangement. According to the position mapping rules, positional connection edges are established between structural nodes with adjacency or positional correspondence. Each graph node is associated with its corresponding node attribute description, and each connection edge is associated with its corresponding connection type. All structural nodes and connection edges are then uniformly organized according to the node connection relationships to generate a publishing content structure graph.

[0076] S2.3: Based on the original resource set of publications and combined with the structural positioning information in the content structure diagram of publications, perform semantic information extraction of the publishing domain on the content of each level of publications to obtain the semantic elements of publications.

[0077] It should be noted that the structural positioning information in the publication content structure diagram includes the location, level depth, subordinate path, sequence relationship of each structural node, as well as the hierarchical connection relationship, sequential connection relationship and positional connection relationship between it and other structural nodes. It is used to indicate which level a certain publication content is located at, which path it is attached to, what sequence it is in, and how it is structurally associated with surrounding publication content.

[0078] Furthermore, based on the location coordinates, subordinate paths, and sequential relationships of the structural nodes in the publication content structure diagram, the original content fragments, such as text paragraphs, data tables, and chart captions corresponding to each level of publication content, are located and extracted from the original publication resource set. For each located original content fragment, semantic information extraction in the publishing field is performed using natural language processing technology. Specifically, this includes: extracting entities such as names of people, places, institutions, and times; extracting subject-specific terms; refining core concepts (such as theories, methods, and conclusions); and identifying attributes (such as authors and publication years) and relationships (such as definitions, descriptions, comparisons, and causality) between concepts. All entities, subject-specific terms, core concepts, and attributes are then compiled as semantic elements for publication.

[0079] It should be noted that natural language processing (NLP) technology is used to transform character sequences in published resources into recognizable, extractable, and organizeable semantic content. NLP technology typically includes processing methods such as word segmentation, part-of-speech tagging, syntactic analysis, named entity recognition, terminology recognition, concept extraction, semantic classification, and relation recognition. It can identify semantic content such as names of people, places, institutions, times, subject-specific terms, theories, methods, and conclusions from text paragraphs, figure captions, annotations, and reference entries, and further extract attribute information and semantic association clues.

[0080] S2.4: Using the structural positioning information in the publication content structure diagram, perform semantic association analysis on the semantic associations between various publishing semantic elements, determine the semantic association direction, association type and association strength, and generate a publishing semantic association chain.

[0081] Furthermore, based on the coordinates of the corresponding structural nodes in the publishing content structure diagram, the spatial distance between any two publishing semantic elements is calculated and compared with a proximity threshold. When the spatial distance is less than or equal to the proximity threshold, the publishing semantic elements are determined to have layout proximity. Based on the subordinate paths of the structural nodes corresponding to the publishing semantic elements, the path strings are compared. If the subordinate path of one publishing semantic element is a prefix of the subordinate path of another publishing semantic element, a hierarchical inclusion relationship is determined. Based on the same-level sequence number in the sequential relationship of the structural nodes corresponding to the publishing semantic elements, if the sequence numbers of two publishing semantic elements are consecutive, a sequential adjacency relationship is determined. When the publishing semantic elements satisfy at least one of layout proximity, hierarchical inclusion relationship, or sequential adjacency relationship, the semantic association between the publishing semantic elements is confirmed by combining the co-occurrence frequency (e.g., co-occurrence frequency exceeds the co-occurrence frequency threshold), semantic similarity (e.g., semantic similarity exceeds the semantic similarity threshold), and syntactic dependency relationship (e.g., subject-verb-object relationship identified through dependency parsing) in the corresponding content fragments of the original publishing resource set.

[0082] The semantic association direction is determined based on hierarchical inclusion relationships (from the containing publishing semantic element to the included publishing semantic element) or sequential adjacency relationships (from the prior publishing semantic element to the subsequent publishing semantic element). The semantic association type (such as definition, description, comparison, or causation) is identified by analyzing patterns in the co-occurrence context (e.g., the presence of "defined as" indicates a definition relationship) and logical dependencies (e.g., the presence of causal conjunctions indicates a causal relationship). The semantic association strength is obtained by weighting the co-occurrence frequency, the reciprocal of the spatial distance, and the semantic similarity. Before calculating the semantic association strength, the co-occurrence frequency, the reciprocal of the spatial distance, and the semantic similarity are normalized respectively, and weight coefficients are assigned to the three based on the association identification accuracy on the validation sample set. The sum of the weight coefficients is always 1. The publishing semantic element pairs with semantic associations and their corresponding semantic association directions, semantic association types, and semantic association strengths are concatenated to generate a publishing semantic association chain.

[0083] It should be noted that the proximity threshold, co-occurrence frequency threshold, and semantic similarity threshold are calculated by iterating through candidate threshold combinations on the validation sample set, and the threshold combination corresponding to the highest association recognition accuracy is used as the set value. For example, the value range of the proximity threshold is 0.05-0.20 times the page height, the value range of the co-occurrence frequency threshold is 2-5 times, and the value range of the semantic similarity threshold is 0.60-0.85.

[0084] S2.5: Graph the semantic elements of publishing into semantic nodes, and graph each semantic association in the semantic association chain of publishing into semantic edges to generate a semantic graph of publishing.

[0085] Furthermore, a unique semantic node identifier is assigned to each publishing semantic element, and the semantic content and type information of the publishing semantic element are recorded as semantic node attributes. Each association record in the publishing semantic association chain is traversed, and a semantic edge is created by taking the semantic node corresponding to the starting publishing semantic element in the association record as the starting point of the directed edge and the semantic node corresponding to the ending publishing semantic element as the ending point of the directed edge. The semantic association type and semantic association strength in the association record are assigned to the semantic edge as attributes. A publishing semantic graph is constructed with semantic nodes as vertices and semantic edges as edges.

[0086] S3. Extract the structural distribution features and content anchoring features of semantic nodes in publishing resources from the publishing semantic graph, and establish anchor edges between the publishing semantic graph and the publishing content structure graph to form a multi-granularity anchor point dual graph.

[0087] S3.1: Backtrack and match the semantic nodes in the publishing semantic graph with the corresponding content fragments in the original publishing resource set, and count the occurrence position, distribution frequency and hierarchical landing point of each semantic node in the publishing content at each level to generate a semantic node structure distribution spectrum.

[0088] Furthermore, based on the semantic content attributes of semantic nodes, all text or data fragments containing semantic content are located and obtained in the original publishing resource set by combining string matching and semantic similarity matching. For each matched content fragment, the specific location of the semantic node in each level of publishing content is determined by associating it with the subordinate path of the corresponding structural node in the publishing content structure diagram, and the location is recorded as a combination of subordinate path and sequence number.

[0089] The number of content fragments matched by the same semantic node in the published content at each level is counted to determine the distribution frequency of the semantic node in the corresponding level, and the level landing point of the semantic node is determined by parsing the level identifier in the subordinate path; the occurrence position, distribution frequency and level landing point of each semantic node are integrated to generate the semantic node structure distribution spectrum.

[0090] S3.2: Perform alignment analysis between semantic nodes and corresponding structural nodes in the publication content structure diagram to determine the content attachment position and structural pointing position of each semantic node in the publication content structure diagram, and generate a semantic node anchoring feature table.

[0091] Furthermore, based on the position of semantic nodes in each level of published content recorded in the semantic node structure distribution spectrum (the combination of subordinate path and sequence number), structural nodes with the same subordinate path and sequence relationship are searched in the published content structure diagram; the specific coordinates of the original content fragments corresponding to the semantic nodes in the area under the jurisdiction of the structural nodes are compared to determine the content attachment position of the semantic nodes in the published content structure diagram (such as the starting relative position, ending relative position, or normalized position interval within the area under the jurisdiction of the structural nodes); the node identifier and node path of the matched structural nodes in the published content structure diagram are recorded as the structural pointing position of the semantic nodes; the content attachment position and structural pointing position of each semantic node are integrated to generate a semantic node anchoring feature table.

[0092] S3.3: Match the semantic node structure distribution spectrum and the semantic node anchoring feature table to determine the target structure node and anchoring connection method corresponding to each semantic node, and generate a dual-graph anchor point mapping chain.

[0093] Furthermore, based on the semantic node identifier, the semantic nodes recorded in the semantic node structure distribution spectrum are associated with the structure pointing positions of the same semantic node recorded in the semantic node anchoring feature table. Based on the structure pointing positions, the corresponding structure nodes are located in the published content structure diagram and determined as the target structure nodes corresponding to the semantic nodes.

[0094] The semantic node anchoring feature table is used to read the positional relationship between the content attachment position of the semantic node and the area under the jurisdiction of the target structural node. If the content attachment position is completely inside the area corresponding to the target structural node, the anchoring connection method is defined as internal inclusion. If the content attachment position and the boundary of the area corresponding to the target structural node intersect, the anchoring connection method is defined as boundary contact. The semantic nodes, target structural nodes and anchoring connection methods are recorded in sequence according to the corresponding relationship to generate a dual-graph anchor point mapping chain.

[0095] It should be noted that in this embodiment, each semantic node corresponds to a single target structure node. When the same semantic node corresponds to multiple structure nodes at the same time, the structure node with the largest content attachment location coverage ratio is selected as the target structure node.

[0096] S3.4: Based on the dual-graph anchor mapping chain, anchor edges are established between semantic nodes and target structural nodes, and the anchor edges, publishing semantic graph and publishing content structural graph are associated and fused to generate a multi-granularity anchor dual graph.

[0097] Furthermore, anchor edges are established between each semantic node and its corresponding target structural node, and the anchoring connection method is written into the anchor edge attribute to represent the connection form between the semantic node and the target structural node. Each anchor edge is merged with the semantic nodes and semantic edges in the publishing semantic graph, as well as the structural nodes and connecting edges in the publishing content structural graph, so that the anchor edges connect the semantic nodes in the publishing semantic graph and the target structural nodes in the publishing content structural graph. This unifies the anchor edges, the publishing semantic graph, and the publishing content structural graph, generating a multi-granularity anchor point bi-graph.

[0098] It should be noted that the multi-granularity anchor point dual graph achieves collaborative mapping and bidirectional tracing of semantic relationships and content carrying positions by linking and integrating the publishing semantic graph and the publishing content structure graph through anchor point edges, enabling knowledge services to be located and organized synchronously along the semantic and structural dimensions.

[0099] S4. Perform service intent recognition and slot extraction on the user's natural language request, obtain the intent slot structure, and map it to a multi-granularity anchor point dual graph to generate a graph service execution plan.

[0100] S4.1: Perform semantic parsing and element decomposition on user natural language requests, and extract request semantic elements that represent service requirements.

[0101] Furthermore, the text of the user's natural language request is segmented to obtain a word sequence, and part-of-speech tagging and named entity recognition are performed on the word sequence. The grammatical category of each word is labeled and the personal names, place names, organization names and domain terms are identified. Dependency parsing is performed on the labeled word sequence to construct a dependency parsing tree that reflects the subject-verb, verb-object and attributive-head relations between words, and the semantic roles such as agent, patient, time and place are labeled.

[0102] Based on the dominance relationship between predicates and noun components in dependency syntax, the noun phrases representing the query target are identified and determined as the target entity. Based on the noun phrases, quantitative phrases, locative phrases, and prepositional phrases modifying the target entity in dependency syntax, the descriptive content representing the filtering requirements and scope is identified and determined as attribute conditions. Based on the verbs, prepositions, and relational phrases connecting the target entity with other semantic components in dependency syntax, the descriptive content representing the query targeting and association requirements is identified and determined as relational predicates. Based on the enumeration and generalization expressions appearing in the user's natural language request, the descriptive content representing the result presentation method is identified and determined as the output format. The target entity, attribute conditions, relational predicates, and output format are uniformly defined as request semantic elements representing service requirements.

[0103] S4.2: Perform semantic classification and slot mapping on the semantic elements of the request, determine the service goal, service object, constraints and output requirements in the user's natural language request, and generate the intent slot structure.

[0104] Furthermore, the target entities, attribute conditions, relational predicates, and output formats in the request semantic elements are semantically categorized. Request semantic elements representing the query focus or service direction are categorized into service targets; target entities representing the retrieved, located, or associated content are categorized into service objects; attribute conditions representing scope limitations, attribute filtering, hierarchical limitations, and path restrictions are categorized into constraint conditions; and output formats representing the presentation, organization, and style of the results are categorized into output requirements. Slot mapping is then performed on the categorized request semantic elements, writing different categories of request semantic elements into corresponding slots and associating them according to the semantic dependency order and combination relationship in the user's natural language request to generate an intent slot structure.

[0105] S4.3: Map and match the intent slot structure with the corresponding nodes, associated paths and anchor relationships in the multi-granularity anchor dual graph to determine the target node range and service path constraints that are compatible with the user's natural language request, and generate a graph service mapping table.

[0106] Furthermore, based on the service objects in the intent slot structure, nodes that match the semantic content of the service objects are searched in the semantic nodes and structural nodes of the multi-granularity anchor point dual graph, and the target node range is initially determined.

[0107] Based on constraints (such as level, time, and author), the initially determined range of target nodes is filtered. Specifically, the hierarchical restrictions are checked by examining the subordinate paths and locations of structural nodes to see if they meet the constraints. The time and author restrictions are checked by combining semantic node attributes with structural node attributes to see if they meet the constraints. At the same time, the associated nodes that are consistent with the constraints are retained by combining the anchoring connection relationships corresponding to the anchor edges, thereby determining the final range of target nodes.

[0108] Based on the service objective, the associated path type and anchor point relationship type corresponding to the service objective are determined in the multi-granularity anchor point dual graph and defined as service path constraints; the target node range, service path constraints and output requirements are organized accordingly to generate a graph service mapping table.

[0109] It should be noted that when the service objective is to query the definition, the service path constraint is limited to extending along the semantic edges of the definition class or description class, and locating the corresponding structural node through the anchor edge; when the service objective is to trace the source, the service path constraint is limited to tracing back along the semantic edges related to author relationship, citation relationship or publication information.

[0110] S4.4: Based on the graph service mapping table, arrange the calling method, node retrieval order, path generation method, and service organization method of the multi-granularity anchor point dual graph, and generate the graph service execution plan.

[0111] Furthermore, if the target node range is primarily composed of semantic nodes, the publishing semantic graph and its corresponding anchor edges are prioritized. If the target node range is primarily composed of structural nodes, the publishing content structure graph and its corresponding anchor edges are prioritized. If both semantic and structural nodes within the target node range satisfy the retrieval conditions, the publishing semantic graph and the publishing content structure graph are invoked in parallel, and the expansion results of the two types of nodes are merged according to service path constraints. Based on the target node range and service path constraints in the graph service mapping table, the node retrieval order is determined for nodes within the target node range according to their hierarchical depth, sequential relationship, or path proximity. Based on service path constraints, the path generation method for path expansion and traversal along specific types of semantic edges, structural edges, and anchor edges in the multi-granularity anchor point dual graph is specified. The service organization method is arranged according to the output requirements, the organization order and presentation form of the retrieved content in the output are determined, and the invocation method, node retrieval order, path generation method, and service organization method are integrated accordingly to generate a graph service execution plan.

[0112] S5. Based on the graph service execution plan, perform dual-graph collaborative retrieval and service path generation, obtain a knowledge service candidate set, and perform service-oriented sorting and arrangement to generate knowledge service content in the publishing field.

[0113] S5.1: Convert the map service execution plan into retrieval control parameters, and determine the initial retrieval entry point in the multi-granularity anchor point dual map based on the retrieval control parameters.

[0114] Furthermore, the invocation method, node retrieval order, path generation method, and service organization method recorded in the graph service execution plan are broken down into corresponding retrieval control items. Specifically, the invocation method used to determine the starting graph type for retrieval is converted into graph invocation parameters, the node retrieval order used to determine the order of candidate nodes is converted into node priority parameters, the path generation method used to limit the path expansion range is converted into path constraint parameters, and the service organization method used to limit the output organization form is converted into result organization parameters. The initial invocation graph type in the multi-granularity anchor point dual graph is determined based on the graph invocation parameters, candidate entry nodes are screened based on the node priority parameters, and the edge type reachability, hierarchical reachability, and anchor point connectivity of the candidate entry nodes are checked in conjunction with the path constraint parameters to see if they meet the subsequent path expansion conditions, thereby determining the initial retrieval entry point.

[0115] S5.2: Starting from the initial retrieval entry point, extract semantic nodes, structural nodes, and anchor relationships corresponding to the service target from the multi-granularity anchor point dual graph, and generate a dual graph candidate node table.

[0116] Furthermore, starting from the initial retrieval entry point, based on the allowed edge types specified by the path constraint parameters in the graph service execution plan, path expansion is performed in the multi-granularity anchor point dual graph: traversing along the semantic edges in the publishing semantic graph to obtain associated semantic nodes, traversing along the connection edges in the publishing content structure graph to obtain associated structure nodes, and traversing along the anchor point edges connecting the publishing semantic graph and the publishing content structure graph to obtain another graph node that has an anchor point relationship with the current node;

[0117] During the traversal, the node attributes of the semantic nodes and structural nodes obtained are checked in real time, and matching is performed according to the service target in the graph service mapping table. For example, when the service target is query definition, nodes whose semantic content contains definitional expressions are matched; when the service target is location, structural nodes that correspond to the target entity and whose location meets the constraints are matched. The node identifier, node type, graph to which it belongs and relationship type of the successfully matched nodes are recorded respectively, and a dual-graph candidate node table is generated.

[0118] S5.3: Based on the candidate node table of the two graphs, perform path splicing and constraint filtering along the anchor edge between the publishing semantic graph and the publishing content structure graph to determine the knowledge service path that meets the service requirements and generate a knowledge service candidate set.

[0119] Furthermore, taking the semantic nodes and structural nodes with anchor point relationships as the starting point of connection, sequentially splicing them along the semantic edges in the publishing semantic graph, the structural edges in the publishing content structural graph, and the anchor point edges between the two graphs, to form candidate paths from semantic nodes to structural nodes or from structural nodes to semantic nodes.

[0120] Based on the path constraint parameters in the graph service execution plan and the service path constraints in the graph service mapping table, the candidate paths are screened item by item according to the node type, edge type, hierarchical position, subordinate path and anchor connection method, and candidate paths that meet the service objectives, constraints and output requirements are retained. The semantic nodes, structural nodes, anchor relationships and path order covered by each candidate path after screening are sorted out to generate a knowledge service candidate set.

[0121] S5.4: Arrange the candidate content in the knowledge service candidate set in order according to the service path sequence of the graph service execution plan, and generate a knowledge service sorting table.

[0122] Furthermore, according to the service path order of the graph service execution plan, the path order, node arrangement order, and structural hierarchy order corresponding to each candidate content in the knowledge service candidate set are read; the candidate content is sorted in the first round according to the path order, and the candidate content corresponding to the first knowledge service path is arranged first, and the candidate content corresponding to the subsequent knowledge service paths is arranged in turn.

[0123] When multiple candidate contents correspond to the same path priority, the node arrangement order, the hierarchical position of the structural nodes in the published content structure diagram, and the connection order of the anchor edges in the knowledge service path are read. The nodes are compared level by level according to the node arrangement order, hierarchical position, and connection order. When the node arrangement order, hierarchical position, and connection order are all the same, the structural nodes are sorted according to their page position on the page. If the page position order is still the same, the content is sorted according to the type of the content it carries. The order of each candidate content under the same path priority is determined. The sorted candidate contents, corresponding path priority, node arrangement order, and hierarchical position are recorded to generate a knowledge service sorting table.

[0124] S5.5: Perform service path merging and service integration on candidate content according to the knowledge service sorting table to generate knowledge service content in the publishing field.

[0125] Furthermore, candidate content belonging to the same knowledge service path and with consistent content direction is merged into service paths to form a content chain that unfolds continuously in the order of service paths. During the merging process, redundant candidate content and duplicate path segments are eliminated to ensure the continuity of the content chain.

[0126] If the output requirement in the graph service mapping table is a list, then the identifier, content summary, and position information of the corresponding node of the candidate content in the published content structure graph are extracted sequentially from each content chain and compiled in the form of entries; if the output requirement is a summary, then the semantic content covered by each content chain is summarized and associated with the corresponding structural node chapter information to form a continuous expression of content.

[0127] The compiled content chains are organized and output according to the overall order determined by the knowledge service sorting table to generate knowledge service content in the publishing field.

[0128] In summary, this invention achieves coordinated mapping of semantic relationships and content carrying positions by establishing anchor edges between the publishing semantic graph and the publishing content structure graph, forming a multi-granularity anchor point dual graph, thereby enhancing the accuracy of knowledge organization and service orientation; and by mapping the intent slot structure to the multi-granularity anchor point dual graph to generate a graph service execution plan, it realizes the orderly transformation of service requirements into retrieval paths and organization methods, improving the coherence and adaptability of knowledge service content generation in the publishing field.

[0129] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for constructing knowledge services in the publishing field based on knowledge graphs, characterized in that, include: Multi-granularity content structure analysis is performed on the original set of publishing resources to identify the hierarchical organization of content and content carriers in the publishing resources, and to generate a set of publishing content structure nodes; Based on the set of structural nodes of the published content, a structural graph of the published content is constructed, and semantic information extraction and semantic association analysis in the publishing domain are performed in combination with the original set of publishing resources to obtain a semantic graph of the publishing. Extract the structural distribution features and content anchoring features of semantic nodes in publishing resources from the publishing semantic graph, and establish anchor edges between the publishing semantic graph and the publishing content structure graph to form a multi-granularity anchor point bi-graph; The system performs service intent recognition and slot extraction on user natural language requests, obtains the intent slot structure, maps it to a multi-granularity anchor point dual graph, and generates a graph service execution plan. Based on the graph service execution plan, the system performs dual-graph collaborative retrieval and service path generation to obtain a knowledge service candidate set, and then performs service-oriented sorting and arrangement to generate knowledge service content in the publishing field.

2. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 1, characterized in that, The steps for generating the published content structure node set are as follows: Extract directory hierarchy information, page layout information, and content boundary information that characterize the organization of publishing resources from the original publishing resources to generate a publishing content structure analysis framework; Based on the content structure analysis framework, the publishing resources are hierarchically divided, the hierarchical and sequential adjacency relationships between the publishing content at each level are identified, and a hierarchical relationship chain of publishing content is generated. Based on the hierarchical relationship chain of publishing content, objectify the publishing content at each level, determine the content carrier, location, subordinate path and sequence relationship of each level of publishing content, and generate a set of publishing content object identifiers; Based on the object identifier set of published content and the hierarchical relationship chain of published content, structural node identifiers are assigned to each level of published content and its carrier object, and node attribute descriptions are established to generate a set of structural nodes for published content.

3. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 1, characterized in that, The steps for constructing the publication content structure diagram are as follows: Based on the location, subordinate path, and sequence relationship of each structural node in the publishing content structure node set, establish the structural mapping relationship between each structural node and generate the publishing content structure networking rules; Based on the network rules for publishing content structure, the structural nodes in the publishing content structure node set are connected in a graph to generate a publishing content structure graph.

4. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 3, characterized in that, The steps for obtaining the publishing semantic graph are as follows: Based on the original set of publishing resources and combined with the structural positioning information in the publishing content structure diagram, semantic information of the publishing domain is extracted from the publishing content at each level to obtain publishing semantic elements; Using the structural positioning information in the publication content structure diagram, semantic association analysis is performed on the semantic associations between various publishing semantic elements to determine the semantic association direction, association type and association strength, and to generate a publishing semantic association chain. The semantic elements of publishing are graphed into semantic nodes, and the semantic associations of publishing in the semantic association chain are graphed into semantic edges, thus generating a semantic graph of publishing.

5. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 1, characterized in that, The steps for extracting the structural distribution features and content anchoring features of semantic nodes in publishing resources are as follows: The semantic nodes in the publishing semantic graph are backtracked and matched with the corresponding content fragments in the original publishing resource set. The occurrence position, distribution frequency and hierarchical landing point of each semantic node in the publishing content at each level are counted to generate a semantic node structure distribution spectrum. By performing alignment analysis between semantic nodes and corresponding structural nodes in the publication content structure diagram, the content attachment position and structural pointing position of each semantic node in the publication content structure diagram are determined, and a semantic node anchoring feature table is generated.

6. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 5, characterized in that, The steps for forming a multi-granularity anchor point bi-graph are as follows: The semantic node structure distribution spectrum and the semantic node anchoring feature table are associated and matched to determine the target structure node and anchoring connection method corresponding to each semantic node, and generate a dual-graph anchor point mapping chain. Based on the dual-graph anchor mapping chain, anchor edges are established between semantic nodes and target structural nodes, and the anchor edges, publishing semantic graph and publishing content structural graph are associated and fused to generate a multi-granularity anchor dual graph.

7. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 1, characterized in that, The steps for obtaining the intent slot structure are as follows: Semantic parsing and element decomposition are performed on user natural language requests, and request semantic elements that represent service needs are extracted; Semantic classification and slot mapping are performed on the semantic elements of the request to determine the service goals, service objects, constraints and output requirements in the user's natural language request, and to generate the intent slot structure.

8. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 7, characterized in that, The execution plan for the generated atlas service includes the following steps: The intent slot structure is mapped and matched with the corresponding nodes, associated paths and anchor relationships in the multi-granularity anchor dual graph to determine the target node range and service path constraints that are adapted to the user's natural language request, and to generate a graph service mapping table. Based on the graph service mapping table, the calling method, node retrieval order, path generation method, and service organization method of the multi-granularity anchor point dual graph are arranged to generate a graph service execution plan.

9. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 1, characterized in that, The steps for obtaining the knowledge service candidate set are as follows: The map service execution plan is converted into retrieval control parameters, and the initial retrieval entry point is determined in the multi-granularity anchor point dual map based on the retrieval control parameters; Starting from the initial retrieval entry point, semantic nodes, structural nodes, and anchor relationships corresponding to the service target are extracted from the multi-granularity anchor point dual graph to generate a dual graph candidate node table. Based on the candidate node table of the two graphs, path splicing and constraint filtering are performed along the anchor edge between the publishing semantic graph and the publishing content structure graph to determine the knowledge service path that meets the service requirements and generate a knowledge service candidate set.

10. The method for constructing knowledge services in the publishing field based on knowledge graphs as described in claim 1 or 9, characterized in that, The steps for generating knowledge service content in the publishing field are as follows: According to the service path sequence of the knowledge service execution plan, the candidate contents in the knowledge service candidate set are arranged in order to generate a knowledge service ranking table; According to the knowledge service ranking table, the candidate content is merged into service paths and integrated into services to generate knowledge service content in the publishing field.

Citation Information

Patent Citations

  • Cabin active recommendation system and method based on knowledge graph and semantic reasoning

    CN120578813A

  • Digital publishing process intelligent management system based on knowledge graph

    CN121504116A

  • Method for automatically checking consistency of soft materials based on mapping knowledge domain

    CN121659288A

  • Method and device for improving RAG recall effect

    CN121705417A