Periodical knowledge graph construction method, periodical knowledge question-answering method and periodical knowledge graph construction device

Through sequence labeling method and priority strategy alignment disambiguation and knowledge inference completion, the problems of data integration and verification in journal knowledge graph are solved, and high-quality journal knowledge graph construction is achieved.

CN120373432APending Publication Date: 2025-07-25YANGTZE UNIVERSITY

Patent Information

Application Number
CN202510283824.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology fails to effectively integrate and verify semi-structured and unstructured data when constructing journal knowledge graphs, resulting in difficulty in ensuring data quality and being unable to build high-quality journal knowledge graphs.

Method used

The sequence annotation method is used to extract knowledge, align knowledge entities of different data sources through vocabulary network policies, and use priority strategies to disambiguate and knowledge reasoning to integrate and verify journal knowledge.

Benefits of technology

The data quality of journal knowledge graph is improved, and a high-quality journal knowledge graph is constructed by effectively integrating and verifying journal knowledge from different sources, eliminating ambiguities between entities and completing missing values.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373432A_ABST
    Figure CN120373432A_ABST
Patent Text Reader

Abstract

The invention provides a periodical knowledge graph construction method, a periodical knowledge question-answering method and a periodical knowledge graph construction device. The periodical knowledge graph construction method comprises the following steps: performing sequence labeling method knowledge extraction on an obtained periodical text to be processed to obtain a knowledge labeling sequence; performing vocabulary network strategy entity alignment on the knowledge labeling sequence to obtain an aligned knowledge sequence, performing priority strategy entity disambiguation on the aligned knowledge sequence to obtain a disambiguation knowledge sequence, and performing knowledge reasoning completion on the disambiguation knowledge sequence to obtain integrated journal knowledge; and performing knowledge graph construction according to the integrated periodical knowledge to obtain a periodical knowledge graph. Primary knowledge extraction is carried out through a sequence labeling method, knowledge entities of different data sources are aligned through a vocabulary network strategy, periodical knowledge of different sources can be effectively integrated, ambiguous content between entities is eliminated through priority strategy entity disambiguation, missing values and blank values are complemented through knowledge reasoning, the knowledge content can be effectively verified, and the accuracy of the periodical knowledge is improved. Therefore, the data quality of the periodical knowledge graph is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method for constructing a periodical knowledge graph, a method and device for periodical knowledge question answering. Background Technique

[0002] A knowledge graph is a way to describe concepts, entities and their relationships in the objective world in a structured form, which can express information in a form closer to the human cognitive world and provides an ability to better organize, manage and understand the vast amount of information on the Internet. For the periodical field, the storage method of the knowledge graph can effectively reflect the relationships between periodical knowledge and facilitate the query and use by users.

[0003] Currently, the construction of a periodical knowledge graph relies relatively heavily on structured periodical data, and the structured periodical data can be transformed into a knowledge graph in the same structured form through relatively simple processing steps. However, in the periodical field, there are more semi-structured and unstructured periodical data. These semi-structured and unstructured periodical data usually come from a large number of different data sources, and the storage formats between these data sources are often different. These multi-source data sets are prone to data chaos during the structuring process. At the same time, there may also be some content conflicts, content omissions or blank values in the periodical data from different sources. In the process of constructing a knowledge graph in the prior art, the structured data content cannot be effectively integrated and verified, it is difficult to ensure the data quality, and a high-quality periodical knowledge graph cannot be constructed.

[0004] Therefore, there are technical problems in the prior art that in the process of constructing a knowledge graph, the structured data content cannot be effectively integrated and verified, it is difficult to ensure the data quality, and a high-quality periodical knowledge graph cannot be constructed, which needs to be improved. Summary of the Invention

[0005] In view of this, it is necessary to provide a method for constructing a periodical knowledge graph, a method and device for periodical knowledge question answering, so as to effectively integrate and verify data during the process of constructing a periodical knowledge graph using semi-structured and unstructured data, ensure the data quality, and construct a high-quality periodical knowledge graph.

[0006] To solve the above technical problems, in a first aspect, the present invention provides a method for constructing a periodical knowledge graph, including: Performing knowledge extraction on the obtained periodical text to be processed by a sequence annotation method to obtain a knowledge annotation sequence; Performing entity alignment on the knowledge annotation sequence by a lexical network strategy to obtain an aligned knowledge sequence, performing entity disambiguation on the aligned knowledge sequence by a priority strategy to obtain a disambiguated knowledge sequence, and performing knowledge reasoning and complementation on the disambiguated knowledge sequence to obtain integrated periodical knowledge; Construct a journal knowledge graph according to the integrated journal knowledge; Among them, the entity disambiguation of the priority strategy includes entity disambiguation of the knowledge annotation sequence based on the priority of each entity, and the knowledge inference and completion includes filling in the missing values of the knowledge annotation sequence based on the preset inference rules.

[0007] In some possible implementation manners, the journal text to be processed is obtained by filtering the text of the crawler framework for semi-structured journal data and unstructured journal data.

[0008] In some possible implementation manners, knowledge extraction is performed on the obtained journal text to be processed by using the sequence annotation method to obtain a knowledge annotation sequence, including: Perform bidirectional feature encoding on the journal text to be processed to obtain word vectors; Extract context features from the word vectors by using a bidirectional long short-term memory network to obtain context features; Perform conditional random field output decoding on the context features to obtain a knowledge annotation sequence.

[0009] In some possible implementation manners, entity disambiguation of the priority strategy is performed on the aligned knowledge sequence to obtain a disambiguated knowledge sequence, including: Determine the priority of each knowledge entity in the aligned knowledge sequence based on the preset priority sequence and the data source category corresponding to the aligned knowledge sequence; Perform entity disambiguation on the aligned knowledge sequence according to the priority of the knowledge entity to obtain a disambiguated knowledge sequence.

[0010] In some possible implementation manners, knowledge inference and completion is performed on the disambiguated knowledge sequence to obtain integrated journal knowledge, including: Determine the complementary association entity and the corresponding association value of each missing entity in the disambiguated knowledge sequence based on the preset complementary association rules; Perform knowledge inference according to the complementary association entity and the association value to obtain the missing value of each missing entity, and complete each missing entity according to the missing value to obtain integrated journal knowledge.

[0011] In some possible implementation manners, a journal knowledge graph is constructed according to the integrated journal knowledge, including: Write the integrated journal knowledge into a preset graph database based on the data layer update mode to obtain a journal knowledge graph.

[0012] In a second aspect, the present invention provides a method for answering journal knowledge questions, including: Input the journal knowledge question to be answered into a large language model for natural language understanding to obtain a preliminary query semantics; Perform knowledge query on the journal knowledge graph according to the preliminary query semantics to obtain relevant journal knowledge; Obtain the journal knowledge Q&A output result by performing large language model Q&A output based on the preliminary query semantics and relevant journal knowledge; Among them, the journal knowledge graph is obtained according to the journal knowledge graph construction method of any one of the above.

[0013] In some possible implementation manners, the journal knowledge Q&A method further includes: Optimize the journal knowledge Q&A output result according to the obtained user portrait to obtain an optimized Q&A output result; Among them, the user portrait includes user type, user historical behavior, and user preference.

[0014] In a third aspect, the present invention provides a journal knowledge graph construction device, including: A knowledge extraction unit, configured to perform knowledge extraction on the obtained journal text to be processed to obtain a knowledge annotation sequence; A knowledge integration unit, configured to perform entity alignment on the knowledge annotation sequence by using a lexical network strategy to obtain an aligned knowledge sequence, perform entity disambiguation on the aligned knowledge sequence by using a priority strategy to obtain a disambiguated knowledge sequence, and perform knowledge inference and complementation on the disambiguated knowledge sequence to obtain integrated journal knowledge; A knowledge graph construction unit, configured to construct a knowledge graph according to the integrated journal knowledge to obtain a journal knowledge graph; Among them, entity disambiguation by using a priority strategy includes performing entity disambiguation on the knowledge annotation sequence based on the priority of each entity, and knowledge inference and complementation includes performing missing value complementation on the knowledge annotation sequence based on a preset inference rule.

[0015] In a fourth aspect, the present invention provides a journal intelligent Q&A device, including a processor and a memory. Among them, The memory is used to store programs; The processor is coupled to the memory and is configured to execute the programs stored in the memory to implement the steps in the journal knowledge graph construction method of any one of the above and / or the steps in the journal knowledge Q&A method of any one of the above.

[0016] In a fifth aspect, the present invention provides a computer-readable storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps in the journal knowledge graph construction method of any one of the above and / or the steps in the journal knowledge Q&A method of any one of the above can be implemented.

[0017] The beneficial effects of the present invention are as follows: In the method for constructing a journal knowledge graph provided by the present invention, preliminary knowledge extraction of knowledge entities in journal data is performed through knowledge extraction by sequence annotation method. The vocabulary network strategy aligns knowledge entities from different data sources, which can effectively integrate journal knowledge from different sources. The priority strategy entity disambiguation eliminates ambiguous content between entities, and knowledge reasoning complements missing values and blank values, which can effectively verify knowledge content, thereby improving the data quality of the journal knowledge graph. Brief Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of an embodiment of the method for constructing a journal knowledge graph provided by the present invention; Figure 2 It is a flowchart of knowledge extraction by sequence annotation method in the embodiment of the present invention; Figure 3 It is a flowchart of priority strategy entity disambiguation in the embodiment of the present invention; Figure 4 It is a flowchart of knowledge reasoning and complement in the embodiment of the present invention; Figure 5 It is a flowchart of an embodiment of a journal knowledge Q&A method provided by the present invention; Figure 6 It is a structural schematic diagram of an embodiment of the journal knowledge graph construction device provided by the present invention; Figure 7 It is a structural schematic diagram of an embodiment of the journal intelligent Q&A device provided by the present invention. Detailed Description of the Embodiments

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0021] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of the present invention, may add one or more other operations to the flowchart or remove one or more operations from the flowchart. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.

[0022] In the embodiments of the present invention, the descriptions such as "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one of such features.

[0023] The mention of "embodiment" in this document means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0024] The present invention provides a method for constructing a journal knowledge graph, a method for journal knowledge answering, and an apparatus, which will be described separately below.

[0025] Figure 1 is a schematic flowchart of an embodiment of the method for constructing a journal knowledge graph provided by the present invention. As Figure 1 shown, the method for constructing a journal knowledge graph includes: S101. Perform knowledge extraction on the obtained journal text to be processed by the sequence annotation method to obtain a knowledge annotation sequence; Among them, according to the different degrees of data structuring, these data can be divided into three categories: structured data, semi-structured data, and unstructured data. The journal text to be processed includes the filtered semi-structured journal data and unstructured journal data. These journal data contain a large amount of professional content, and there are complex association relationships between entities. Effectively extracting these entities and complex relationships is the key point of knowledge extraction. In this regard, the embodiment adopts a sequence annotation method for knowledge extraction to perform preliminary knowledge extraction on the journal text to be processed. The sequence annotation method for knowledge extraction includes converting the text data into vector content through a neural network model and extracting the context relationship therein, and finally obtaining the annotation sequence of the journal text through decoding to indicate the knowledge entities and entity relationships therein. In this process, the embodiment utilizes the efficient semantic extraction ability of the neural network to effectively extract and retain the content of the knowledge entities and the relationships between entities in the journal text to be processed.

[0026] S102. Perform entity alignment on the knowledge annotation sequence using a lexical network strategy to obtain an aligned knowledge sequence, perform entity disambiguation on the aligned knowledge sequence using a priority strategy to obtain a disambiguated knowledge sequence, and perform knowledge reasoning and completion on the disambiguated knowledge sequence to obtain integrated journal knowledge; Among them, entity disambiguation using a priority strategy includes performing entity disambiguation on the knowledge annotation sequence based on the priority of each entity, and knowledge reasoning and completion includes performing missing value completion on the knowledge annotation sequence based on a preset reasoning rule.

[0027] Among them, after performing preliminary knowledge extraction, to integrate the journal knowledge entities of each data source, the embodiment performs entity alignment by adopting a context-based BERT model. Using its bidirectional encoding mechanism, BERT effectively captures the surrounding context information of the entity, thereby reducing the ambiguity inherent in the dictionary-based matching method. For the conflicts existing between entities (such as value conflicts, structural conflicts, and semantic conflicts), the embodiment eliminates the conflicting content through entity disambiguation using a priority strategy. For each knowledge entity, the embodiment sets different priorities according to its reliability. When there is a conflict, the content with a high priority is retained, and the content with a low priority is corrected for the conflict.

[0028] For the missing values and blank values in the knowledge graph, the embodiment performs completion through a preset reasoning rule. The preset reasoning rule sets a completion association rule for the entity. When a missing entity appears, the embodiment determines the completion association entity and the corresponding association value corresponding to the missing value through the completion association rule, and then searches for the journal data in other journal data where the missing entity does not appear and the value of the completion association entity is the association value, and determines the missing value corresponding to the missing entity based on the found journal data, so as to complete the missing part.

[0029] Through entity alignment, entity disambiguation, and missing value completion steps, the embodiments can effectively integrate and verify knowledge content during the structuring process of semi-structured and unstructured journal data, ensure data quality, and provide effective data support for constructing a journal knowledge graph.

[0030] S103. Construct a journal knowledge graph based on the integrated journal knowledge.

[0031] Among them, during the construction process of the knowledge graph, preferably, the embodiments adopt a semi-automatic construction method. The machine extracts entities from the integrated journal knowledge and combines manual correction during the process to construct a high-quality ontology. Compared with the completely manual and fully automated methods, semi-automatic ontology construction can balance entity quality and construction efficiency.

[0032] It should be understood that the method for constructing a journal knowledge graph in this embodiment can also process structured journal data. Compared with the complex semi-structured and unstructured data processing steps, structured journal data can be processed with more or fewer steps according to the data quality, which is not limited in this embodiment.

[0033] Compared with the prior art, in the method for constructing a journal knowledge graph provided by the present invention, preliminary knowledge extraction of knowledge entities in journal data is performed through sequence annotation method knowledge extraction, and the knowledge entities from different data sources are aligned by the lexical network strategy, which can effectively integrate journal knowledge from different sources. Entity disambiguation is performed through the priority strategy to eliminate ambiguous content between entities, and knowledge reasoning is used to complete missing values and blank values, which can effectively verify knowledge content, thereby improving the data quality of the journal knowledge graph.

[0034] In some embodiments of the present invention, the journal text to be processed is obtained by filtering the text of the semi-structured journal data and unstructured journal data through a crawler framework.

[0035] Specifically, for data with different degrees of structuring, structured data generally comes from databases and third-party journal websites, has a clear format and a standardized structure, and is convenient for direct analysis and processing. Although semi-structured data has a certain structure, it does not fully conform to the standardized format. Unstructured data has no fixed format and usually exists in the form of natural language text. For semi-structured data and unstructured data, the embodiments use the Scrapy crawler framework to obtain and present them in the HTML and XML formats. Among them, for discrete data, the Scrapy crawler framework can first delete advertisements, navigation bars, footnotes, and duplicate content by parsing web pages, filter out irrelevant information. Then, XPath is used to locate and extract the required content, regular expressions are used to match specific patterns, and the text is filtered to obtain the journal text to be processed after filtering according to the rules.

[0036] In some embodiments of the present invention, Figure 2 is a schematic flow chart of knowledge extraction by the sequence annotation method according to the embodiments of the present invention. As Figure 2 shown, knowledge extraction by the sequence annotation method is performed on the obtained journal text to be processed to obtain a knowledge annotation sequence, including: S201. Perform bidirectional feature encoding on the journal text to be processed to obtain word vectors; Among them, there are complex association relationships between entities in the journal dataset. Effectively extracting these relationships from unstructured text is the basis for constructing a high-quality knowledge graph. In the initial extraction process of journal knowledge, the embodiment uses the BERT-BiLSTM-CRF model for relationship extraction. First, the embodiment uses the BERT (Bidirectional Encoder Representations from Transformers) pre-trained model to convert the input text into word vectors. During the text vectorization process, the BERT pre-trained model effectively considers the context relationship and can effectively capture the context semantic relevance therein.

[0037] S202. Perform context extraction on the word vectors through a bidirectional long short-term memory network to obtain context features; Among them, after the word vectors are extracted, the embodiment further extracts context features and performs modeling through BiLSTM (Bidirectional Long Short-Term Memory). BiLSTM is a recurrent neural network with forward and backward directions and can effectively capture the semantic relationship between contexts in the text.

[0038] S203. Perform conditional random field output decoding on the context features to obtain a knowledge annotation sequence.

[0039] Finally, the embodiment provides CRF (conditional random field) to decode the obtained context features. The CRF conditional random field is a discriminative probabilistic undirected graph model. By using CRF to annotate the knowledge sequence, the interaction between entities can be considered to ensure the extraction of the optimal annotation sequence in entity recognition.

[0040] In some embodiments of the present invention, Figure 3 is a schematic flow chart of entity disambiguation by the priority strategy according to the embodiments of the present invention. As Figure 3 shown, entity disambiguation by the priority strategy is performed on the aligned knowledge sequence to obtain a disambiguated knowledge sequence, including: S301. Determine the priorities of the knowledge entities in the aligned knowledge sequence based on the preset priority sequence and the data source category corresponding to the aligned knowledge sequence; S302. Disambiguate the aligned knowledge sequence according to the priorities of knowledge entities to obtain a disambiguated knowledge sequence.

[0041] Among them, for the possible data chaos in the data processing process and the possible conflicts in the journal data from different sources, the embodiments perform entity disambiguation by setting a priority strategy. Among them, the embodiments set a priority sequence for different data sources in advance to sort the priorities of the data. For example: in the priority sequence, the data from the official website of the journal with the most reliable source is taken as the first priority, the authoritative websites and databases that collect each journal (such as CNKI, Wanfang Database, and Web of Science) are taken as the second priority, and other third-party journal websites are taken as the third priority. In addition, the priorities can be further subdivided according to the historical data accuracy ratio of each journal website. Then, according to the set priority sequence, for the conflicting entity content, the content with a higher priority is adopted and retained.

[0042] In some embodiments of the present invention, Figure 4 is a schematic flowchart of knowledge inference and completion in an embodiment of the present invention. As Figure 4 shown, perform knowledge inference and completion on the disambiguated knowledge sequence to obtain integrated journal knowledge, including: S401. Determine the complementary associated entities and corresponding associated values of each missing entity in the disambiguated knowledge sequence based on a preset complementary association rule; S402. Perform knowledge inference based on the complementary associated entities and associated values to obtain the missing values of each missing entity, and complete each missing entity according to the missing values to obtain integrated journal knowledge.

[0043] Specifically, for journal knowledge, there is often a strong correlation between some entities. For example, there is often a correspondence between a publisher and a place of publication. The embodiments establish a preset complementary association rule based on these entities with strong correlations. When there is a missing entity, according to the preset complementary association rule, find the complementary associated entity corresponding to the missing entity and the associated value of the complementary associated entity, and then find the data in other journal data that has the complementary associated entity and the value is also the associated value. At this time, the missing value corresponding to the missing entity can be found according to this journal data. Taking the journal IET Computer Vision as an example, this journal is published by Wiley and the location is England. Therefore, in the case where the place of publication of another journal IET Intelligent Transport Systems is unknown, since both are published by Wiley, it can also be inferred that its place of publication should be England.

[0044] It should be noted that during the process of knowledge inference and completion, if there are multiple inferred missing values and there are conflicts among these missing values, the above priority strategy can also be used for selection.

[0045] In some embodiments of the present invention, a journal knowledge graph is constructed according to the integrated journal knowledge, including: Writing the integrated journal knowledge into a preset graph database based on the data layer update mode to obtain a journal knowledge graph.

[0046] Specifically, during the construction of the journal knowledge graph, the storage and visualization of the journal knowledge graph adopt the Neo4j graph database. This database represents entities and their relationships through property graphs, and has higher performance and query efficiency compared to other databases. Considering the timeliness of some journal information, a knowledge update process needs to be introduced during the construction of the journal knowledge graph to ensure the timeliness and accuracy of the knowledge in the knowledge graph. The traditional knowledge graph update process includes schema layer and data layer updates. Among them, the schema layer update is to update the data structure of the knowledge graph, and the data layer update is to add new entities or update the relationships of existing entities. In the embodiments, considering the actual needs of journal knowledge update, after the schema layer is set up in advance based on the requirements, it can no longer be adjusted, and the knowledge update process of the journal knowledge graph only involves the data layer update.

[0047] It should be noted that in the journal knowledge graph provided by the present invention, the type of knowledge may not be limited to text type, and can also integrate multi-modal data such as images, videos, and audios to further improve the intelligence and richness of the system. For example, in the academic field, some charts, research videos, or information in academic lectures can also be used as the basis for answering questions. Valuable information can be extracted from multi-modal data through technologies such as vision and speech processing.

[0048] In addition, considering that in journal knowledge question answering, the question answering system based on the traditional search engine usually relies on keyword matching. Although it is suitable for some simple question answering, when facing complex and professional questions that require reasoning, it often cannot give accurate answers. The question answering system based on the large language model is usually trained for a wide range of fields. Although it has a certain reasoning ability, when facing professional questions in a specific field, it is prone to inaccurate or incomplete answers due to the lack of domain-specific knowledge and usually cannot give in-depth answers. By retraining or fine-tuning the large language model, although the accuracy of the model's answers to professional questions in a specific field can be improved, due to the continuous expansion of the data scale of journal knowledge and the relatively fast update and iteration speed, the methods of retraining and fine-tuning the model require continuous model update training, and the efficiency is significantly insufficient, making it difficult to give accurate answer results for newly added knowledge.

[0049] Therefore, in the journal knowledge Q&A tasks, both the current traditional search engine Q&A system and the large language model Q&A system are difficult to give accurate and in-depth Q&A results. For this reason, the present invention also provides a journal knowledge Q&A method to achieve accurate and in-depth Q&A results in the journal knowledge Q&A tasks.

[0050] Figure 5 FIG. is a schematic flowchart of an embodiment of a journal knowledge Q&A method provided by the present invention, as Figure 5 shown, the journal knowledge Q&A method includes: S501. Input the journal knowledge question sentence to be answered into the large language model for natural language understanding to obtain a preliminary query semantics; Among them, to improve the accuracy and depth of journal knowledge Q&A, the embodiment provides a Q&A method that combines a large language model and a journal knowledge graph. In this method, the embodiment first performs natural language understanding on the journal knowledge question sentence to be answered through the large language model, so as to extract the key semantic content therein and obtain the preliminary query semantics. The large language model is relatively good at understanding natural language, can accurately extract the key semantic information in the user's question sentence, identify the question intention, and lay a foundation for the subsequent retrieval of journal knowledge.

[0051] S502. Perform knowledge query on the journal knowledge graph according to the preliminary query semantics to obtain relevant journal knowledge; Among them, the embodiment searches for relevant content in the knowledge graph according to the obtained preliminary query semantics. The knowledge graph has unique advantages in processing structured knowledge and logical reasoning, can clearly display the association and reasoning path between knowledge, and can be updated over time to ensure the real-time nature of the knowledge base. In addition, through the visualization of the graph structure, a large amount of knowledge and their mutual relationships can be intuitively presented, helping users better understand the complex knowledge network.

[0052] S503. Perform large language model Q&A output according to the preliminary query semantics and the relevant journal knowledge to obtain the journal knowledge Q&A output result.

[0053] Among them, finally, the embodiment inputs the preliminary query semantics and the searched relevant journal knowledge into the large language model again. The large language model generates a more accurate and domain-depth answer according to the supplemented relevant journal knowledge.

[0054] In addition, if no corresponding entity is matched in the knowledge base, then the large language model generates a default answer to avoid the situation of no response.

[0055] Finally, comprehensively sort the retrieved information and the generated content, and output the optimal answer.

[0056] For example, for the user's question "I have an article in the field of databases and want to submit it to a journal in the first zone of SCI. Can you recommend two for me?", when the large language model journal knowledge graph joint Q&A model receives this question, it first performs semantic analysis on the question to understand the user's intention and needs, including the theme of "database", the task of recommending SCI first-zone journals suitable for submission, and the need to recommend two journals. After parsing by the large language model, a preliminary query semantics is generated to ensure that the recommended journals belong to the first zone of SCI and are related to the database theme. After receiving the preliminary query semantics, the intelligent agent (Agent) determines that this question requires further interaction with the journal knowledge base to provide a more accurate answer. Since the internal knowledge base of the large language model cannot be updated in real time, the joint system queries the journal knowledge graph database to retrieve information related to "database" and "SCI first-zone journals". Finally, the retrieved information and the preliminary query semantics are sorted out through Q&A by the large language model again to output the journal knowledge Q&A output result.

[0057] In addition, when the question raised by the user involves a scope outside the knowledge base, the joint Q&A model can first generate an answer in its own vector knowledge base, extract knowledge from the generated answer content, and extract new knowledge points through semantic analysis. After these new knowledge points are reviewed and confirmed, they can be incorporated into the knowledge base to expand its content and improve the accuracy and efficiency of subsequent Q&A.

[0058] In some embodiments of the present invention, the journal knowledge Q&A method further includes: Optimizing the journal knowledge Q&A output result according to the obtained user portrait to obtain an optimized Q&A output result; Wherein, the user portrait includes user type, user historical behavior, and user preference.

[0059] Among them, to further combine the Q&A output result of the joint Q&A model, the embodiment can also optimize the output result according to the obtained user portrait. By combining the user's historical behavior and preference, a personalized intelligent Q&A system is customized, and through a detailed analysis of the user's needs, a tailored Q&A is provided. For example: for different user groups (such as scholars, students, or industry experts), the answers to different questions are optimized.

[0060] In summary, in the journal knowledge graph construction method provided by the present invention, through preliminary knowledge extraction of knowledge entities in journal data by sequence annotation method knowledge extraction, and alignment of knowledge entities from different data sources by the vocabulary network strategy, journal knowledge from different sources can be effectively integrated. By eliminating ambiguous content between entities through the priority strategy entity disambiguation and complementing missing values and blank values through knowledge reasoning, the knowledge content can be effectively verified, thereby improving the data quality of the journal knowledge graph.

[0061] To better implement the journal knowledge graph construction method in the embodiments of the present invention, correspondingly, on the basis of the journal knowledge graph construction method, as Figure 6 shown, the embodiments of the present invention further provide a journal knowledge graph construction device. The journal knowledge graph construction device 600 includes: A knowledge extraction unit 601, configured to perform knowledge extraction on the obtained journal text to be processed to obtain a knowledge annotation sequence; A knowledge integration unit 602, configured to perform entity alignment on the knowledge annotation sequence by using a lexical network strategy to obtain an aligned knowledge sequence, perform entity disambiguation on the aligned knowledge sequence by using a priority strategy to obtain a disambiguated knowledge sequence, and perform knowledge inference and completion on the disambiguated knowledge sequence to obtain integrated journal knowledge; A knowledge graph construction unit 603, configured to construct a journal knowledge graph according to the integrated journal knowledge; Among them, entity disambiguation by using a priority strategy includes performing entity disambiguation on the knowledge annotation sequence based on the priority of each entity, and knowledge inference and completion includes filling in missing values for the knowledge annotation sequence based on preset inference rules.

[0062] The above-described journal knowledge graph construction device 600 provided in the above embodiments can implement the technical solutions described in the embodiments of the above journal knowledge graph construction method. The specific implementation principles of the above modules or units can be referred to the corresponding content in the embodiments of the above journal knowledge graph construction method, and will not be elaborated here.

[0063] As Figure 7 shown, the present invention correspondingly further provides a journal intelligent question answering device 700. The journal intelligent question answering device 700 includes a processor 701, a memory 702, and a display 703. Figure 7 Only some components of the journal intelligent question answering device 700 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0064] In some embodiments, the processor 701 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is configured to run program codes stored in the memory 702 or process data, such as the journal knowledge graph construction method and / or the journal knowledge question answering method in the present invention.

[0065] In some embodiments, the processor 701 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 701 may be local or remote. In some embodiments, the processor 701 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-cloud, etc., or any combination of the above.

[0066] In some embodiments, the memory 702 may be an internal storage unit of the periodical intelligent Q&A device 700, such as the hard disk or memory of the periodical intelligent Q&A device 700. In some other embodiments, the memory 702 may also be an external storage device of the periodical intelligent Q&A device 700, such as a plug-in hard disk equipped on the periodical intelligent Q&A device 700, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0067] Furthermore, the memory 702 may also include both the internal storage unit and the external storage device of the periodical intelligent Q&A device 700. The memory 702 is used to store the application software installed in the periodical intelligent Q&A device 700 and various types of data.

[0068] In some embodiments, the display 703 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 703 is used to display the information of the periodical intelligent Q&A device 700 and to display a visual user interface. The components 701 - 703 of the periodical intelligent Q&A device 700 communicate with each other through a system bus.

[0069] In one embodiment, when the processor 701 executes the periodical knowledge graph construction program in the memory 702, the following steps may be implemented: Performing knowledge extraction on the obtained periodical text to be processed by the sequence annotation method to obtain a knowledge annotation sequence; Performing entity alignment on the knowledge annotation sequence by the lexical network strategy to obtain an aligned knowledge sequence, performing entity disambiguation on the aligned knowledge sequence by the priority strategy to obtain a disambiguated knowledge sequence, and performing knowledge reasoning and complementation on the disambiguated knowledge sequence to obtain integrated periodical knowledge; Performing knowledge graph construction based on the integrated periodical knowledge to obtain a periodical knowledge graph.

[0070] In another embodiment, when the processor 701 executes the periodical knowledge Q&A program in the memory 702, the following steps may be implemented: Inputting the periodical knowledge question sentence to be answered into a large language model for natural language understanding to obtain a preliminary query semantics; Performing knowledge query on the periodical knowledge graph according to the preliminary query semantics to obtain relevant periodical knowledge; Performing large language model Q&A output according to the preliminary query semantics and the relevant periodical knowledge to obtain the periodical knowledge Q&A output result.

[0071] It should be understood that when the processor 701 executes the periodical knowledge graph construction program in the memory 702, in addition to the above functions, other functions can also be realized. For details, reference can be made to the description of the corresponding method embodiments above.

[0072] Furthermore, the embodiment of the present invention does not specifically limit the type of the mentioned periodical intelligent question answering device 700. The periodical intelligent question answering device 700 can be a portable electronic device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop, etc. Exemplary embodiments of the portable electronic device include, but are not limited to, portable electronic devices equipped with IOS, android, microsoft or other operating systems. The above portable electronic device can also be other portable electronic devices, such as a laptop with a touch-sensitive surface (such as a touch panel). It should also be understood that in some other embodiments of the present invention, the periodical intelligent question answering device 700 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (such as a touch panel).

[0073] Correspondingly, the embodiment of the present application also provides a computer-readable storage medium, which is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps or functions in the periodical knowledge graph construction method and / or the periodical knowledge question answering method provided by the above method embodiments can be realized.

[0074] Those skilled in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory or a random access memory, etc.

[0075] The above has introduced in detail the periodical knowledge graph construction method, the periodical knowledge question answering method, the device, the intelligent recognition device and the storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for constructing a journal knowledge graph, characterized in that, Including: Performing knowledge extraction on the obtained journal text to be processed using the sequence annotation method to obtain a knowledge annotation sequence; Performing entity alignment on the knowledge annotation sequence using the lexical network strategy to obtain an aligned knowledge sequence, performing entity disambiguation on the aligned knowledge sequence using the priority strategy to obtain a disambiguated knowledge sequence, and performing knowledge inference and completion on the disambiguated knowledge sequence to obtain integrated journal knowledge; Constructing a knowledge graph based on the integrated journal knowledge to obtain a journal knowledge graph; Among them, the entity disambiguation using the priority strategy includes performing entity disambiguation on the knowledge annotation sequence based on the priority of each entity, and the knowledge inference and completion includes performing missing value completion on the knowledge annotation sequence based on preset inference rules.

2. The method for constructing a journal knowledge graph according to claim 1, wherein The journal text to be processed is obtained by filtering the semi-structured journal data and unstructured journal data through a crawler framework text filter.

3. The method for constructing a periodical knowledge graph according to claim 1, wherein The performing knowledge extraction on the obtained journal text to be processed using the sequence annotation method to obtain a knowledge annotation sequence includes: Performing bidirectional feature encoding on the journal text to be processed to obtain word vectors; Performing bidirectional long short-term memory network context extraction on the word vectors to obtain context features; Performing conditional random field output decoding on the context features to obtain a knowledge annotation sequence.

4. The method for constructing a journal knowledge graph according to claim 1, wherein The performing entity disambiguation on the aligned knowledge sequence using the priority strategy to obtain a disambiguated knowledge sequence includes: Determining the priority of each knowledge entity in the aligned knowledge sequence based on a preset priority sequence and the data source category corresponding to the aligned knowledge sequence; Performing entity disambiguation on the aligned knowledge sequence according to the priority of the knowledge entity to obtain a disambiguated knowledge sequence.

5. The method for constructing a journal knowledge graph according to claim 1, wherein The performing knowledge inference and completion on the disambiguated knowledge sequence to obtain integrated journal knowledge includes: Determining the complementary associated entity and the corresponding associated value of each missing entity in the disambiguated knowledge sequence based on a preset complementary association rule; Performing knowledge inference according to the complementary associated entity and the associated value to obtain the missing value of each missing entity, and completing each missing entity according to the missing value to obtain integrated journal knowledge.

6. The method for constructing a journal knowledge graph according to claim 1, wherein The constructing a knowledge graph based on the integrated journal knowledge to obtain a journal knowledge graph includes: Writing the integrated journal knowledge into a preset graph database based on the data layer update mode to obtain a journal knowledge graph.

7. A method for journal knowledge Q&A, characterized in that, Including: Inputting the journal knowledge question to be answered into a large language model for natural language understanding to obtain a preliminary query semantics; Performing knowledge query on the journal knowledge graph according to the preliminary query semantics to obtain relevant journal knowledge; Performing large language model question answering output according to the preliminary query semantics and the relevant journal knowledge to obtain a journal knowledge question answering output result; Among them, the journal knowledge graph is obtained according to the journal knowledge graph construction method described in any one of claims 1 to 6.

8. The periodical knowledge Q&A method according to claim 7, wherein, The method further includes: Optimizing the journal knowledge question answering output result according to the obtained user portrait to obtain an optimized question answering output result; Among them, the user portrait includes user type, user historical behavior, and user preference.

9. A device for constructing a periodical knowledge graph, characterized in that, Including: A knowledge extraction unit for performing knowledge extraction on the obtained journal text to be processed to obtain a knowledge annotation sequence; A knowledge integration unit, which is used to perform entity alignment of the lexical network strategy on the knowledge annotation sequence to obtain an aligned knowledge sequence, perform disambiguation of priority strategy entities on the aligned knowledge sequence to obtain a disambiguated knowledge sequence, and perform knowledge reasoning and completion on the disambiguated knowledge sequence to obtain integrated journal knowledge; A knowledge graph construction unit, which is used to construct a knowledge graph based on the integrated journal knowledge to obtain a journal knowledge graph; Among them, the disambiguation of priority strategy entities includes entity disambiguation of the knowledge annotation sequence based on the priority of each entity, and the knowledge reasoning and completion includes missing value completion of the knowledge annotation sequence based on preset reasoning rules.

10. An intelligent Q&A device for periodicals, characterized in that, It includes a processor and a memory. Among them, The memory is used to store programs; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the journal knowledge graph construction method described in any one of claims 1 to 6 above and / or the steps in the journal knowledge question and answer method described in any one of claims 7 to 8 above.

Citation Information

Patent Citations

  • Cyclically updated and iterated periodical literature knowledge graph construction method

    CN111209412A

  • Method and system of Chinese knowledge graph question-answering system based on semantic joint modeling

    CN115640391A

  • Power system knowledge graph construction method and device, computer equipment, readable storage medium and program product

    CN118674034A

  • Multi-source knowledge graph construction method and system for chronic disease diagnosis and treatment

    CN118820486A

Cited By

  • Data annotation overview method and device based on AI large model and medium

    CN121031819A

  • Periodical information retrieval method based on big data driving

    CN121980017A