A method, device and medium for extracting document knowledge elements based on large models
Through the document knowledge element extraction method based on large-models, the problems of poor generalization ability of knowledge element extraction and difficulty in field migration in the existing technology are solved, and more efficient and accurate knowledge element extraction and field adaptability are achieved.
Patent Information
- Application Number
- CN202510412502.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When faced with complex and changeable text content and multi-source data, the existing knowledge element extraction method has poor generalization capabilities and cannot effectively adapt to the diversity of different fields, and has problems such as poor interpretability and difficulty in field migration.
The document knowledge element extraction method based on large models is used to obtain original documents through multi-source interfaces, perform part-of-speech annotation and named entity recognition, build a domain knowledge graph, and map entities and relationships to low-dimensional vector space for data amplification. Then, based on the strategy of combining transfer learning and active learning, the big model is fine-tuned to obtain a domain adaptation model, which is used to extract knowledge elements and visually display it.
It improves the accuracy and generalization ability of knowledge element extraction, enhances the domain adaptability and interpretability of the model, and can more effectively process complex and changeable text content and multi-source data.
Smart Images

Figure CN119938946B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of knowledge graphs, and particularly to a method, device, and medium for extracting document knowledge elements based on large models. Background Art
[0002] In the application scenarios of knowledge management and natural language processing, accurately extracting knowledge elements from documents is the key foundation for realizing intelligent information processing. Currently, knowledge element extraction technology is widely used in multiple fields, such as intelligence analysis, academic research, business insights, etc. With the development of the digital age, the data sources are becoming increasingly rich and diverse, and the professional knowledge in different fields varies greatly. Therefore, how to accurately complete the extraction of knowledge elements has become a very challenging problem.
[0003] Existing knowledge element extraction methods, such as rule-based methods, although having high interpretability, are difficult to handle complex and variable text content, have poor generalization ability, and cannot adapt to the diversity brought by multi-source data. Machine learning-based methods rely on a large number of manually labeled features and have limited processing capabilities for long texts and complex contexts, being inefficient and ineffective when facing large-scale, multi-domain data. Deep learning-based methods, although having the ability of automatic feature learning and strong generalization ability, have problems such as poor interpretability and sensitivity to domain-specific data, and are difficult to be flexibly migrated and applied between different fields. Summary of the Invention
[0004] To solve the above technical problems, one or more embodiments of this specification provide a method, device, and medium for extracting document knowledge elements based on large models.
[0005] One or more embodiments of this specification adopt the following technical solutions:
[0006] One or more embodiments of this specification provide a method for extracting document knowledge elements based on large models, the method including:
[0007] Obtaining original documents in a specified field based on multi-source interfaces, and performing part-of-speech tagging and named entity recognition processing according to the document content and document field of each original document to obtain standardized documents; wherein, the document field is included within the specified field range;
[0008] Collecting domain terms corresponding to the document field, constructing a knowledge graph for the corresponding field, and mapping the entities and relationships in the knowledge graph for the corresponding field to a low-dimensional vector space for data augmentation;
[0009] Inputting the augmented knowledge graph for the field into a pre-set large model, and performing domain fine-tuning on the pre-set large model based on a strategy combining transfer learning and active learning to obtain a domain-adapted model;
[0010] Input the document to be extracted in the specified field into the domain adaptation model to extract the knowledge elements of the document to be extracted, and screen important knowledge elements based on the attention scores corresponding to each knowledge element for visual display.
[0011] Optionally, in one or more embodiments of the present specification, obtaining the original document in the specified field based on the multi-source interface specifically includes:
[0012] Obtain multi-modal data in the specified field based on the multi-source interface, and perform correlation merging on the multi-modal data based on the key identifiers corresponding to the specified field to obtain a multi-modal data set corresponding to the same key identifier;
[0013] Based on the data formats corresponding to each modal data in the multi-modal data set, call the corresponding format conversion tool to convert the data formats corresponding to each modal data into the corresponding standard format;
[0014] Match the corresponding data storage structure and storage method based on the data volume of each type of corresponding standard format data in each modal data;
[0015] Construct the original document in the specified field based on the modal data in the standard format, the data storage structure, and the storage method.
[0016] Optionally, in one or more embodiments of the present specification, perform part-of-speech tagging and named entity recognition processing on the document content and document field of each original document to obtain a standardized document, specifically including:
[0017] Obtain the specific word segmentation corresponding to the document field of the original document, and construct a custom dictionary corresponding to the original document based on the specific word segmentation;
[0018] Perform data cleaning on each original document based on regular expressions to obtain the cleaned original document, and perform character query on the cleaned original document to determine whether there are full-width characters or incorrect encoding formats in the cleaned original document;
[0019] Convert and correct the full-width characters and the incorrect encoding formats to obtain the original document to be segmented;
[0020] Perform word segmentation processing on the original document to be segmented through a preset word segmentation tool and the custom dictionary to obtain a plurality of word segmentation data;
[0021] Perform part-of-speech tagging on the word segmentation data based on a pre-trained language model, and perform entity tagging on the specific entities of the original document to be segmented based on the custom dictionary and a preset named entity recognition model;
[0022] Output the original document to be segmented, which has been processed by part-of-speech tagging and named entity recognition, based on a preset unified format to form a standardized document.
[0023] Optionally, in one or more embodiments of this specification, collect domain terms corresponding to the document domain and construct a knowledge graph for the corresponding domain, specifically including:
[0024] Based on the general knowledge graph corresponding to the document domain, determine the domain boundary and the general causal relationship between entities corresponding to the domain knowledge graph;
[0025] Collect text data corresponding to the document domain to extract fixed-pattern terms corresponding to the document domain from the text data based on regular expressions;
[0026] Based on the co-occurrence frequency of each term in the text data, construct a co-occurrence matrix to extract domain terms corresponding to the text data based on the co-occurrence matrix;
[0027] Based on the domain boundary, determine the knowledge graph type information corresponding to the document domain; wherein, the knowledge graph type information includes: entity type, relationship type, attribute type;
[0028] Extract knowledge corresponding to the knowledge graph type information from the fixed-pattern terms and the domain terms, and based on the relevant parts of the knowledge and the general knowledge graph, realize the fusion of the knowledge and the general knowledge graph to obtain a corresponding domain knowledge graph.
[0029] Optionally, in one or more embodiments of this specification, map the entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation, specifically including:
[0030] Based on the graph structure and relationship type of the domain knowledge graph, determine an embedding model that matches the domain knowledge graph; wherein, the embedding model includes: TransE, DistMult, ComplEx;
[0031] Based on the matching embedding model, map the entities and relationships of the domain knowledge graph to a low-dimensional vector space;
[0032] Obtain the entity vectors of each entity in the domain knowledge graph in the low-dimensional vector space, and based on the operation characteristics of the entity vectors, augment the entity vectors to obtain new entity vectors and determine the relationships corresponding to the new entity vectors based on the operation characteristics corresponding to the new entity vectors;
[0033] Based on the relationship between the newly added entity and the corresponding new entity vector, expand the knowledge graph of the corresponding field to obtain an expanded knowledge graph of the field.
[0034] Optionally, in one or more embodiments of this specification, input the expanded knowledge graph of the field into a pre-set large model, and perform field fine-tuning on the pre-set large model based on a strategy that combines transfer learning and active learning to obtain a field adaptation model. Specifically, it includes:
[0035] Collect existing knowledge graphs in different fields, map the existing knowledge graphs and the expanded knowledge graph of the field to a low-dimensional vector space, and determine semantically similar entities and relation-similar entities in different fields based on the similarity between vectors in the low-dimensional vector space;
[0036] Perform contrastive learning on the semantically similar entities and relation-similar entities based on a pre-set contrastive learning objective function to obtain knowledge data in each existing knowledge graph that is associated with the expanded knowledge graph of the field, and transfer the knowledge data to the expanded knowledge graph of the field based on the vector representation of the knowledge data to obtain a transferred knowledge graph of the field;
[0037] Parse and transform the transferred knowledge graph of the field based on the input format of the pre-set large model to adapt the transferred knowledge graph of the field to the input layer of the pre-set large model;
[0038] Perform prediction on the transformed knowledge graph of the field based on the existing model parameters of the pre-set large model, and determine the uncertain samples in the prediction results as candidate samples for active learning by calculating the entropy value of the prediction probability;
[0039] Train the pre-set large model based on the candidate samples to iteratively update the existing model parameters of the pre-set large model and obtain a field adaptation model that meets the requirements.
[0040] Optionally, in one or more embodiments of this specification, input the document to be extracted in the specified field into the field adaptation model to extract the knowledge elements of the document to be extracted. Specifically, it includes:
[0041] Preprocess the document to be extracted in the specified field, and determine the feature words of the document to be extracted based on the inverse document frequency value of each word segment in the document to be extracted in the document;
[0042] Construct a feature vector according to the inverse document frequency value corresponding to the feature words, and determine the document field of the document to be extracted through the matching degree between the feature vector and the pre-set domain term list;
[0043] Determine the domain adaptation model corresponding to the document domain of the document to be extracted, and convert the format of the document to be extracted based on the input requirements of the domain adaptation model, so as to input the converted document to be extracted into the domain adaptation model;
[0044] Perform knowledge element extraction on the converted document to be extracted based on the corresponding module of the domain adaptation model, so as to obtain the entities, entity relationships and category labels of each entity in the document to be extracted;
[0045] Integrate the entities, entity relationships and category labels of each entity in the document to be extracted to obtain the knowledge representation of the document to be extracted;
[0046] Output the knowledge representation based on a preset standard format to obtain the knowledge elements of the document to be extracted.
[0047] Optionally, in one or more embodiments of this specification, important knowledge elements are screened based on the attention scores corresponding to each knowledge element and visualized, which specifically includes:
[0048] Based on the attention mechanism module of the domain adaptation model, obtain the attention scores corresponding to each knowledge element, and determine the knowledge element threshold corresponding to the document to be extracted according to the attention score distribution of each knowledge element;
[0049] Determine the important knowledge elements corresponding to the document to be extracted through the knowledge element threshold, so as to obtain the confidence levels of the entities and entity relationships in the important knowledge elements, and visualize the important knowledge elements and the confidence levels based on a preset display tool.
[0050] One or more embodiments of this specification provide a device for extracting document knowledge elements based on a large model. The device includes:
[0051] At least one processor; and,
[0052] A memory communicatively connected to the at least one processor; wherein,
[0053] The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can: execute any one of the above methods.
[0054] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to: be able to execute any one of the above methods.
[0055] One or more of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:
[0056] Retrieve the original document based on multi-source interfaces, which can collect rich and diverse data, making the data sources more comprehensive and reducing the possibility of information loss. Construct a knowledge graph by collecting domain terms, which can structure the scattered knowledge in the domain, clarify the relationships between entities, and facilitate the understanding and utilization of domain knowledge. Map the entities and relationships in the knowledge graph to a low-dimensional vector space and perform data augmentation, reducing the data complexity while increasing data diversity, which helps to improve the generalization ability of subsequent models. Adopt a strategy that combines transfer learning and active learning to fine-tune the pre-set large model for the domain, enabling the model to quickly adapt to the domain characteristics and improve the accuracy and performance in specific domain tasks. Use the domain adaptation model to extract knowledge elements, which can process domain-specific texts more accurately and better identify domain terms. Visualize the selected important knowledge elements to present the abstract knowledge intuitively, helping users understand the model decision-making process and basis, and enhancing the interpretability. Description of the Drawings
[0057] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0058] Figure 1 It is a schematic flowchart of a method for extracting document knowledge elements based on a large model provided by an embodiment of this specification;
[0059] Figure 2 It is a schematic structural diagram of a device for extracting document knowledge elements based on a large model provided by an embodiment of this specification;
[0060] Figure 3 It is a schematic structural diagram of a non-volatile storage medium provided by an embodiment of this specification. Detailed Implementation Modes
[0061] Embodiments of this specification provide a method, device, and medium for extracting document knowledge elements based on a large model.
[0062] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0063] As Figure 1 shown, the embodiment of this specification provides a schematic flow chart of a method for extracting document knowledge elements based on a large model. As Figure 1 can be seen, in one or more embodiments of this specification, a method for extracting document knowledge elements based on a large model includes the following steps:
[0064] S101: Obtain the original documents of a specified domain based on multiple source interfaces, and perform part-of-speech tagging and named entity recognition processing according to the document content and document domain of each original document to obtain standardized documents; wherein, the document domain is included within the scope of the specified domain.
[0065] In real-world knowledge management and natural language processing application scenarios, they are complex and diverse, and multiple factors need to be considered comprehensively. Different data interfaces often cover information in different aspects or angles. For example, in the medical field, electronic medical record systems, medical literature databases, clinical research reports, etc. are all important data sources. Obtaining original documents from multiple such data sources can make the extracted knowledge elements more abundant and comprehensive, and avoid information loss caused by single data. Therefore, in the embodiment of this specification, the original problems of the specified domain will be obtained based on multiple source interfaces. In addition, the original documents from different sources may have format differences. For example, some are in plain text format, some are in web HTML format, and some are in document formats such as PDF. Therefore, in the embodiment of this specification, part-of-speech tagging and named entity recognition processing will be performed according to the document content and document domain of each original document to obtain standardized documents. Among them, it should be noted that the document domain is included within the scope of the specified domain. In this process, part-of-speech tagging can clarify the part of speech of each word in a sentence, such as noun, verb, adjective, etc. This helps the model better understand the grammatical structure and semantic information of the sentence, and provides a basis for subsequent tasks such as knowledge element extraction and relationship recognition. Named entity recognition can identify key entities such as person names, place names, organization names, dates, and amounts in the document. These entities are important components of knowledge elements, and accurately identifying them helps subsequent operations such as constructing a knowledge graph and extracting knowledge element relationships.
[0066] Specifically, in one or more embodiments of this specification, obtaining the original documents of a specified domain based on multiple source interfaces specifically includes:
[0067] First, obtain multimodal data in a specified domain based on multi-source interfaces, and then perform correlation and merging on the multimodal data according to the key identifiers corresponding to the specified domain to obtain a set of multimodal data corresponding to the same key identifier. For example, in the medical field, the key identifier is the patient ID. Through the patient ID, the medical record text of the same patient, the tabular data in the inspection report, the medical images, and the audio records of the doctor's diagnosis can be associated and integrated together. It can be seen that by using a unified key identifier such as a document ID or an event number, data from different sources and different modalities such as text, images, and tables are correlated and merged to ensure the complete aggregation of multi-dimensional information of the same entity. Moreover, it realizes the elimination of data islands and improves the collaborative analysis ability of cross-modal data. At the same time, it helps to enhance the accuracy of subsequent knowledge element extraction and the context understanding ability.
[0068] Then, in the embodiments of this specification, according to the data formats corresponding to the modal data in the multimodal data set, the corresponding format conversion tools are called to convert the data formats corresponding to the modal data into the corresponding standard formats. Call the corresponding conversion tools according to the formats of the modal data and convert them into standard formats. This not only improves the data compatibility, enabling different modal data to be processed under a unified framework. Then, based on the data volumes of various types of data in the standard formats corresponding to the modal data, the corresponding data storage structures and storage methods are matched, and then the modal data in the standard format is used to construct the original documents in the specified domain based on the data storage structure and the storage method.
[0069] Specifically, in one or more embodiments of this specification, the process of part-of-speech tagging and named entity recognition is performed on the document content and document domain of each original document to obtain standardized documents, which specifically includes the following processes:
[0070] There are significant differences in language expressions and vocabulary usage across different fields. By obtaining specific word segmentations corresponding to the fields of the original documents to construct a custom dictionary, the word segmentation tool can more accurately identify professional terms, specific names, etc. within the field. In the processing of medical documents, professional terms such as "coronary atherosclerosis" and "magnetic resonance imaging" can be accurately segmented, avoiding word segmentation errors caused by ordinary word segmentation rules. Therefore, first, specific word segmentations corresponding to the document fields of the original documents will be obtained, and then a custom dictionary corresponding to the original documents will be constructed based on the specific word segmentations. Data cleaning will be performed on each original document according to regular expressions to obtain the cleaned original document, and character queries will be performed on the cleaned original document to determine whether there are full-width characters or incorrect encoding formats. The full-width characters and incorrect encoding formats will be converted and corrected to obtain the original document to be word segmented. By using regular expressions to perform data cleaning on the original documents, irrelevant content such as special characters and HTML tags can be removed, effectively reducing noise interference. And the conversion and correction of full-width characters and incorrect encoding formats ensure unified text format and correct encoding, avoiding affecting subsequent processing due to format and encoding problems, and improving the quality and usability of the data. The original document to be word segmented will be word segmented by using a pre-set word segmentation tool and the custom dictionary to obtain multiple word segmentation data. The word segmentation data will be part-of-speech tagged based on a pre-trained language model, and specific entities in the original document to be word segmented will be entity tagged based on the custom dictionary and the pre-set named entity recognition model. The original document to be word segmented after part-of-speech tagging and named entity recognition processing will be output based on the pre-set unified format to form a standardized document. In this process, part-of-speech tagging the word segmentation data based on a pre-trained language model can clarify the part of speech of each word, such as nouns, verbs, adjectives, etc. This helps to deeply understand the grammatical structure and semantic information of the text, and provides richer semantic features for subsequent tasks such as information extraction and text analysis. Combining the custom dictionary and the pre-set named entity recognition model can accurately identify specific entities in the original document, which helps to quickly locate and extract key information, and improve the efficiency of understanding and analyzing the document content.
[0071] In a certain application scenario, part-of-speech tagging and named entity recognition are performed on the content and domain of each original document to obtain a standardized document, which can also be achieved based on the following process: First, clean the original document, that is, use regular expressions to remove irrelevant content such as special characters, HTML tags, and line breaks; normalize the document format, such as converting full-width characters to half-width characters; at the same time, handle encoding problems to ensure that the text is in a unified format. Then perform word segmentation and removal of stop words. First, select a suitable word segmentation tool, such as the jieba word segmentation tool for Chinese or the SpaCy word segmentation tool for English, to perform word segmentation; build a domain-specific custom dictionary to enhance the accuracy of word segmentation; remove words that are irrelevant to semantic understanding according to the stop word list, such as "de", "shi", "zai", etc. Then perform part-of-speech tagging and named entity recognition, that is, use a pre-trained language model (such as BERT, ERNIE) for part-of-speech tagging; combine the domain term list and the pre-trained NER model to identify specific entities for documents in different domains. For example: when the original document is a legal document, the specific entities identified include: identifying legal clauses, regulations names, court names, judgment results, etc.; when the original document is a medical document, the specific entities identified include: identifying drug names, disease names, medical terms; when the original document is a financial document, the specific entities identified include: identifying company names, stock codes, financial events, etc.
[0072] S102: Collect domain terms corresponding to the document domain, construct a corresponding domain knowledge graph, and map the entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation.
[0073] In a specified domain, knowledge is often scattered in various documents, databases, or different application systems, and in various forms. Therefore, in order to integrate fragmented knowledge and at the same time solve the problems of insufficient data or insufficient data features, this specification will collect domain terms corresponding to the document domain, construct a corresponding domain knowledge graph to structurally integrate the originally scattered and disordered knowledge in this domain, and map the entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation. Mapping the entities and relationships in the knowledge graph to a low-dimensional vector space realizes data dimensionality reduction, and then augmenting the knowledge graph in the low-dimensional vector space can generate more feature combinations, enabling subsequent model learning to learn more comprehensive knowledge patterns and enhancing the generalization ability of the model, so that it can better process and predict when facing new and unseen data.
[0074] Specifically, in one or more embodiments of this specification, collecting domain terms corresponding to the document domain and constructing a corresponding domain knowledge graph specifically includes the following steps:
[0075] Based on the general knowledge graph corresponding to the document domain, determine the domain boundary of the domain knowledge graph and the general causal relationships between entities. It can be understood that the general knowledge graph is a collection of knowledge with general applicability in this field, which contains a large number of entities, entity attributes, and various relationships between entities. This knowledge is collected, organized, and structured, with a certain degree of generality and comprehensiveness. In addition, when determining the domain boundary, it is oriented towards the document domain, and relevant knowledge content in the general knowledge graph is screened out. For example, for the medical document domain, knowledge related to human physiology, diseases, treatment methods, etc. is selected from the general knowledge graph. By analyzing the screened knowledge, clarify the core concepts, themes, and key entities included in this domain. For example, in the medical field, the core concepts may include disease names, symptoms, drugs, etc.; the key entities may include various disease entities, medical research institution entities, etc. Through these core concepts and entities, determine the scope boundary of the domain, that is, clarify which knowledge belongs to this domain and which does not.
[0076] In addition, for the general causal relationships between entities, since there is a large amount of relationship information between entities in the general knowledge graph. Analyze the associations between the entities involved in the document domain and look for possible causal relationship clues. For example, in the medical knowledge graph, there may be a causal relationship clue between the "smoking" entity and the "lung cancer" entity, because in general knowledge, smoking is considered an important factor leading to lung cancer. The process of determining the domain boundary and the general causal relationships between entities provides an important basis and framework for constructing or improving the domain knowledge graph. The domain boundary clarifies the scope and focus of the knowledge graph, making the knowledge graph more focused and accurate; the general causal relationships between entities enrich the semantic information of the knowledge graph, helping to better understand the knowledge logic and internal connections within the domain.
[0077] Then, collect the text data corresponding to the document domain, and extract the fixed-pattern terms corresponding to the document domain from the text data according to the regular expression. For example, in the legal domain, the legal clause "Article X, Paragraph Y" belongs to the fixed-pattern term. Then, based on the co-occurrence frequency of each term in the text data, construct a co-occurrence matrix, and extract the domain terms corresponding to the text data according to the co-occurrence matrix. Determine the type information of the knowledge graph corresponding to the document domain according to the domain boundary. Among them, the type information of the knowledge graph includes: entity type, relationship type, and attribute type. For example, when constructing a knowledge graph in the financial domain, clarify that "company", "stock", etc. are entity types, "hold", "issue", etc. are relationship types, and "company market value", "stock price", etc. are attribute types, so that the knowledge graph can accurately and orderly represent the knowledge in the financial domain. By extracting the knowledge corresponding to the type information of the knowledge graph from the fixed-pattern terms and domain terms, and according to the knowledge and the relevant parts of the general knowledge graph, realize the integration of the knowledge and the general knowledge graph, and obtain the knowledge graph corresponding to the domain.
[0078] In the above process, the general knowledge graph is used to determine the domain boundary of the domain knowledge graph and the general causal relationship between entities, which points out the direction for the entire construction process. This can accurately locate the scope of domain knowledge, avoid interference from irrelevant information, and ensure that the constructed domain knowledge graph focuses on the core content of a specific domain. Extracting fixed-pattern terms based on regular expressions can quickly screen out terms that conform to the domain characteristics from a large amount of text data. Combining the co-occurrence matrix to extract domain terms can efficiently and accurately obtain the domain terms corresponding to the document domain from a large amount of text data, ensuring the efficiency and accuracy of term extraction in the construction process of the domain knowledge graph. Integrating the extracted knowledge with the general knowledge graph not only makes full use of the rich resources and prior knowledge of the general knowledge graph, but also combines the professional knowledge of a specific domain, making the constructed domain knowledge graph both generally applicable and accurately reflect the uniqueness of the domain.
[0079] Specifically, in one or more embodiments of this specification, map the entities and relationships in the knowledge graph corresponding to the domain to a low-dimensional vector space for data augmentation, which specifically includes:
[0080] A domain knowledge graph consists of entities, the relationships between entities, and the graph structure formed by them. Different embedding models have different characteristics and applicable scenarios, and are suitable for knowledge graphs with different structures and relationship types. For example, the TransE model is simple and intuitive, and is suitable for processing knowledge graphs with simple relationships; the DistMult model performs well in processing symmetric relationships; the ComplEx model can handle more complex relationships. Therefore, in the embodiments of this specification, an embedding model matching the domain knowledge graph will be determined based on the graph structure and relationship type of the domain knowledge graph. Then, based on the matching embedding model, the entities and relationships of the domain knowledge graph are mapped into a low-dimensional vector space. In the low-dimensional vector space, each entity and relationship is represented by a vector, and these vectors contain the semantic information of the entities and relationships, realizing the transformation of the complex structure and relationships in the knowledge graph into operations between vectors. Then, after obtaining the entity vectors of each entity in the domain knowledge graph in the low-dimensional vector space, the entity vectors are amplified by using the operation characteristics of the entity vectors, such as addition, multiplication, dot product, etc. of the vectors, and new entity vectors are obtained. Based on the operation characteristics corresponding to the new entity vectors, the relationships corresponding to the new entity vectors are determined. According to the new entities and the relationships corresponding to the new entity vectors, the original domain knowledge graph is amplified, and the new entities and relationships are added to the knowledge graph to form an amplified domain knowledge graph. In this way, the scale and content of the knowledge graph are expanded, containing more semantic information and knowledge, and can better reflect various entities and relationships in the domain.
[0081] In this process, using the embedding model to map the entities and relationships of the domain knowledge graph into a low-dimensional vector space can reduce the data dimension while retaining the structure and semantic information of the knowledge graph. The low-dimensional vector representation makes the calculation more efficient, can quickly process large-scale knowledge graph data, and reduce storage and calculation costs. By amplifying the entity vectors based on the operation characteristics of the entity vectors, potential relationships that were not originally explicitly represented in the domain knowledge graph can be discovered. The determination of the new entity vectors and their corresponding relationships enriches the semantic information of the knowledge graph, helps to reveal the hidden connections between entities, and improves the integrity and accuracy of the knowledge graph. And amplifying the domain knowledge graph according to the new entities and the relationships corresponding to the new entity vectors realizes the automatic update and expansion of the knowledge graph.
[0082] S103: Input the amplified domain knowledge graph into a pre-set large model, and perform domain fine-tuning on the pre-set large model based on a strategy combining transfer learning and active learning to obtain a domain-adapted model.
[0083] Different fields have their unique characteristics such as terms, concepts, semantic relationships, and data distributions. In order to enable the model to adjust its own parameters according to domain knowledge, better adapt to the characteristics of a specific domain, and improve the accuracy and performance of the model in tasks of that domain. In the embodiments of this specification, the pre-trained large model will be input with the augmented domain knowledge graph, and then the pre-trained large model will be fine-tuned for the domain according to the strategy combining transfer learning and active learning to obtain a domain-adapted model.
[0084] Specifically, in one or more embodiments of this specification, the pre-trained large model is input with the augmented domain knowledge graph to perform domain fine-tuning on the pre-trained large model based on the strategy combining transfer learning and active learning to obtain a domain-adapted model, which specifically includes the following processes:
[0085] First, existing knowledge graphs in different fields are collected, and these knowledge graphs contain rich knowledge in their respective fields. Then, these existing knowledge graphs and the augmented target domain knowledge graph are mapped into a low-dimensional vector space together. In the low-dimensional vector space, each entity and relationship is represented by a vector. By calculating the similarity between vectors, entities with similar semantics and entities with similar relationships in different fields can be found. For example, the "disease" entity in the medical knowledge graph and the "pathological state" entity in the biological knowledge graph may be entities with similar semantics. Then, based on the pre-set contrastive learning objective function, contrastive learning is performed on the found entities with similar semantics and entities with similar relationships. The purpose of contrastive learning is to enable the model to learn the differences and commonalities between these similar entities and relationships, so as to obtain the knowledge data in each existing knowledge graph that is associated with the augmented domain knowledge graph. These knowledge data are represented in vector form and then migrated to the augmented domain knowledge graph, enriching the content of the domain knowledge graph and obtaining the migrated domain knowledge graph. For example, the knowledge data about the relationship between gene regulation and disease occurrence in the biological knowledge graph is migrated to the medical knowledge graph to improve the part of the medical knowledge graph about disease mechanisms.
[0086] Then, based on the input format of the pre - set large - model, the migrated domain knowledge graph is parsed and transformed to adapt the migrated domain knowledge graph to the input layer of the pre - set large - model, ensuring that the knowledge graph can be passed as valid input to the large - model for processing. The transformed domain knowledge graph is predicted using the existing model parameters of the pre - set large - model. During the prediction process, the uncertainty of the prediction result is measured by calculating the entropy value of the prediction probability. The larger the entropy value, the higher the uncertainty of the prediction result. At the same time, the samples with larger entropy values in the prediction results are determined as uncertain samples and used as candidate samples for active learning. The pre - set large - model is trained based on the determined candidate samples, and the existing model parameters of the pre - set large - model are iteratively updated through the training process. As the training continues with the candidate samples, the model gradually learns more features and rules about the domain knowledge graph, thereby improving the model's performance in this domain and finally obtaining a domain - adapted model that meets the requirements. This model can better handle relevant tasks in the target domain.
[0087] In this process, by mapping the existing knowledge graphs of different domains and the amplified domain knowledge graph to a low - dimensional vector space and determining semantically similar entities and relation - similar entities, the knowledge of multiple domains can be effectively integrated, enriching the content of the target - domain knowledge graph, providing more comprehensive information for the model, helping the model learn more general and rich features, and improving the model's performance and generalization ability in the target domain. Based on the pre - set contrastive learning objective function, contrastive learning is performed on the similar entities, which can specifically obtain the knowledge data related to the amplified domain knowledge graph within each existing knowledge graph. This method can automatically discover the potential connections between different - domain knowledge, accurately extract the knowledge valuable to the target domain, avoid the noise and interference brought by blindly fusing knowledge, and improve the accuracy and effectiveness of knowledge migration. By calculating the entropy value of the prediction probability to determine the candidate samples for active learning, it is possible to focus on the uncertain samples in the model's prediction results. These samples usually contain information that is difficult for the model to understand or classify. Only manually annotating these key samples can greatly reduce the annotation workload and improve the annotation efficiency compared to annotating a large number of random samples. At the same time, it can also make more effective use of human annotation resources and enhance the effect of model training.
[0088] Obtaining a domain adaptation model in a certain application scenario can also be achieved based on the following process: First, collect domain terms, crawl professional literature and databases to obtain relevant data, then use methods such as TF-IDF and Word2Vec to automatically extract high-frequency terms, and optimize the term list through manual review to obtain domain terms. Then, construct a domain knowledge graph, that is, use rule-based or machine learning methods to construct entities and their relationships. For example, in the medical field, relationships such as [disease - symptom] and [drug - indication] can be constructed. Next, perform data augmentation. In this process, domain synonyms can be used for replacement to achieve data replacement, generate new sentences based on templates to achieve data synthesis, add interferences such as spelling mistakes and synonymous replacements, and improve the robustness of the model to achieve noise injection. Then, the domain fine-tuning process is as follows: Manually annotate a small amount of high-quality domain data to achieve data annotation; use large models such as BERT and GPT and transfer learning techniques to fine-tune on the basis of a general pre-trained model, and use the Adam optimizer and learning rate adjustment strategy; then, based on evaluations such as F1-score, precision, recall, and K-fold cross-validation optimization, perform model evaluation to obtain a domain adaptation model.
[0089] S104: Input the document to be extracted in the specified domain into the domain adaptation model to extract the knowledge elements of the document to be extracted, and screen important knowledge elements based on the attention scores corresponding to each knowledge element for visual display.
[0090] After obtaining the domain adaptation model based on the above process, by inputting the document to be extracted in the specified domain into the domain adaptation model, the knowledge elements of the document to be extracted can be extracted, and important knowledge elements can be screened based on the attention scores corresponding to each knowledge element for visual display.
[0091] Specifically, in one or more embodiments of this specification, the document to be extracted in a specified domain is input into the domain adaptation model to extract the knowledge elements of the document to be extracted, which specifically includes: preprocessing the document to be extracted in the specified domain to determine the feature words of the document to be extracted based on the term frequency-inverse document frequency values of each word segment in the document to be extracted. Then, according to the term frequency-inverse document frequency values corresponding to the feature words, a feature vector is constructed, and the document domain of the document to be extracted is determined by the matching degree between the feature vector and the preset domain term list. The domain adaptation model corresponding to the document domain of the document to be extracted is determined, and the document to be extracted is format-converted based on the input requirements of the domain adaptation model to input the converted document to be extracted into the domain adaptation model. Based on the corresponding module of the domain adaptation model, knowledge element extraction is performed on the converted document to be extracted to obtain the entities, entity relationships, and category labels of each entity in the document to be extracted. The entities, entity relationships, and category labels of each entity in the document to be extracted are integrated to obtain the knowledge representation of the document to be extracted. Then, the knowledge representation is output based on the preset standard format to obtain the knowledge elements of the document to be extracted.
[0092] In this process, determining the feature words of the document to be extracted through the term frequency-inverse document frequency values can effectively filter out common but meaningless words, highlight the important and distinctive words in the document, and make the extracted feature words better represent the core content of the document, providing a solid foundation for subsequent domain judgment and knowledge extraction. For example, in academic documents, some commonly used function words, etc., will be filtered out, while key contents such as professional terms will be retained as feature words. Constructing a feature vector based on the TF-IDF values of the feature words and matching it with the preset domain term list to determine the document domain can more accurately classify the document into the appropriate domain. In addition, determining the corresponding domain adaptation model according to the document domain and performing format conversion according to the model input requirements ensure that the document to be extracted can be well adapted to the model. This enables the model to fully exert its performance, more effectively perform knowledge element extraction, avoids the problem of poor model performance caused by format incompatibility or domain mismatch, and improves the efficiency and accuracy of knowledge extraction. Integrating the extracted knowledge elements and outputting them based on the preset standard format improves the standardization and generality of knowledge.
[0093] In a certain application scenario, the extraction of knowledge elements is to use a fine-tuned large model to perform knowledge element extraction based on semantic understanding and context analysis, that is, first perform domain classification: use TF-IDF + Naive Bayes / BERT classifier to identify the document domain; calculate the document domain score based on the keyword matching in the domain term list.
[0094] Then, multi-label classification is performed: entity classification is carried out based on a pre-trained classification model (such as BERT-MLC); few-shot learning is adopted to improve the classification ability of rare categories. Sequence labeling is performed: the BILSTM-CRF or Transformers model is used for word-by-word labeling; the accuracy is improved by combining domain terms and a pre-trained NER model. Relation extraction is carried out: when based on rules, syntactic analysis such as dependency parsing is used to extract entity relations. When based on machine learning, a two-tower model is used to calculate entity similarity; an attention mechanism is adopted to obtain context dependencies. Nested entity processing is carried out: a hierarchical decoding method is adopted, where large entities are identified first and then sub-entities are refined; SpanBERT is used to improve the cross-sentence entity recognition ability. Complex relation processing is carried out: a graph convolutional network is used to model multi-entity relations; combined with rule post-processing to ensure logical consistency.
[0095] Specifically, in one or more embodiments of this specification, important knowledge elements are screened based on the attention scores corresponding to each knowledge element and visualized, which specifically includes:
[0096] Based on the attention mechanism module of the domain adaptation model, the attention scores corresponding to each knowledge element are obtained to determine the knowledge element threshold corresponding to the document to be extracted according to the attention score distribution of each knowledge element. The attention score reflects the degree of attention of the model to this knowledge element during the extraction process. The higher the score, the more important the model considers this knowledge element in the document. Therefore, the attention score of each knowledge element is compared with the determined threshold, and the knowledge elements greater than or equal to the threshold are identified as important knowledge elements. For the entities and entity relations involved in the important knowledge elements, their confidence levels are further obtained. Then, using a pre-set display tool, the determined important knowledge elements and their corresponding confidence levels are presented in an intuitive manner.
[0097] In a certain application scenario, the process of improving the interpretability of the model through the attention mechanism and visualization technology can also be implemented based on the following methods: Attention weight visualization: Use the Attention Score in the Transformer model to visualize important text segments; calculate the text contribution score to identify key text. Confidence scoring: Use Softmax probability to calculate entity and relation confidence levels; use Bayesian uncertainty estimation to improve the credibility assessment. Visualization tool integration: Use TensorBoard to monitor the model's Attention Map; develop a custom Web tool to display knowledge elements and their relationship graphs.
[0098] As Figure 2 shown, an embodiment of this specification provides a schematic structural diagram of a device for extracting document knowledge elements based on a large model. From Figure 2 it can be seen that in one or more embodiments of this specification, a device for extracting document knowledge elements based on a large model includes:
[0099] at least one processor; and,
[0100] a memory communicatively connected to the at least one processor; wherein,
[0101] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: execute any one of the above-mentioned methods.
[0102] As Figure 3 shown, an embodiment of this specification provides a schematic structural diagram of a non-volatile storage medium. It can be seen from Figure 3 that in one or more embodiments of this specification, a non-volatile storage medium stores computer-executable instructions 301, and the computer-executable instructions are capable of: executing any one of the above-mentioned methods.
[0103] The various embodiments in this specification are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0104] The above specifically describes certain embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0105] The above is only one or more embodiments of this specification and is not used to limit this specification. For those skilled in the art, there can be various modifications and changes to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A document knowledge element extraction method based on a large model, characterized in that: The method comprises: Acquire original documents in a specified field based on a multi-source interface, perform part-of-speech tagging and named entity recognition processing according to the document content and document field of each original document, so as to obtain a standardized document; wherein the document field is included in the scope of the specified field; Collect domain terms corresponding to the document domain, construct a knowledge graph of the corresponding domain, and map entities and relationships in the knowledge graph of the corresponding domain to a low-dimensional vector space for data augmentation; Input a preset large model according to the amplified domain knowledge graph, and fine-tune the preset large model in the domain by combining transfer learning with active learning to obtain a domain adaptation model; Inputting the documents to be extracted in the specified field into the field adaptation model to extract knowledge elements of the documents to be extracted, and selecting important knowledge elements based on the attention scores corresponding to the knowledge elements for visual display; According to the augmented domain knowledge graph, the preset large model is input, and the preset large model is fine-tuned in the domain by a strategy based on combining transfer learning with active learning to obtain a domain adaptation model, which specifically includes: Collect existing knowledge graphs in different fields to map the existing knowledge graphs and the amplified domain knowledge graphs into a low-dimensional vector space, and determine semantically similar entities and relationally similar entities in different fields based on the similarity between vectors in the low-dimensional vector space; Based on a preset comparative learning objective function, the semantically similar entities and relationally similar entities are comparatively learned to obtain knowledge data associated with the augmented domain knowledge graph in each existing knowledge graph, and the knowledge data is migrated to the augmented domain knowledge graph based on the vector representation of the knowledge data to obtain the migrated domain knowledge graph; Based on the input format of the preset large model, the migrated domain knowledge graph is parsed and converted to adapt the migrated domain knowledge graph to the input layer of the preset large model; Predicting the converted domain knowledge graph based on the existing model parameters of the preset large model, so as to determine the uncertain samples in the prediction results as candidate samples for active learning by calculating the entropy value of the prediction probability; The preset large model is trained based on the candidate samples to iteratively update the existing model parameters of the preset large model to obtain a domain adaptation model that meets the requirements.
2. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Obtain original documents in the specified field based on multi-source interfaces, including: Acquire multimodal data of a specified field based on a multi-source interface, and associate and merge the multimodal data based on a key identifier corresponding to the specified field to obtain a multimodal data set corresponding to the same key identifier; Based on the data format corresponding to each modal data in the multimodal data set, calling a corresponding format conversion tool to convert the data format corresponding to each modal data into a corresponding standard format; Based on the amount of data in each corresponding standard format in each modality data, match the corresponding data storage structure and storage method; The modal data in the standard format is used to construct the original document of the specified field based on the data storage structure and the storage method.
3. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Part-of-speech tagging and named entity recognition are performed based on the document content and document domain of each original document to obtain standardized documents, including: Acquire specific word segments corresponding to the document field of the original document, so as to construct a custom dictionary corresponding to the original document based on the specific word segments; Performing data cleaning on each original document based on a regular expression to obtain a cleaned original document, and performing a character query on the cleaned original document to determine whether the cleaned original document contains full-width characters or an erroneous encoding format; Convert and correct the full-width characters and the erroneous encoding format to obtain the original document to be segmented; Performing word segmentation processing on the original document to be segmented by using a preset word segmentation tool and the custom dictionary to obtain a plurality of word segmentation data; Perform part-of-speech tagging on the word segmentation data based on the pre-trained language model, and identify specific entities in the original document to be segmented for entity tagging based on the custom dictionary and the preset named entity recognition model; The original document to be segmented after part-of-speech tagging and named entity recognition is output based on a preset unified format to form a standardized document.
4. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Collect domain terms corresponding to the document domain and construct a corresponding domain knowledge graph, including: Based on the general knowledge graph corresponding to the document domain, determining the general causal relationship between the domain boundary and entities corresponding to the domain knowledge graph; Collecting text data corresponding to the document domain to extract fixed pattern terms corresponding to the document domain from the text data based on regular expressions; Based on the co-occurrence frequency of each term in the text data, a co-occurrence matrix is constructed to extract the field terms corresponding to the text data based on the co-occurrence matrix; Determine the knowledge graph type information corresponding to the document domain based on the domain boundary; wherein the knowledge graph type information includes: entity type, relationship type, and attribute type; Extract the knowledge corresponding to the knowledge graph type information from the fixed pattern terms and the domain terms, so as to achieve the fusion of the knowledge and the general knowledge graph based on the relevant parts of the knowledge and the general knowledge graph, and obtain the corresponding domain knowledge graph.
5. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Mapping entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation includes: Based on the graph structure and relationship type of the domain knowledge graph, determine an embedding model that matches the domain knowledge graph; wherein the embedding model includes: TransE, DistMult, ComplEx; Based on the matched embedding model, mapping entities and relationships of the domain knowledge graph to a low-dimensional vector space; Obtaining entity vectors of each entity in the domain knowledge graph in the low-dimensional vector space, amplifying the entity vectors based on the computational characteristics of the entity vectors, obtaining newly added entity vectors, and determining the relationship corresponding to the newly added entity vectors based on the computational characteristics corresponding to the newly added entity vectors; Based on the relationship between the newly added entity and the newly added entity vector, the corresponding domain knowledge graph is amplified to obtain an amplified domain knowledge graph.
6. The document knowledge element extraction method based on a large model according to claim 1 is characterized in that: Inputting the document to be extracted in the specified domain into the domain adaptation model to extract the knowledge element of the document to be extracted specifically includes: Preprocessing the documents to be extracted in the specified field to determine the characteristic words of the documents to be extracted based on the word frequency inverse document frequency value of each word segment in the documents to be extracted; According to the word frequency inverse document frequency value corresponding to the feature word, a feature vector is constructed to determine the document domain of the document to be extracted through the matching degree between the feature vector and a preset domain term list; Determine a domain adaptation model corresponding to the document domain of the document to be extracted, and convert the format of the document to be extracted based on the input requirements of the domain adaptation model, so as to input the converted document to be extracted into the domain adaptation model; Extracting knowledge elements from the converted document to be extracted based on the corresponding module of the domain adaptation model to obtain entities, entity relationships and category labels of each entity of the document to be extracted; Integrate the entities, entity relationships and category labels of the document to be extracted to obtain the knowledge representation of the document to be extracted; The knowledge representation is output based on a preset standardized format to obtain the knowledge element of the document to be extracted.
7. The document knowledge element extraction method based on a large model according to claim 1 is characterized in that: Based on the attention scores corresponding to each knowledge element, important knowledge elements are selected and visualized, including: Based on the attention mechanism module of the domain adaptation model, the attention score corresponding to each knowledge element is obtained, so as to determine the knowledge element threshold corresponding to the document to be extracted according to the attention score distribution of each knowledge element; The important knowledge element corresponding to the document to be extracted is determined by the knowledge element threshold to obtain the confidence of the entity and entity relationship in the important knowledge element, so as to visualize the important knowledge element and the confidence based on a preset display tool.
8. A document knowledge element extraction device based on a large model, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: execute any of the methods described in claims 1-7.
9. A non-volatile storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Unstructured text data knowledge extraction method based on large language model
CN118036734A
Enterprise-level knowledge base construction method based on large model
CN119622040A