Document knowledge element extraction method and device based on large model and medium
Through the document knowledge element extraction method based on large-models, the problems of poor generalization ability and interpretability in complex, variable texts and multi-source data are solved, and more accurate and efficient knowledge element extraction and domain adaptability are achieved.
Patent Information
- Application Number
- CN202510412502.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When facing complex and changeable text content and multi-source data, the existing knowledge element extraction methods have poor generalization capabilities, cannot adapt to diversity, and have problems such as poor interpretability and sensitivity to domain-specific data.
The document knowledge element extraction method based on large models is used to obtain original documents through multi-source interfaces, perform part-of-speech annotation and named entity recognition, build a domain knowledge graph, and map entities and relationships to low-dimensional vector space for data amplification. Then, based on a strategy combining transfer learning and active learning, the big model is fine-tuned in domains, and a domain adaptation model is obtained, which is used to extract knowledge elements and visually display it.
It improves the accuracy and generalization ability of knowledge element extraction, enhances the domain adaptability and interpretability of the model, and can better handle complex and diverse text data.
Smart Images

Figure CN119938946A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of knowledge graph technology, and in particular to a document knowledge element extraction method, device and medium based on a large model. Background Art
[0002] In the application scenarios of knowledge management and natural language processing, accurately extracting knowledge elements from documents is the key foundation for realizing intelligent information processing. Currently, knowledge element extraction technology is widely used in many fields, such as intelligence analysis, academic research, business insights, etc. With the development of the digital age, data sources are becoming more and more diverse, and professional knowledge in different fields varies greatly. Therefore, how to accurately extract knowledge elements has become a very challenging problem.
[0003] Although existing knowledge element extraction methods such as rule-based methods have high interpretability, they are difficult to cope with complex and changing text content, have poor generalization ability, and cannot adapt to the diversity brought by multi-source data. Machine learning-based methods rely on a large number of manually annotated features and have limited processing capabilities for long texts and complex contexts. They are inefficient and ineffective when faced with large-scale, multi-domain data. Although deep learning-based methods have automatic feature learning capabilities and strong generalization capabilities, they have problems such as poor interpretability and sensitivity to domain-specific data, making them difficult to flexibly migrate and apply between different fields. Summary of the invention
[0004] In order to solve the above technical problems, one or more embodiments of this specification provide a document knowledge element extraction method, device and medium based on a large model.
[0005] One or more embodiments of this specification adopt the following technical solutions: One or more embodiments of this specification provide a document knowledge element extraction method based on a large model, the method comprising: Acquire original documents in a specified field based on a multi-source interface, perform part-of-speech tagging and named entity recognition processing according to the document content and document field of each original document, so as to obtain a standardized document; wherein the document field is included in the scope of the specified field; Collect domain terms corresponding to the document domain, construct a knowledge graph of the corresponding domain, and map entities and relationships in the knowledge graph of the corresponding domain to a low-dimensional vector space for data augmentation; Input a preset large model according to the amplified domain knowledge graph, and fine-tune the preset large model in the domain by combining transfer learning with active learning to obtain a domain adaptation model; The documents to be extracted in the specified field are input into the field adaptation model to extract the knowledge elements of the documents to be extracted, and important knowledge elements are screened based on the attention scores corresponding to the knowledge elements for visual display.
[0006] Optionally, in one or more embodiments of the present specification, obtaining original documents in a specified field based on a multi-source interface specifically includes: Acquire multimodal data of a specified field based on a multi-source interface, and associate and merge the multimodal data based on a key identifier corresponding to the specified field to obtain a multimodal data set corresponding to the same key identifier; Based on the data format corresponding to each modal data in the multimodal data set, calling a corresponding format conversion tool to convert the data format corresponding to each modal data into a corresponding standard format; Based on the amount of data in each corresponding standard format in each modality data, match the corresponding data storage structure and storage method; The modal data in the standard format is used to construct the original document of the specified field based on the data storage structure and the storage method.
[0007] Optionally, in one or more embodiments of the present specification, part-of-speech tagging and named entity recognition are performed according to the document content and document domain of each original document to obtain a standardized document, specifically including: Acquire specific word segments corresponding to the document field of the original document, so as to construct a custom dictionary corresponding to the original document based on the specific word segments; Performing data cleaning on each original document based on a regular expression to obtain a cleaned original document, and performing a character query on the cleaned original document to determine whether the cleaned original document contains full-width characters or an erroneous encoding format; Convert and correct the full-width characters and the erroneous encoding format to obtain the original document to be segmented; Performing word segmentation processing on the original document to be segmented by using a preset word segmentation tool and the custom dictionary to obtain a plurality of word segmentation data; Perform part-of-speech tagging on the word segmentation data based on the pre-trained language model, and identify specific entities in the original document to be segmented for entity tagging based on the custom dictionary and the preset named entity recognition model; The original document to be segmented after part-of-speech tagging and named entity recognition is output based on a preset unified format to form a standardized document.
[0008] Optionally, in one or more embodiments of the present specification, collecting domain terms corresponding to the document domain and constructing a corresponding domain knowledge graph specifically includes: Based on the general knowledge graph corresponding to the document domain, determining the general causal relationship between the domain boundary and entities corresponding to the domain knowledge graph; Collecting text data corresponding to the document domain to extract fixed pattern terms corresponding to the document domain from the text data based on regular expressions; Based on the co-occurrence frequency of each term in the text data, a co-occurrence matrix is constructed to extract the field terms corresponding to the text data based on the co-occurrence matrix; Determine the knowledge graph type information corresponding to the document domain based on the domain boundary; wherein the knowledge graph type information includes: entity type, relationship type, and attribute type; Extract the knowledge corresponding to the knowledge graph type information from the fixed pattern terms and the domain terms, so as to achieve the fusion of the knowledge and the general knowledge graph based on the relevant parts of the knowledge and the general knowledge graph, and obtain the corresponding domain knowledge graph.
[0009] Optionally, in one or more embodiments of the present specification, mapping entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation specifically includes: Based on the graph structure and relationship type of the domain knowledge graph, determine an embedding model that matches the domain knowledge graph; wherein the embedding model includes: TransE, DistMult, ComplEx; Based on the matched embedding model, mapping entities and relationships of the domain knowledge graph to a low-dimensional vector space; Obtaining entity vectors of each entity in the domain knowledge graph in the low-dimensional vector space, amplifying the entity vectors based on the computational characteristics of the entity vectors, obtaining newly added entity vectors, and determining the relationship corresponding to the newly added entity vectors based on the computational characteristics corresponding to the newly added entity vectors; Based on the relationship between the newly added entity and the newly added entity vector, the corresponding domain knowledge graph is amplified to obtain an amplified domain knowledge graph.
[0010] Optionally, in one or more embodiments of the present specification, a preset large model is input according to the augmented domain knowledge graph, and the preset large model is fine-tuned in the domain based on a strategy combining transfer learning and active learning to obtain a domain adaptation model, specifically including: Collect existing knowledge graphs in different fields to map the existing knowledge graphs and the amplified domain knowledge graphs into a low-dimensional vector space, and determine semantically similar entities and relationally similar entities in different fields based on the similarity between vectors in the low-dimensional vector space; Based on a preset comparative learning objective function, the semantically similar entities and relationally similar entities are comparatively learned to obtain knowledge data associated with the augmented domain knowledge graph in each existing knowledge graph, and the knowledge data is migrated to the augmented domain knowledge graph based on the vector representation of the knowledge data to obtain the migrated domain knowledge graph; Based on the input format of the preset large model, the migrated domain knowledge graph is parsed and converted to adapt the migrated domain knowledge graph to the input layer of the preset large model; Predicting the converted domain knowledge graph based on the existing model parameters of the preset large model, so as to determine the uncertain samples in the prediction results as candidate samples for active learning by calculating the entropy value of the prediction probability; The preset large model is trained based on the candidate samples to iteratively update the existing model parameters of the preset large model to obtain a domain adaptation model that meets the requirements.
[0011] Optionally, in one or more embodiments of the present specification, inputting the document to be extracted in the specified field into the field adaptation model to extract the knowledge element of the document to be extracted specifically includes: Preprocessing the documents to be extracted in the specified field to determine the characteristic words of the documents to be extracted based on the word frequency inverse document frequency value of each word segment in the documents to be extracted; According to the word frequency inverse document frequency value corresponding to the feature word, a feature vector is constructed to determine the document domain of the document to be extracted through the matching degree between the feature vector and a preset domain term list; Determine a domain adaptation model corresponding to the document domain of the document to be extracted, and convert the format of the document to be extracted based on the input requirements of the domain adaptation model, so as to input the converted document to be extracted into the domain adaptation model; Extracting knowledge elements from the converted document to be extracted based on the corresponding module of the domain adaptation model to obtain entities, entity relationships and category labels of each entity of the document to be extracted; Integrate the entities, entity relationships and category labels of the document to be extracted to obtain the knowledge representation of the document to be extracted; The knowledge representation is output based on a preset standardized format to obtain the knowledge element of the document to be extracted.
[0012] Optionally, in one or more embodiments of the present specification, important knowledge elements are selected based on the attention scores corresponding to each knowledge element, and are visualized, specifically including: Based on the attention mechanism module of the domain adaptation model, the attention score corresponding to each knowledge element is obtained, so as to determine the knowledge element threshold corresponding to the document to be extracted according to the attention score distribution of each knowledge element; The important knowledge element corresponding to the document to be extracted is determined by the knowledge element threshold to obtain the confidence of the entity and entity relationship in the important knowledge element, so as to visualize the important knowledge element and the confidence based on a preset display tool.
[0013] One or more embodiments of this specification provide a document knowledge element extraction device based on a large model, the device comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: execute any of the above-mentioned methods.
[0014] One or more embodiments of the present specification provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute any of the above-described methods.
[0015] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: Obtaining original documents based on multi-source interfaces can collect rich and diverse data, making the data source more comprehensive and reducing the possibility of information missing. Collecting domain terms to build a knowledge graph can structure the scattered knowledge in the field, clarify the relationship between entities, and facilitate the understanding and use of domain knowledge. Mapping entities and relationships in the knowledge graph to a low-dimensional vector space and performing data augmentation can reduce data complexity while increasing data diversity, which helps to improve the generalization ability of subsequent models. Using a strategy that combines transfer learning and active learning to fine-tune the domain of the preset large model can enable the model to quickly adapt to the characteristics of the domain and improve the accuracy and performance in specific domain tasks. Using a domain adaptation model to extract knowledge elements can more accurately process domain-specific texts and better identify domain terms. Visualizing the important knowledge elements after screening can intuitively present abstract knowledge, help users understand the model decision-making process and basis, and improve interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings: Figure 1 A schematic diagram of a method flow of a document knowledge element extraction method based on a large model provided in an embodiment of this specification; Figure 2 A schematic diagram of the structure of a document knowledge element extraction device based on a large model provided in an embodiment of this specification; Figure 3 A schematic diagram of the structure of a non-volatile storage medium provided in an embodiment of this specification. DETAILED DESCRIPTION
[0017] The embodiments of this specification provide a method, device and medium for extracting document knowledge elements based on a large model.
[0018] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0019] like Figure 1 As shown, the embodiment of this specification provides a method flow chart of a document knowledge element extraction method based on a large model. Figure 1 It can be seen that in one or more embodiments of this specification, a document knowledge element extraction method based on a large model includes the following steps: S101: Acquire original documents of a specified domain based on a multi-source interface, perform part-of-speech tagging and named entity recognition according to the document content and document domain of each original document, so as to obtain a standardized document; wherein the document domain is included in the scope of the specified domain.
[0020] The application scenarios of knowledge management and natural language processing in reality are complex and diverse, and many factors need to be considered comprehensively. Different data interfaces often cover information from different aspects or angles. For example, in the medical field, electronic medical record systems, medical literature databases, clinical research reports, etc. are all important data sources. Obtaining original documents from multiple such data sources can make the extracted knowledge elements richer and more comprehensive, and avoid information loss caused by single data. Therefore, in the embodiments of this specification, the original questions in the specified field will be obtained based on the multi-source interface. In addition, there may be format differences in original documents from different sources, such as plain text format, web page HTML format, and PDF and other document formats. Therefore, in the embodiments of this specification, part-of-speech tagging and named entity recognition will be processed according to the document content and document field of each original document to obtain standardized documents. Among them, it should be noted that the document field is included in the scope of the specified field. In this process, part-of-speech tagging can clarify the part of speech of each word in the sentence, such as noun, verb, adjective, etc. This helps the model to better understand the grammatical structure and semantic information of the sentence, and provides a basis for subsequent tasks such as knowledge element extraction and relationship recognition. Named entity recognition can identify key entities in documents, such as names of people, places, names of organizations, dates, amounts, etc. These entities are important components of knowledge elements, and accurately identifying them helps in subsequent operations such as building knowledge graphs and extracting knowledge element relationships.
[0021] Specifically, in one or more embodiments of the present specification, obtaining original documents in a specified field based on a multi-source interface specifically includes: First, the multimodal data of the specified field is obtained based on the multi-source interface, and then the multimodal data is associated and merged according to the key identifier corresponding to the specified field to obtain a multimodal data set corresponding to the same key identifier. For example, in the medical field, the key identifier is the patient ID. Through the patient ID, the medical record text, the table data in the examination report, the medical image, and the audio record of the doctor's diagnosis of the same patient can be associated and integrated together. It can be seen that through unified key identifiers such as document IDs, event numbers, etc., data from different sources and different modalities such as text, images, tables, etc. are associated and merged to ensure the complete aggregation of multi-dimensional information of the same entity. It also achieves the elimination of data silos and improves the collaborative analysis capabilities of cross-modal data. At the same time, it helps to enhance the accuracy of subsequent knowledge element extraction and context understanding capabilities.
[0022] Then, in the embodiment of this specification, the corresponding format conversion tool will be called according to the data format corresponding to each modal data in the multimodal data set to convert the data format corresponding to each modal data into the corresponding standard format. The corresponding conversion tool is called according to the format of each modal data to convert it into a standard format. This not only improves the compatibility of the data, but also enables data of different modalities to be processed under a unified framework. Then, based on the amount of data in each corresponding standard format in each modal data, the corresponding data storage structure and storage method are matched, and then the modal data in the standard format is constructed based on the data storage structure and the storage method to construct the original document of the specified field.
[0023] Specifically, in one or more embodiments of the present specification, part-of-speech tagging and named entity recognition are performed according to the document content and document domain of each original document to obtain a standardized document, which specifically includes the following process: The language expression and vocabulary usage in different fields vary greatly. By obtaining the specific segmentation words of the original document corresponding to the field to build a custom dictionary, the segmentation tool can more accurately identify professional terms, specific names, etc. in the field. In medical document processing, professional terms such as "coronary atherosclerosis" and "magnetic resonance imaging" can be accurately segmented to avoid segmentation errors caused by ordinary segmentation rules. Therefore, first of all, the specific segmentation words corresponding to the document field of the original document will be obtained, and a custom dictionary corresponding to the original document will be built based on the specific segmentation words. Data cleaning is performed on each original document according to regular expressions to obtain the cleaned original document, and character query is performed on the cleaned original document to determine whether the cleaned original document contains full-width characters or incorrect encoding formats. Full-width characters and incorrect encoding formats are converted and corrected to obtain the original document to be segmented. By using regular expressions to clean the original document, irrelevant content such as special characters and HTML tags can be removed, effectively reducing noise interference. Conversion and correction of full-width characters and incorrect encoding formats ensures that the text format is unified and the encoding is correct, avoiding the impact of format and encoding problems on subsequent processing and improving the quality and availability of data. The original document to be segmented is segmented through the preset word segmentation tool and the custom dictionary to obtain multiple word segmentation data. The word segmentation data is tagged with parts of speech based on the pre-trained language model, and the specific entities of the original document to be segmented are identified for entity tagging based on the custom dictionary and the preset named entity recognition model. The original document to be segmented after the part-of-speech tagging and named entity recognition processing is output based on the preset unified format to form a standardized document. In this process, the word segmentation data is tagged with parts of speech based on the pre-trained language model, which can clarify the part of speech of each word, such as noun, verb, adjective, etc. This helps to deeply understand the grammatical structure and semantic information of the text, and provide richer semantic features for subsequent information extraction, text analysis and other tasks. Combined with the custom dictionary and the preset named entity recognition model, specific entities in the original document can be accurately identified, which helps to quickly locate and extract key information and improve the efficiency of understanding and analyzing the document content.
[0024] In a certain application scenario, part-of-speech tagging and named entity recognition are performed on the document content and document fields of each original document to obtain a standardized document, which can also be achieved based on the following process: First, clean the original document, that is, use regular expressions to remove irrelevant content such as special characters, HTML tags, and line breaks; normalize the document format, such as converting full-width characters to half-width characters; at the same time, handle encoding problems to ensure that the text is in a unified format. Then perform word segmentation and removal of stop words. First, select a suitable word segmentation tool, such as the jieba word segmentation tool suitable for Chinese or the SpaCy word segmentation tool suitable for English for word segmentation; build a domain-specific custom dictionary to enhance the accuracy of word segmentation; remove words irrelevant to semantic understanding according to the stop word list, such as "de", "shi", "zai", etc. Then perform part-of-speech tagging and named entity recognition, that is, use a pre-trained language model (such as BERT, ERNIE) for part-of-speech tagging; combine the domain term list and the pre-trained NER model to identify specific entities for documents in different fields. For example: when the original document is a legal document, the specific entities identified include: identifying legal clauses, regulations names, court names, judgment results, etc.; when the original document is a medical document, the specific entities identified include: identifying drug names, disease names, medical terms; when the original document is a financial document, the specific entities identified include: identifying company names, stock codes, financial events, etc.
[0025] S102: Collect domain terms corresponding to the document field, construct a corresponding domain knowledge graph, and map the entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation.
[0026] In a specified domain, knowledge is often scattered in various documents, databases, or different application systems, and in various forms. Therefore, in order to integrate fragmented knowledge and at the same time solve the problems of insufficient data or insufficient data features, this specification will collect domain terms corresponding to the document field, construct a corresponding domain knowledge graph to structurally integrate the originally scattered and disordered knowledge in this field, and map the entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation. Mapping the entities and relationships in the knowledge graph to a low-dimensional vector space realizes data dimensionality reduction, and then augmenting the knowledge graph in the low-dimensional vector space can generate more feature combinations, enabling the subsequent model learning to learn a more comprehensive knowledge pattern and improving the generalization ability of the model, so that it can better process and predict when facing new and unseen data.
[0027] Specifically, in one or more embodiments of this specification, collecting domain terms corresponding to the document field and constructing a corresponding domain knowledge graph specifically include the following steps: Based on the general knowledge graph corresponding to the document domain, the general causal relationship between the domain boundary and entities corresponding to the domain knowledge graph is determined. It can be understood that the general knowledge graph is a knowledge set with general applicability in this field, which contains a large number of entities, entity attributes and various relationships between entities. This knowledge is collected, organized and structured, and has certain generality and comprehensiveness. In addition, when determining the domain boundary, the document domain is used as a guide to screen out knowledge content related to the field from the general knowledge graph. For example, in the field of medical documents, knowledge related to human physiology, diseases, treatment methods, etc. is selected from the general knowledge graph. By analyzing the selected knowledge, the core concepts, themes and key entities contained in the field are clarified. For example, in the medical field, core concepts may include disease names, symptoms, drugs, etc.; key entities may include various disease entities, medical research institution entities, etc. Through these core concepts and entities, the scope boundary of the field is determined, that is, which knowledge belongs to the field and which does not.
[0028] In addition, for the general causal relationship between entities, since there is a large amount of relationship information between entities in the general knowledge graph. For the entities involved in the document field, analyze the associations between them and look for possible causal clues. For example, in the medical knowledge graph, there may be causal clues between the "smoking" entity and the "lung cancer" entity, because in general knowledge, smoking is considered to be an important factor causing lung cancer. In this process, determining the domain boundaries and the general causal relationship between entities provides an important foundation and framework for building or improving the domain knowledge graph. The domain boundaries clarify the scope and focus of the knowledge graph, making the knowledge graph more focused and accurate; the general causal relationship between entities enriches the semantic information of the knowledge graph, which helps to better understand the knowledge logic and internal connections within the domain.
[0029] Then, collect text data corresponding to the document domain, and extract fixed pattern terms corresponding to the document domain in the text data according to regular expressions. For example, the legal clause "Article X, Paragraph Y" in the legal field belongs to fixed pattern terms. Then, based on the frequency of co-occurrence of each term in the text data, construct a co-occurrence matrix, and extract the domain terms corresponding to the text data according to the co-occurrence matrix. Determine the knowledge graph type information corresponding to the document domain according to the domain boundary. Among them, the knowledge graph type information includes: entity type, relationship type, and attribute type. For example, when constructing a knowledge graph in the financial field, it is clear that "company" and "stock" are entity types, "holding" and "issuance" are relationship types, and "company market value" and "stock price" are attribute types, so that the knowledge graph can accurately and orderly represent the knowledge in the financial field. By extracting the knowledge corresponding to the knowledge graph type information in the fixed pattern terms and domain terms, the knowledge and the general knowledge graph are integrated according to the relevant parts of the knowledge and the general knowledge graph to obtain the corresponding domain knowledge graph.
[0030] In the above process, the general knowledge graph is used to determine the domain boundaries of the domain knowledge graph and the general causal relationship between entities, which provides direction for the entire construction process. This can accurately locate the scope of domain knowledge, avoid interference from irrelevant information, and ensure that the constructed domain knowledge graph focuses on the core content of a specific field. Fixed-pattern terms can be extracted based on regular expressions, which can quickly filter out terms that meet the characteristics of the domain from a large amount of text data. Combined with the co-occurrence matrix to extract domain terms, it is possible to efficiently and accurately obtain domain terms corresponding to the document domain from a large amount of text data, ensuring the efficiency and accuracy of term extraction during the construction of the domain knowledge graph. The extracted knowledge is integrated with the general knowledge graph, which not only makes full use of the rich resources and prior knowledge of the general knowledge graph, but also combines the professional knowledge of a specific field, so that the constructed domain knowledge graph has both broad versatility and can accurately reflect the uniqueness of the field.
[0031] Specifically, in one or more embodiments of the present specification, mapping entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation specifically includes: The domain knowledge graph is composed of entities, the relationships between entities, and the graph structure they constitute. Different embedding models have different characteristics and applicable scenarios, and are suitable for knowledge graphs with different structures and relationship types. For example, the TransE model is simple and intuitive, and is suitable for processing knowledge graphs with simple relationships; the DistMult model performs better when processing symmetric relationships; and the ComplEx model can handle more complex relationships. Therefore, in the embodiments of this specification, an embedding model that matches the domain knowledge graph will be determined based on the graph structure and relationship type of the domain knowledge graph. Then, based on the matching embedding model, the entities and relationships of the domain knowledge graph are mapped to a low-dimensional vector space. In the low-dimensional vector space, each entity and relationship is represented by a vector. These vectors contain the semantic information of the entities and relationships, and realize the conversion of complex structures and relationships in the knowledge graph into operations between vectors. Then, after obtaining the entity vectors of each entity in the domain knowledge graph in the low-dimensional vector space, the entity vectors are amplified using the operation characteristics of the entity vectors, such as vector addition, multiplication, dot product, etc., to obtain the newly added entity vectors and determine the relationship corresponding to the newly added entity vectors based on the operation characteristics corresponding to the newly added entity vectors. According to the newly added entities and the relationships corresponding to the newly added entity vectors, the original domain knowledge graph is amplified, and the new entities and relationships are added to the knowledge graph to form an amplified domain knowledge graph. In this way, the scale and content of the knowledge graph are expanded, containing more semantic information and knowledge, and can better reflect the various entities and relationships in the domain.
[0032] In this process, an embedding model is used to map the entities and relationships of the domain knowledge graph to a low-dimensional vector space, which can reduce the data dimension while retaining the structure and semantic information of the knowledge graph. Low-dimensional vector representation makes calculations more efficient, can quickly process large-scale knowledge graph data, and reduce storage and computing costs. By amplifying entity vectors based on the operational characteristics of entity vectors, potential relationships that were not explicitly expressed in the domain knowledge graph can be discovered. The determination of new entity vectors and their corresponding relationships enriches the semantic information of the knowledge graph, helps to reveal hidden connections between entities, and improves the integrity and accuracy of the knowledge graph. The domain knowledge graph is amplified according to the relationships corresponding to the new entities and new entity vectors, realizing the automatic update and expansion of the knowledge graph.
[0033] S103: Input a preset large model according to the amplified domain knowledge graph, and perform domain fine-tuning on the preset large model based on a strategy combining transfer learning and active learning to obtain a domain adaptation model.
[0034] Different fields have their own unique terms, concepts, semantic relationships, data distribution and other characteristics. In order to enable the model to adjust its own parameters according to the domain knowledge, better adapt to the characteristics of a specific field, and improve the accuracy and performance of the model on tasks in this field. In the embodiments of this specification, the preset large model will be input according to the amplified domain knowledge graph, so as to fine-tune the preset large model in the field according to the strategy of combining transfer learning with active learning to obtain a domain adaptation model.
[0035] Specifically, in one or more embodiments of the present specification, the preset large model is input according to the augmented domain knowledge graph, and the preset large model is fine-tuned in the domain based on a strategy combining transfer learning and active learning to obtain a domain adaptation model, which specifically includes the following processes: First, we collect existing knowledge graphs in different fields, which contain rich knowledge in their respective fields. Then, we map these existing knowledge graphs together with the augmented target domain knowledge graph to a low-dimensional vector space. In the low-dimensional vector space, each entity and relationship is represented by a vector. By calculating the similarity between each vector, we can find entities with similar semantics and entities with similar relationships between different fields. For example, the "disease" entity in the medical knowledge graph and the "pathological state" entity in the biological knowledge graph may be semantically similar entities. Then, based on the preset contrastive learning objective function, we conduct contrastive learning on the semantically similar entities and relationally similar entities found. The purpose of contrastive learning is to allow the model to learn the differences and commonalities between these similar entities and relationships, so as to obtain the knowledge data associated with the augmented domain knowledge graph in each existing knowledge graph. These knowledge data are represented in the form of vectors, and then migrated to the augmented domain knowledge graph, enriching the content of the domain knowledge graph and obtaining the migrated domain knowledge graph. For example, the knowledge data on the relationship between gene regulation and disease occurrence in the biological knowledge graph is migrated to the medical knowledge graph, thereby improving the part of the medical knowledge graph on the disease mechanism.
[0036] Then, based on the input format of the preset large model, the migrated domain knowledge graph is parsed and converted to adapt the migrated domain knowledge graph to the input layer of the preset large model to ensure that the knowledge graph can be passed to the large model as a valid input for processing. The converted domain knowledge graph is predicted using the existing model parameters of the preset large model. During the prediction process, the uncertainty of the prediction result is measured by calculating the entropy value of the prediction probability. The larger the entropy value, the higher the uncertainty of the prediction result. At the same time, the samples with larger entropy values in the prediction results are determined as uncertain samples and used as candidate samples for active learning. The preset large model is trained based on the determined candidate samples, and the existing model parameters of the preset large model are iteratively updated through the training process. With the continuous use of candidate samples for training, the model gradually learns more about the characteristics and laws of the domain knowledge graph, thereby improving the performance of the model in this field, and finally obtaining a domain adaptation model that meets the requirements, which can better handle related tasks in the target field.
[0037] In this process, by mapping the existing knowledge graphs of different fields and the augmented domain knowledge graphs to low-dimensional vector space, and determining semantically similar entities and relationally similar entities, the knowledge of multiple fields can be effectively integrated, the content of the target domain knowledge graph can be enriched, and more comprehensive information can be provided to the model, which helps the model learn more general and rich features and improves the performance and generalization ability of the model in the target domain. Based on the preset comparative learning objective function, the similar entities are compared and learned, and the knowledge data associated with the augmented domain knowledge graph in each existing knowledge graph can be obtained in a targeted manner. This method can automatically discover the potential connections between knowledge in different fields, accurately extract valuable knowledge for the target field, avoid the noise and interference caused by blindly integrating knowledge, and improve the accuracy and effectiveness of knowledge transfer. By calculating the entropy value of the prediction probability to determine the candidate samples for active learning, it is possible to focus on the uncertain samples in the model prediction results, which usually contain information that is difficult for the model to understand or classify. Only manually annotating these key samples can greatly reduce the annotation workload and improve the annotation efficiency compared to annotating a large number of random samples. At the same time, it can also more effectively utilize manual annotation resources and improve the effect of model training.
[0038] In a certain application scenario, the domain adaptation model can also be obtained based on the following process: first, collect domain terms, crawl professional literature and databases to obtain relevant data, then use TF-IDF, Word2Vec and other methods to automatically extract high-frequency terms, and combine manual review to optimize the term list to obtain domain terms. Then construct the domain knowledge graph, that is, use rule-based or machine learning methods to construct entities and their relationships. For example, in the medical field, [disease-symptom], [drug-indication] and other relationships can be constructed. Then perform data enhancement. In this process, domain synonyms can be used for replacement to achieve data replacement, and new sentences can be generated based on templates to achieve data synthesis, increase interference such as spelling errors and synonym replacement, and improve model robustness to achieve noise injection. Then the domain fine-tuning process is: manually annotate a small amount of high-quality domain data to achieve data annotation; use large models such as BERT and GPT to use transfer learning technology, fine-tune on the basis of general pre-training models, use Adam optimizer, and learning rate adjustment strategy; then evaluate the model based on F1-score, precision, recall evaluation and K-fold cross-validation optimization to obtain a domain adaptation model.
[0039] S104: Inputting the documents to be extracted in the specified field into the field adaptation model to extract knowledge elements of the documents to be extracted, and screening important knowledge elements based on the attention scores corresponding to the knowledge elements for visual display.
[0040] After obtaining the domain adaptation model based on the above process, the knowledge elements of the documents to be extracted in the specified domain can be extracted by inputting the documents to be extracted into the domain adaptation model, and important knowledge elements can be screened based on the attention scores corresponding to each knowledge element for visual display.
[0041] Specifically, in one or more embodiments of the present specification, the document to be extracted in the specified field is input into the domain adaptation model to extract the knowledge element of the document to be extracted, which specifically includes: preprocessing the document to be extracted in the specified field to determine the feature words of the document to be extracted based on the word frequency inverse document frequency value of each word segment in the document to be extracted. Then, according to the word frequency inverse document frequency value corresponding to the feature word, a feature vector is constructed to determine the document field of the document to be extracted through the matching degree of the feature vector with the preset domain term list. Determine the domain adaptation model corresponding to the document field of the document to be extracted, and convert the format of the document to be extracted based on the input requirements of the domain adaptation model to input the converted document to be extracted into the domain adaptation model. Extract knowledge elements from the converted document to be extracted based on the corresponding module of the domain adaptation model to obtain the entity, entity relationship and category label of each entity of the document to be extracted. Integrate the entity, entity relationship and category label of each entity of the document to be extracted to obtain the knowledge representation of the document to be extracted. Then, output the knowledge representation based on the preset standardized format to obtain the knowledge element of the document to be extracted.
[0042] In this process, the feature words of the document to be extracted are determined by the word frequency inverse document frequency value, which can effectively filter out common but meaningless words, highlight the important and distinguishing words in the document, so that the extracted feature words can better represent the core content of the document, and provide a solid foundation for subsequent domain judgment and knowledge extraction. For example, in academic documents, some commonly used function words will be filtered, while key content such as professional terms will be retained as feature words. The feature vector is constructed based on the TF-IDF value of the feature word, and the document domain is determined by matching it with the preset domain term list, which can accurately classify the document into the appropriate domain. In addition, the corresponding domain adaptation model is determined according to the document domain, and the format is converted according to the model input requirements, ensuring that the document to be extracted can be well adapted to the model. This enables the model to give full play to its performance and extract knowledge elements more effectively, avoiding the problem of poor model effect caused by format incompatibility or domain mismatch, and improving the efficiency and accuracy of knowledge extraction. The extracted knowledge elements are integrated and output based on the preset standardized format, which improves the standardization and universality of knowledge.
[0043] In a certain application scenario, knowledge elements are extracted by using a fine-tuned large model based on semantic understanding and context analysis, that is, domain classification is first performed: TF-IDF+Naive Bayes / BERT classifier is used to identify the document domain; the document domain score is calculated based on keyword matching in the domain terminology table.
[0044] Then perform multi-label classification: perform entity classification based on pre-trained classification models (such as BERT-MLC); use few-shot learning to improve the ability to classify rare categories. Perform sequence annotation: use BILSTM-CRF or Transformers models for word-by-word annotation; combine domain terminology and pre-trained NER models to improve accuracy. Perform relationship extraction: use syntactic analysis such as dependency syntax to extract entity relationships when based on rules. Use the twin tower model to calculate entity similarity when based on machine learning; use the attention mechanism to obtain contextual dependencies. Perform nested entity processing: use a hierarchical decoding method to first identify large entities and then refine sub-entities; use SpanBERT to improve cross-sentence entity recognition capabilities. Perform complex relationship processing: use graph convolutional networks to model multi-entity relationships; combine rule post-processing to ensure logical consistency.
[0045] Specifically, in one or more embodiments of the present specification, important knowledge elements are selected based on the attention scores corresponding to each knowledge element, and are visualized, specifically including: Based on the attention mechanism module of the domain adaptation model, the attention score corresponding to each knowledge element is obtained, so as to determine the knowledge element threshold corresponding to the document to be extracted according to the attention score distribution of each knowledge element. The attention score reflects the degree of attention paid by the model to the knowledge element during the extraction process. The higher the score, the more important the model believes this knowledge element is in the document. Therefore, the attention score of each knowledge element is compared with the determined threshold, and the knowledge element greater than or equal to the threshold is identified as an important knowledge element. For the entities and entity relationships involved in the important knowledge element, their confidence is further obtained. Then, the preset display tool is used to present the determined important knowledge elements and their corresponding confidence in an intuitive way.
[0046] In a certain application scenario, the process of improving the interpretability of the model through attention mechanism and visualization technology can also be achieved in the following ways: Attention weight visualization: Use the Attention Score in the Transformer model to visualize important text fragments; calculate the text contribution score to identify key texts. Confidence scoring: Use Softmax probability to calculate entity and relationship confidence; Use Bayesian uncertainty estimation to improve credibility assessment. Visualization tool integration: Use TensorBoard to monitor the model Attention Map; Develop a custom Web tool to display knowledge elements and their relationship diagrams.
[0047] like Figure 2 As shown, the embodiment of this specification provides a structural schematic diagram of a document knowledge element extraction device based on a large model. Figure 2 It can be seen that in one or more embodiments of this specification, a document knowledge element extraction device based on a large model includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: execute any of the above-mentioned methods.
[0048] like Figure 3 As shown, the present specification provides a schematic diagram of the structure of a non-volatile storage medium. Figure 3 It can be seen that in one or more embodiments of the present specification, a non-volatile storage medium stores computer executable instructions 301, and the computer executable instructions can: execute any of the above-mentioned methods.
[0049] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0050] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0051] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.
Claims
1. A document knowledge element extraction method based on a large model, characterized in that: The method comprises: Acquire original documents in a specified field based on a multi-source interface, perform part-of-speech tagging and named entity recognition processing according to the document content and document field of each original document, so as to obtain a standardized document; wherein the document field is included in the scope of the specified field; Collect domain terms corresponding to the document domain, construct a knowledge graph of the corresponding domain, and map entities and relationships in the knowledge graph of the corresponding domain to a low-dimensional vector space for data augmentation; Input a preset large model according to the amplified domain knowledge graph, and fine-tune the preset large model in the domain by combining transfer learning with active learning to obtain a domain adaptation model; The documents to be extracted in the specified field are input into the field adaptation model to extract the knowledge elements of the documents to be extracted, and important knowledge elements are screened based on the attention scores corresponding to the knowledge elements for visual display.
2. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Obtain original documents in the specified field based on multi-source interfaces, including: Acquire multimodal data of a specified field based on a multi-source interface, and associate and merge the multimodal data based on a key identifier corresponding to the specified field to obtain a multimodal data set corresponding to the same key identifier; Based on the data format corresponding to each modal data in the multimodal data set, calling a corresponding format conversion tool to convert the data format corresponding to each modal data into a corresponding standard format; Based on the amount of data in each corresponding standard format in each modality data, match the corresponding data storage structure and storage method; The modal data in the standard format is used to construct the original document of the specified field based on the data storage structure and the storage method.
3. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Part-of-speech tagging and named entity recognition are performed based on the document content and document domain of each original document to obtain standardized documents, including: Acquire specific word segments corresponding to the document field of the original document, so as to construct a custom dictionary corresponding to the original document based on the specific word segments; Performing data cleaning on each original document based on a regular expression to obtain a cleaned original document, and performing a character query on the cleaned original document to determine whether the cleaned original document contains full-width characters or an erroneous encoding format; Convert and correct the full-width characters and the erroneous encoding format to obtain the original document to be segmented; Performing word segmentation processing on the original document to be segmented by using a preset word segmentation tool and the custom dictionary to obtain a plurality of word segmentation data; Perform part-of-speech tagging on the word segmentation data based on the pre-trained language model, and identify specific entities in the original document to be segmented for entity tagging based on the custom dictionary and the preset named entity recognition model; The original document to be segmented after part-of-speech tagging and named entity recognition is output based on a preset unified format to form a standardized document.
4. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Collect domain terms corresponding to the document domain and construct a corresponding domain knowledge graph, including: Based on the general knowledge graph corresponding to the document domain, determining the general causal relationship between the domain boundary and entities corresponding to the domain knowledge graph; Collecting text data corresponding to the document domain to extract fixed pattern terms corresponding to the document domain from the text data based on regular expressions; Based on the co-occurrence frequency of each term in the text data, a co-occurrence matrix is constructed to extract the field terms corresponding to the text data based on the co-occurrence matrix; Determine the knowledge graph type information corresponding to the document domain based on the domain boundary; wherein the knowledge graph type information includes: entity type, relationship type, and attribute type; Extract the knowledge corresponding to the knowledge graph type information from the fixed pattern terms and the domain terms, so as to achieve the fusion of the knowledge and the general knowledge graph based on the relevant parts of the knowledge and the general knowledge graph, and obtain the corresponding domain knowledge graph.
5. According to the large model-based document knowledge element extraction method of claim 1, it is characterized in that: Mapping entities and relationships in the corresponding domain knowledge graph to a low-dimensional vector space for data augmentation includes: Based on the graph structure and relationship type of the domain knowledge graph, determine an embedding model that matches the domain knowledge graph; wherein the embedding model includes: TransE, DistMult, ComplEx; Based on the matched embedding model, mapping entities and relationships of the domain knowledge graph to a low-dimensional vector space; Obtaining entity vectors of each entity in the domain knowledge graph in the low-dimensional vector space, amplifying the entity vectors based on the computational characteristics of the entity vectors, obtaining newly added entity vectors, and determining the relationship corresponding to the newly added entity vectors based on the computational characteristics corresponding to the newly added entity vectors; Based on the relationship between the newly added entity and the newly added entity vector, the corresponding domain knowledge graph is amplified to obtain an amplified domain knowledge graph.
6. The document knowledge element extraction method based on a large model according to claim 1 is characterized in that: According to the augmented domain knowledge graph, the preset large model is input, and the preset large model is fine-tuned in the domain by a strategy based on combining transfer learning with active learning to obtain a domain adaptation model, which specifically includes: Collect existing knowledge graphs in different fields to map the existing knowledge graphs and the amplified domain knowledge graphs into a low-dimensional vector space, and determine semantically similar entities and relationally similar entities in different fields based on the similarity between vectors in the low-dimensional vector space; Based on a preset comparative learning objective function, the semantically similar entities and relationally similar entities are comparatively learned to obtain knowledge data associated with the augmented domain knowledge graph in each existing knowledge graph, and the knowledge data is migrated to the augmented domain knowledge graph based on the vector representation of the knowledge data to obtain the migrated domain knowledge graph; Based on the input format of the preset large model, the migrated domain knowledge graph is parsed and converted to adapt the migrated domain knowledge graph to the input layer of the preset large model; Predicting the converted domain knowledge graph based on the existing model parameters of the preset large model, so as to determine the uncertain samples in the prediction results as candidate samples for active learning by calculating the entropy value of the prediction probability; The preset large model is trained based on the candidate samples to iteratively update the existing model parameters of the preset large model to obtain a domain adaptation model that meets the requirements.
7. The document knowledge element extraction method based on a large model according to claim 1 is characterized in that: Inputting the document to be extracted in the specified domain into the domain adaptation model to extract the knowledge element of the document to be extracted specifically includes: Preprocessing the documents to be extracted in the specified field to determine the characteristic words of the documents to be extracted based on the word frequency inverse document frequency value of each word segment in the documents to be extracted; According to the word frequency inverse document frequency value corresponding to the feature word, a feature vector is constructed to determine the document domain of the document to be extracted through the matching degree between the feature vector and a preset domain term list; Determine a domain adaptation model corresponding to the document domain of the document to be extracted, and convert the format of the document to be extracted based on the input requirements of the domain adaptation model, so as to input the converted document to be extracted into the domain adaptation model; Extracting knowledge elements from the converted document to be extracted based on the corresponding module of the domain adaptation model to obtain entities, entity relationships and category labels of each entity of the document to be extracted; Integrate the entities, entity relationships and category labels of the document to be extracted to obtain the knowledge representation of the document to be extracted; The knowledge representation is output based on a preset standardized format to obtain the knowledge element of the document to be extracted.
8. The document knowledge element extraction method based on a large model according to claim 1 is characterized in that: Based on the attention scores corresponding to each knowledge element, important knowledge elements are selected and visualized, including: Based on the attention mechanism module of the domain adaptation model, the attention score corresponding to each knowledge element is obtained, so as to determine the knowledge element threshold corresponding to the document to be extracted according to the attention score distribution of each knowledge element; The important knowledge element corresponding to the document to be extracted is determined by the knowledge element threshold to obtain the confidence of the entity and entity relationship in the important knowledge element, so as to visualize the important knowledge element and the confidence based on a preset display tool.
9. A document knowledge element extraction device based on a large model, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: execute any of the methods described in claims 1-8.
10. A non-volatile storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Knowledge acquisition and representation method for power grid risk field
CN116739081A
Unstructured text data knowledge extraction method based on large language model
CN118036734A
Enterprise-level knowledge base construction method based on large model
CN119622040A
Entity relation mining method based on biomedical literature
US20230007965A1
Medical plan recommendation system and method based on knowledge graph representation learning
WO2021189971A1
Cited By
Generative knowledge object extraction method and system based on context learning
CN120930750A
Document element rapid extraction system based on pre-training large model
CN121189458A
A Fast Document Feature Extraction System Based on Pre-trained Large Models
CN121189458B
Multi-modal detection knowledge graph construction method
CN121390250A