Knowledge graph construction method and device based on large model, computer equipment, storage medium and computer program product
Through the knowledge graph construction method based on large models, data is automatically processed and decoded, and combined with field databases and professional knowledge optimization construction process, the problem of low efficiency in knowledge graph construction in the existing technology is solved, and rapid and accurate knowledge graph construction and document information integration are achieved.
Patent Information
- Application Number
- CN202510021267.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology requires a lot of manual intervention when building a knowledge graph. The construction cycle is long and inefficient, so it is impossible to quickly build a knowledge graph to achieve rapid integration and utilization of document information.
The knowledge graph construction method based on the big model is adopted, and the original data in different formats in the same archive system that is pre-collected, and the data is encoded and decoded using the pre-trained graph to automatically extract entities, attributes and relationships, and a preliminary knowledge graph structure is constructed, and optimized and updated by combining the field database and knowledge data provided by professionals.
The integration and unification of original data in different formats is achieved, automatic information extraction reduces human errors, improves the accuracy and efficiency of information extraction, and the target knowledge graph is richer in structure and content, improves the coverage of domain knowledge, and significantly improves the efficiency of knowledge graph construction.
Smart Images

Figure CN119940501A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of knowledge graph construction, and in particular to a method, apparatus, computer equipment, storage medium and computer program product for constructing a knowledge graph based on a large model. Background Art
[0002] In the current enterprise management process, with the continuous development of enterprise informatization, a large amount of document information has accumulated in the enterprise archive system. Since these document information are scattered in different data sources and have different formats, they are difficult to be efficiently integrated and used. Therefore, how to efficiently organize the document information has become a problem to be solved.
[0003] However, the current way to organize document information is mainly through the construction of knowledge graphs, but the construction process requires a lot of manual intervention, and the construction cycle is long and inefficient, and it is impossible to quickly build a knowledge graph to achieve rapid integration and utilization of document information. Therefore, there is currently a problem of low efficiency in knowledge graph construction. Summary of the invention
[0004] Based on this, it is necessary to provide a knowledge graph construction method, device, computer equipment, computer-readable storage medium and computer program product based on a large model to address the technical problem of low efficiency in constructing the above-mentioned knowledge graph.
[0005] In a first aspect, the present application provides a method for constructing a knowledge graph based on a large model, comprising:
[0006] Process the original data in different formats collected in the same archive system to obtain the target data in the same format;
[0007] Determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data;
[0008] Based on the entities, the attributes and the relationships, construct a preliminary knowledge graph structure;
[0009] The preliminary knowledge graph structure is converted into a target knowledge graph by constructing a large model through the pre-trained graph and combining it with a database related to the field of the archive system and / or knowledge data provided by professionals; the target knowledge graph is continuously updated based on the original data updated in real time.
[0010] In one of the embodiments, the determining the construction requirements of the knowledge graph to be constructed and using the pre-trained graph to construct a large model to encode and decode the target data includes: determining the construction requirements of the knowledge graph to be constructed, and determining the field and subject scope of the knowledge graph to be constructed based on the construction requirements; using the pre-trained graph to construct a large model, combining the field and the subject scope, and a database related to the field where the archive system is located and / or knowledge data provided by professionals to encode and decode the target data.
[0011] In one of the embodiments, the knowledge data provided by the professionals include the construction rules of the knowledge graph to be constructed; the large-model-based knowledge graph construction method also includes: after the initial graph data is extracted from the pre-trained graph construction large model, the initial graph data is compared with the construction rules, and the initial graph data is corrected based on the comparison results to obtain the target graph data.
[0012] In one of the embodiments, after converting the preliminary knowledge graph structure into the target knowledge graph, it also includes: evaluating the entities, the attributes and the relationships contained in the target knowledge graph to obtain evaluation results; according to the evaluation results, adjusting the parameters of the pre-trained graph construction model, as well as the structure and construction process of the target knowledge graph.
[0013] In one of the embodiments, the large model-based knowledge graph construction method further includes: providing a feedback interface and receiving user feedback through the feedback interface; and updating the target knowledge graph based on the received user feedback.
[0014] In one embodiment, the processing of the raw data in different formats collected in the same archive system in advance to obtain target data in the same format includes: performing data cleaning, noise removal, and duplicate data removal on the raw data in different formats collected in the same archive system in advance to obtain preprocessed data; and performing data conversion on the preprocessed data to obtain target data in the same format.
[0015] In a second aspect, the present application also provides a knowledge graph construction device based on a large model, comprising:
[0016] The raw data processing module is used to process the raw data of different formats collected in the same archive system in advance to obtain the target data of the same format;
[0017] The target data processing module is used to determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data;
[0018] A knowledge graph construction module, used to construct a preliminary knowledge graph structure based on the entities, the attributes and the relationships;
[0019] The knowledge graph construction module is also used to construct a large model through the pre-trained graph, combine the database related to the field of the archive system and / or the knowledge data provided by professionals, and convert the preliminary knowledge graph structure into a target knowledge graph; the target knowledge graph is continuously updated based on the real-time updated original data.
[0020] In a third aspect, the present application further provides a computer device, the computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:
[0021] Process the original data in different formats collected in the same archive system to obtain the target data in the same format;
[0022] Determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data;
[0023] Based on the entities, the attributes and the relationships, construct a preliminary knowledge graph structure;
[0024] The preliminary knowledge graph structure is converted into a target knowledge graph by constructing a large model through the pre-trained graph and combining it with a database related to the field of the archive system and / or knowledge data provided by professionals; the target knowledge graph is continuously updated based on the original data updated in real time.
[0025] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0026] Process the original data in different formats collected in the same archive system to obtain the target data in the same format;
[0027] Determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data;
[0028] Based on the entities, the attributes and the relationships, construct a preliminary knowledge graph structure;
[0029] The preliminary knowledge graph structure is converted into a target knowledge graph by constructing a large model through the pre-trained graph and combining it with a database related to the field of the archive system and / or knowledge data provided by professionals; the target knowledge graph is continuously updated based on the original data updated in real time.
[0030] In a fifth aspect, the present application further provides a computer program product, the computer program product comprising a computer program, which implements the following steps when executed by a processor:
[0031] Process the original data in different formats collected in the same archive system to obtain the target data in the same format;
[0032] Determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data;
[0033] Based on the entities, the attributes and the relationships, construct a preliminary knowledge graph structure;
[0034] The preliminary knowledge graph structure is converted into a target knowledge graph by constructing a large model through the pre-trained graph and combining it with a database related to the field of the archive system and / or knowledge data provided by professionals; the target knowledge graph is continuously updated based on the original data updated in real time.
[0035] The above-mentioned knowledge graph construction method, device, computer equipment, storage medium and computer program product based on the big model have the following beneficial effects in the process of constructing the knowledge graph based on the big model: first, the original data of different formats collected in the same archive system in advance are processed to obtain target data with the same format; then the construction requirements of the knowledge graph to be constructed are determined, and the target data is encoded and decoded using the pre-trained graph to build a big model to obtain the entities, attributes and relationships of the target data; then, based on the entities, attributes and relationships, a preliminary knowledge graph structure is constructed; finally, the preliminary knowledge graph structure is converted into a target knowledge graph through the pre-trained graph construction big model, combined with the database related to the field where the archive system is located and / or the knowledge data provided by professionals; the target knowledge graph is continuously updated based on the real-time updated original data. Through the above process, raw data of different formats are processed to obtain target data of the same format, which can achieve data integration and unification; using pre-trained graphs to build large models for encoding and decoding, automated information extraction can reduce human errors and improve the accuracy and efficiency of information extraction; by combining the database responsible for the field and the knowledge data provided by professionals, the target knowledge graph can be made richer in structure and content, and its coverage of domain knowledge can be improved. Therefore, the above process improves the efficiency of knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0037] Figure 1 It is an application environment diagram of a method for constructing a knowledge graph based on a large model in one embodiment;
[0038] Figure 2 A schematic diagram of a process of constructing a knowledge graph based on a large model in one embodiment;
[0039] Figure 3 A schematic diagram of a process for constructing a knowledge graph based on a large model in one embodiment;
[0040] Figure 4 A structural block diagram of a device for constructing a knowledge graph based on a large model in one embodiment;
[0041] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0042] With the continuous development of enterprise informatization, a large amount of various information has accumulated in enterprise archive systems. However, this information is often scattered in different data sources with different formats, making it difficult to quickly and effectively integrate and utilize. Traditional knowledge graph construction methods usually require a lot of manual intervention, have a long construction cycle, and are inefficient. In today's big data era, enterprises have an increasing demand for information management and knowledge services, and there is an urgent need for a method that can quickly build a knowledge graph to improve document association efficiency and achieve rapid integration and utilization of knowledge.
[0043] In recent years, with the rapid development of artificial intelligence technology, large models have achieved remarkable results in fields such as natural language processing. Large models have powerful language understanding and knowledge extraction capabilities, and can automatically analyze and process large amounts of text data. At the same time, automated pipeline technology has also been widely used in fields such as software development, which can achieve efficient task process automation. In the field of knowledge graph construction, related technologies are constantly developing, gradually moving from traditional rule-based and manual construction methods to automation and intelligence. However, there are still some problems, such as the accuracy of knowledge extraction and the consistency of knowledge fusion.
[0044] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0045] The knowledge graph construction method based on a large model provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The server 104 first processes the original data of different formats in the same archive system provided by the terminal 102 that is collected in advance to obtain the target data of the same format; then determines the construction requirements of the knowledge graph to be constructed, and uses the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data; then builds a preliminary knowledge graph structure based on the entities, attributes and relationships; finally, through the pre-trained graph to build a large model, combined with the database associated with the field of the archive system and / or the knowledge data provided by professionals, the preliminary knowledge graph structure is converted into a target knowledge graph; the target knowledge graph is continuously updated based on the real-time updated original data. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices, and the Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, etc. The portable wearable device may be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.
[0046] In an exemplary embodiment, Figure 2 As shown in the figure, a knowledge graph construction method based on a large model is provided, and this method is applied to Figure 1 The server 104 in the example is used as an example to illustrate the method, which includes the following steps S202 to S208. Among them:
[0047] Step S202, processing the original data in different formats collected in the same archive system in advance to obtain target data in the same format.
[0048] Among them, the same archive system refers to an enterprise archive system, and the original data refers to unprocessed data, which can include various types of information in an enterprise archive system, such as OA (Office Automation) documents, business data, financial statements, etc. The above data is accumulated in the daily operation of the enterprise and contains information about the company's employees, projects, documents, etc.
[0049] Optionally, the same format includes a process of format unification, which means that data from different sources and formats are converted into the same structure and type, and may include data cleaning, standardization and conversion operations.
[0050] Step S204, determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data.
[0051] Among them, the knowledge graph is a graphical knowledge representation method that expresses knowledge and its structure through nodes, namely entities, and edges, namely relationships. Through the knowledge graph, the relationship between entities can be more intuitively represented, providing deeper information retrieval and reasoning capabilities; construction requirements refer to clarifying the content, goals and application scenarios to be built before establishing the knowledge graph; it may include determining which entities need to be included, which relationships need to be depicted, and the purpose of using the knowledge graph, such as data analysis, knowledge sharing, etc.; the pre-trained graph construction model is a large-scale machine learning model that has been trained on a large amount of data, can process and convert various complex data formats, and can automatically extract patterns and features in the data.
[0052] As an example, encoding is converting data into a format that the model can understand, and decoding is converting the data output by the model back into a format that humans can understand; encoding can convert text data into vectors, and decoding can reflect the vectors into natural language text or visual content.
[0053] More generally, an entity refers to the main object or element represented by the target data in the knowledge graph, which can be a noun, such as an item, a place, an event, etc.; an attribute refers to the characteristics or descriptive information of an entity, which is frequently used to supplement and explain the properties of the entity; a relationship refers to the mutual connection or association between entities, usually connecting different nodes through edges.
[0054] As an example, an entity can be an information unit with a unique identifier, such as a customer, a product, or a company, which can exist independently in the knowledge graph; for example, the attributes of a "customer" entity may include "name", "contact information", "address" and other information, and the attributes can describe the entity more comprehensively; the types of relationships can be varied, such as "purchase", "located at", "contact", etc., which describe the functions and interactions between entities.
[0055] Step S206, construct a preliminary knowledge graph structure based on entities, attributes and relationships.
[0056] Among them, entities can have multiple types, such as employees, projects, documents, etc.; relationships can have multiple types, such as participation, ownership, association, etc.; attributes can have multiple types, such as employee name, project time, document content, etc.; the preliminary knowledge graph structure refers to the graphical information structure initially formed in the early stage of knowledge graph construction, which is composed of entities, attributes and relationships, and is used to provide a basic framework. More entities and relationships can be added later for further refinement and expansion.
[0057] Step S208, constructing a large model through pre-trained graphs, combining databases related to the field of the archive system and / or knowledge data provided by professionals, converting the preliminary knowledge graph structure into a target knowledge graph; the target knowledge graph is continuously updated based on the real-time updated original data.
[0058] Among them, the target knowledge graph refers to the knowledge graph that meets the specific application requirements after multiple rounds of processing and optimization; real-time update refers to the process in which the target knowledge graph can continuously access and analyze new data and dynamically maintain its content and structure after the initial construction.
[0059] In the above-mentioned knowledge graph construction method based on the big model, the original data of different formats collected in the same archive system are first processed to obtain the target data of the same format; then the construction requirements of the knowledge graph to be constructed are determined, and the pre-trained graph is used to construct the big model to encode and decode the target data to obtain the entities, attributes and relationships of the target data; then based on the entities, attributes and relationships, a preliminary knowledge graph structure is constructed; finally, the preliminary knowledge graph structure is converted into a target knowledge graph by constructing the big model through the pre-trained graph, combining the database associated with the field where the archive system is located and / or the knowledge data provided by professionals; the target knowledge graph is continuously updated based on the real-time updated original data. Through the above process, the original data of different formats are processed to obtain the target data of the same format, which can realize the integration and unification of data; the pre-trained graph is used to construct the big model for encoding and decoding, and the automated information extraction can reduce human errors and improve the accuracy and efficiency of information extraction; by combining the database responsible for the field and the knowledge data provided by professionals, the target knowledge graph can be made richer in structure and content, and its coverage of domain knowledge can be improved. Therefore, the above process improves the efficiency of knowledge graph construction.
[0060] In an exemplary embodiment, step S204, determining the construction requirements of the knowledge graph to be constructed, and using the pre-trained graph to build a large model to encode and decode the target data, includes: determining the construction requirements of the knowledge graph to be constructed, and determining the field and subject scope of the knowledge graph to be constructed based on the construction requirements; using the pre-trained graph to build a large model, combining the field and subject scope, and the database related to the field where the archive system is located and / or the knowledge data provided by professionals, to encode and decode the target data.
[0061] Among them, construction requirements refer to the purposes, goals and expected functions that need to be clarified when building a knowledge graph; domain refers to the specific knowledge or industry field involved in the knowledge graph, such as medical, financial, education, etc. Determining the domain helps focus on relevant data sources and information structures; subject scope refers to the specific subject or topic that the knowledge graph will focus on within the selected field, such as a specific disease or market trend; clarifying the subject scope can help limit the required data.
[0062] As an example, building requirements help ensure that the knowledge graph can meet specific user needs, for example, whether it is used for data retrieval, analytical decision-making, or knowledge sharing, including identifying the type of information, the required level of sophistication, and the user groups involved; databases can provide a structured data source for building knowledge graphs, and the input of professionals can help identify key entities and relationships, provide a deep understanding of specific fields, and thus improve the quality of knowledge graph construction.
[0063] In this embodiment, through the powerful language understanding ability of the big model (the big model constructed by pre-trained graph), various general entities in corporate archives can be more accurately identified, such as professional terms in specific business fields, product names, business process link names, etc., which greatly reduces the omissions and errors in entity recognition and improves the completeness and accuracy of entity information in the knowledge graph; for complex entity relationships, the big model can better understand the text semantics and accurately judge various relationship expressions such as "responsible", "participated", and "associated", so that the relationship construction in the knowledge graph is more in line with actual business logic, providing a more reliable foundation for subsequent knowledge applications.
[0064] Furthermore, in one embodiment, the knowledge data provided by the professionals includes construction rules of the knowledge graph to be constructed; the knowledge graph construction method based on the big model also includes:
[0065] After the pre-trained graph construction model extracts the initial graph data, the initial graph data is compared with the construction rules, and the initial graph data is corrected based on the comparison results to obtain the target graph data.
[0066] Among them, comparison refers to the systematic analysis of the initial graph data and the construction rules; correction refers to adjusting and optimizing the initial graph data according to the comparison results to solve the problems of inconsistency and missing; the target graph data refers to the corrected knowledge graph that conforms to the construction rules and can be used for practical applications.
[0067] As an example, comparing the initial graph data with the construction rules can help identify the gaps between the initial graph data and the preset specifications, such as inconsistent formats, non-compliance with naming rules, or missing required relationships, so that timely adjustments and corrections can be made.
[0068] In this embodiment, by comparing the initial graph data with the construction rules, problems and deficiencies in the data are identified and then corrected to ensure that the final target knowledge graph meets specific construction standards; and, according to the designed target knowledge graph, the fused knowledge is accurately filled into the corresponding positions, so that the structure of the target knowledge graph is more reasonable and the knowledge system is more complete. Both the attribute information of the entity and the relationship between the entities can be correctly organized and stored, providing reliable guarantees for the query and application of knowledge. At the same time, through continuous updating and maintenance, the dynamic consistency of the knowledge graph with the actual business of the enterprise is guaranteed.
[0069] In addition, by combining domain knowledge and rules to ensure rationality, the rules formulated by domain experts are combined with the output of the large model, adding a layer of professional verification mechanism for knowledge extraction and graph construction, ensuring that the extracted knowledge complies with industry standards and the actual business rules of the enterprise, thereby improving the rationality and compliance of knowledge.
[0070] In one embodiment, in step S208, after converting the preliminary knowledge graph structure into the target knowledge graph, the following steps are further included:
[0071] Evaluate the entities, attributes, and relationships contained in the target knowledge graph to obtain evaluation results; based on the evaluation results, adjust the parameters of the pre-trained graph to build a large model, as well as the structure and construction process of the target knowledge graph.
[0072] The evaluation process usually includes checking the accuracy, completeness, consistency and relevance of the data; the evaluation results refer to the feedback and summary obtained through the evaluation process, including the quality evaluation of entities, attributes and relationships; adjustment refers to modifying and optimizing the parameters of the pre-trained graph construction model and the structure and construction process of the knowledge graph based on the evaluation results.
[0073] As an example, adjustments can include updating the parameters of a pre-trained graph construction model to improve extraction accuracy, optimizing the structure of the knowledge graph to enhance data accessibility and understandability, or improving the construction process to improve efficiency and output quality.
[0074] In this embodiment, the evaluation process helps to identify possible problems, such as missing data, incorrect entity classification or relationship errors; and through the continuous improvement process, not only the quality of the target knowledge graph itself is improved, but also the effectiveness and reliability of the target knowledge graph in practical applications are ensured by fine-tuning the parameters of the pre-trained model and optimizing the construction process. Through this dynamic feedback system, the ability to build the knowledge graph is continuously improved.
[0075] Furthermore, in one embodiment, a method for constructing a knowledge graph based on a large model includes:
[0076] Provide a feedback interface and receive user feedback through the feedback interface; update the target knowledge graph based on the received user feedback.
[0077] Among them, the feedback interface refers to a system component that allows users to submit opinions, suggestions or questions in order to continuously improve the knowledge graph; the process of receiving user feedback includes collecting and storing feedback content submitted by users; updating refers to the process of correcting or expanding the content or structure of the target knowledge graph based on user feedback. Updates may include adding new entities, modifying existing attributes, adjusting relationships, and optimizing data structures.
[0078] In this embodiment, by setting up a feedback interface, users can participate in the construction process of the knowledge graph. Users can correct and confirm the extraction results, ensuring that the accuracy of the knowledge is in line with the user's actual business cognition; at the same time, users can also propose new entities and relationships according to their own needs, enriching the content of the knowledge graph and making it closer to actual business scenarios. This not only improves the quality of the knowledge graph, but also enhances users' understanding and trust in the knowledge graph, and promotes the widespread application and promotion of the knowledge graph within the enterprise.
[0079] In an exemplary embodiment, step S202, processing the original data in different formats collected in the same archive system to obtain target data in the same format includes:
[0080] The raw data in different formats collected in the same archive system are cleaned, noise is removed, and duplicate data is removed to obtain preprocessed data; the preprocessed data is converted to obtain target data in the same format.
[0081] Among them, data cleaning can include identifying errors, correcting missing values, and standardizing formats; noise removal refers to identifying and deleting random errors or irrelevant information in the original data; deduplication refers to identifying and deleting redundant entries in the original data to ensure the uniqueness of each data; preprocessed data is data that has undergone data cleaning, denoising, and deduplication, and is high-quality, structured, non-redundant, and meets the format requirements of subsequent operations.
[0082] In this embodiment, by performing data cleaning, noise removal and duplicate data processing on the original data, high-quality pre-processed data can be generated to achieve a unified target data format.
[0083] This application provides a method for constructing a knowledge graph based on a large model. Figure 3 As shown, the specific process of the knowledge graph construction method based on the large model of this application is described in detail below, including the following steps:
[0084] Step S302, obtaining various types of information in the enterprise archive system, which is original data.
[0085] Among them, the original data may include OA documents, business data, financial statements, etc. These data are accumulated during the daily operations of the enterprise and contain information about the company's employees, projects, documents, etc.
[0086] Step S304: the server collects and processes various types of original data in the enterprise archive system to obtain target data.
[0087] Among them, processing can include cleaning and preprocessing to remove noise and duplicate data, such as removing duplicate records through data deduplication algorithms, removing invalid characters and other noise data through regular expressions and other technologies, and then unifying the data format and converting data in different formats into a unified standard format as the target data.
[0088] Step S306, determine the domain and subject scope of the knowledge graph according to the needs of the enterprise, and build a preliminary knowledge graph structure.
[0089] Among them, for the business process topics or product information topics of the enterprise, etc. Design the architecture of the knowledge graph, define entity types such as employees, projects, documents, etc., relationship types such as participation, ownership, association, etc., and attributes such as employee name, project time, document content, etc. This process may require domain experts to provide expertise and experience to ensure that the design of the preliminary knowledge graph structure meets the actual business needs of the enterprise.
[0090] In addition, it also includes the knowledge extraction stage. The server uses a big model (pre-trained graphs to build a big model) to analyze and understand the pre-processed data. For text data, the big model uses natural language processing technology to identify entity mentions, such as names of people, places, and organizations, and relationship expressions, such as "working at", "responsible for", "associated with", etc. Combined with pattern matching, semantic analysis and other technologies and the output of the big model, entity, attribute and relationship information are extracted. For example, extract the name, position and other attributes of employees from OA documents, as well as the participation relationship between employees and projects, and then perform preliminary screening and verification on the extracted results to remove unreasonable or erroneous extraction results, such as by comparing with preset rules or common sense.
[0091] Furthermore, it also includes the knowledge fusion and integration stage, in which the server fuses the knowledge extracted from different data sources. Through context information, attribute comparison and other methods, possible problems such as entity homonyms and synonyms can be solved. For example, when multiple "Zhang San" are found, by analyzing the context information such as their department, projects involved and other attributes, it is determined whether they are the same person. Then, according to the design of the knowledge graph model, the fused knowledge is integrated, and the entities and relationships are filled into the corresponding positions in the knowledge graph to ensure the consistency and integrity of the knowledge and form a preliminary knowledge graph structure.
[0092] Step S308: The server uses the reasoning ability of the large model to perform reasoning based on the preliminary knowledge graph structure, and obtains the target knowledge graph through interaction with external knowledge bases or domain experts.
[0093] The server uses the reasoning ability of the large model to make inferences based on the preliminary knowledge graph structure. For example, based on the employee's position and project responsibilities, it can infer other related projects that the employee may be involved in. At the same time, through interaction with external knowledge bases or domain experts, more knowledge information can be obtained to supplement and improve the knowledge graph. For example, the industry standard knowledge base can be queried for relevant knowledge of a specific business process, or domain experts can be invited to provide professional insights to further enrich the content of the knowledge graph and obtain the target knowledge graph.
[0094] Furthermore, it also includes the knowledge graph evaluation and optimization stage, where the server evaluates the constructed knowledge graph and evaluates the accuracy and completeness of the entities, relationships, and attributes in the knowledge graph. For example, by statistically analyzing the frequency and distribution of entities and relationships, and comparing them with known standards or actual business situations. Based on the evaluation results, the problems and deficiencies in the knowledge graph construction process are analyzed, such as errors in knowledge extraction, inconsistencies in knowledge fusion, etc., and then the parameters of the large model, the structure of the target knowledge graph, and the construction process are optimized and adjusted, such as adjusting the hyperparameters of the large model, optimizing the algorithm for knowledge extraction, etc., to continuously improve the quality of the target knowledge graph.
[0095] Step S310, update the target knowledge graph in real time according to the continuous changes and additions of enterprise archive data.
[0096] Among them, the server applies the constructed target knowledge graph to actual business scenarios, such as intelligent search, intelligent recommendation, decision support, etc. In the intelligent search scenario, when the user enters the query keyword, the server quickly locates the relevant information based on the knowledge graph and returns it to the user; in the intelligent recommendation scenario, the server recommends relevant projects, documents, etc. to the user based on the user's historical behavior and the associated information in the knowledge graph; in the decision support scenario, by analyzing the data relationship in the target knowledge graph, the company's management provides a basis for decision-making. At the same time, the target knowledge graph is updated and maintained regularly. As the company's archival data continues to change and increase, the server extracts new knowledge in a timely manner and updates it to the target knowledge graph. For example, when a new project is established or employee information changes, the target knowledge graph is updated in a timely manner to ensure the timeliness and practicality of the target knowledge graph so that it can continue to create value for the company.
[0097] Through the above embodiments, a continuously optimized and updated enterprise archive subject knowledge graph, namely the target knowledge graph, is finally obtained. The graph contains accurate, complete and timely enterprise-related knowledge, which can provide strong support for various business applications of the enterprise, such as intelligent search results, intelligent recommended content, decision support suggestions, etc.; combined with large model-assisted extraction: utilizing the powerful language understanding and analysis capabilities of the large model, in-depth processing of the data in the enterprise archives can more accurately identify entity mentions and relationship expressions in text data, such as accurately identifying various generalized entities and complex entity relationships, overcoming the difficulties and inaccuracies of knowledge extraction in the prior art, and greatly improving the accuracy and efficiency of knowledge extraction. Through the automated pipeline technology, a complete set of automated processes from data preparation to knowledge graph application and maintenance has been designed, including data cleaning, knowledge extraction, fusion, reasoning, evaluation and optimization, etc., which realizes the automation of the whole process of knowledge graph construction, reduces manual intervention, significantly shortens the knowledge graph construction cycle, improves construction efficiency, and also solves the problem of imperfect end-to-end construction framework in existing technologies; it also includes knowledge fusion and reasoning optimization. In terms of knowledge fusion, it effectively solves the problems of entity homonymy and synonymy through various methods such as context information and attribute comparison, ensuring the consistency and Completeness; in the knowledge reasoning stage, the reasoning ability of the big model is used to discover potential relationships and knowledge, and interact with external knowledge bases or experts to supplement and improve them, further enriching the content of the knowledge graph, improving its quality and value, and making up for the shortcomings of existing technologies in knowledge integration and updating; in addition, by combining the rules formulated by domain experts with the output of the big model, the rationality and compliance of knowledge are improved, and at the same time, an interactive construction method is introduced to allow users to participate in the knowledge graph construction process. Users can correct erroneous knowledge and propose new entities and relationships, which not only improves the quality of the knowledge graph, but also enhances users' understanding and trust in the knowledge graph.
[0098] Further, in one embodiment, the algorithm principle of extracting raw data with the help of a large model is as follows: the large model is based on deep learning technology, and through a large amount of data training, it has learned rich language knowledge and semantic understanding ability, and can encode and decode the input text, identify the semantic information and grammatical structure therein, so as to accurately find entities and relationships. In addition, in the overall process design of the automated assembly line technology, the data preparation link is the starting point of the assembly line, responsible for collecting and preprocessing the raw data, and providing clean, unified format data for subsequent links; the knowledge graph (target knowledge graph) model design link determines the architecture and specifications of the graph according to the needs of the enterprise, and provides a framework for the organization and storage of knowledge; the knowledge extraction link extracts entity, attribute and relationship information from the data based on the large model and other technologies; the knowledge fusion and integration link integrates and removes the knowledge from different data sources to ensure the consistency of knowledge; the knowledge reasoning and supplement link uses large model reasoning and external interaction to enrich the content of the knowledge graph; the knowledge graph evaluation and optimization link evaluates the quality and optimizes the constructed graph; the knowledge graph application and maintenance link applies the knowledge graph to actual business, and continuously updates and maintains it to ensure its timeliness and practicality.
[0099] Moreover, the above process is inseparable from the synergy of various modules. The data processing relationship between the modules is as follows: the data preparation module provides high-quality data input for the knowledge extraction module; the output of the knowledge extraction module is used as the input of the knowledge fusion and integration module, and the preliminary knowledge graph structure is obtained after fusion and integration; the knowledge reasoning and supplement module performs reasoning and supplement based on the existing knowledge graph, and the results are fed back to the knowledge graph for updating; the knowledge graph evaluation and optimization module evaluates and optimizes according to the current status of the knowledge graph, and the optimized results affect the subsequent knowledge graph application and maintenance; the problems and new data discovered by the knowledge graph application and maintenance module during the application process prompt the aforementioned links to make corresponding adjustments and updates, forming a closed-loop process.
[0100] In addition, combining domain knowledge and rules, which are specific constraints and patterns summarized based on industry experience and professional knowledge, with the big model can increase the accuracy of specific domain knowledge and compliance judgment based on the general language understanding ability of the big model, and improve the accuracy and rationality of knowledge extraction and graph construction. Among them, the rules are formulated by domain experts based on the business characteristics and industry norms of the enterprise. For example, in the financial field, the rules such as "the repayment period must be within a reasonable range" for loan business; after the knowledge is extracted from the big model, the extracted results are compared and verified with these rules. If the extracted knowledge does not meet the rules, such as the repayment period exceeds the reasonable range, it will be regarded as an unreasonable or wrong result for screening and correction.
[0101] In one embodiment, the application of the target knowledge graph may include intelligent search scenarios and intelligent recommendation scenarios.
[0102] Among them, intelligent search scenarios include product side and technical side. For example, the product side is the front-end interface, including interface display content, that is, when the user enters keywords in the search box, the interface will display a list of related search results, including document titles, abstracts, related project information, employee information, etc. related to the keywords; at the same time, the relationship between related entities may be displayed in the form of a chart or knowledge graph, so that users can intuitively understand the relationship between information. For example, when a user searches for "Project A", the interface will display detailed information about Project A, such as project time, person in charge, participants, etc., as well as related information about other projects, documents, business processes, etc. related to Project A, in the form of nodes and lines of the knowledge graph. The search box is located at the top of the interface, and below it is the search result list area. Each result item in the list contains a title, abstract, and related icons, indicating the document type, project type, etc. There is a knowledge graph display area on the right or below. When a search result item is clicked, the relevant knowledge graph will be expanded in this area, with nodes of different colors representing different types of entities and lines representing relationships.
[0103] The technical side is responsible for background data processing, including execution subject and timing. When the user enters the search keyword in the front end, the search request is sent to the server. The server first quickly locates the entities and relationships related to the keyword based on the knowledge graph index, and then extracts relevant information from the knowledge graph, including the attributes of the entity and other entity information related to it. After sorting and sorting this information, it is returned to the front-end interface for display. For example: the user enters the search keyword in the front end and clicks the search button; the search request is sent to the server; the server searches for entities and relationships related to the keyword in the knowledge graph index; based on the entities and relationships found, the server obtains relevant information from the knowledge graph, such as entity attributes, associated entities, etc.; sorts and sorts the obtained information; returns the sorted results to the front-end interface; the front-end interface displays the search results.
[0104] More importantly, smart recommendation scenarios also include product and technology sides. For example, the product side is the front-end interface, including the interface display content: based on the user's historical operation behavior and current scenario, the interface will recommend relevant projects, documents, employees and other information. For example, when a user views a project page, the interface will recommend related similar projects, other documents that may be of interest, or expert employees in related fields. Recommended information is displayed in the form of a list or card, including a brief introduction to the recommended content and related links. In the sidebar or at the bottom of the project details page, there is a "Recommendation" area, which displays recommended content in the form of cards. Each card has the title, picture (if any), brief description and "View Details" button of the recommended project or document. When the user hovers over or clicks on the card, more relevant information will be displayed or a small window will pop up to display a detailed introduction.
[0105] The technical side is used for background data processing, including execution subjects and timing: the server monitors the user's operation behavior on the front end in real time, such as browsing projects, viewing documents, etc., and records these behavior data. When the user performs a specific operation or is on a specific page, the server uses the recommendation algorithm to calculate the relevant recommended content based on the user's behavior history and the associated information in the knowledge graph. For example, by analyzing the types of projects and related document topics that the user has browsed, combined with the similarity relationship between projects in the knowledge graph, the association relationship between documents, and the professional field relationship of employees, find out other content related to the user's interests, and then send the recommendation results to the front-end interface for display. For example, the user performs operations on the front end, such as browsing projects, viewing documents, etc.; the front end sends the user's operation behavior data to the server; the server records the user's behavior data and calculates it based on the knowledge graph and recommendation algorithm; the server obtains relevant recommended content information from the knowledge graph; the server organizes the recommendation results and sends them to the front-end interface; the front-end interface displays the recommended content.
[0106] Through the above embodiments, in the intelligent search scenario, based on the optimized knowledge graph, it is possible to quickly and accurately locate the relevant information of the user's query, and display the relationship between related entities in an intuitive way. Users can obtain the required documents, project information, etc. more quickly, which improves the efficiency and accuracy of information retrieval, saves the time cost of users to find information, and improves work efficiency; in the intelligent recommendation scenario, based on the user's behavior and the associated information of the knowledge graph, users are provided with more personalized and practical recommendation content, for example, recommending relevant project experience documents, potential partners, etc. to project managers, and recommending relevant training courses, business opportunities, etc. to employees, which helps to improve the utilization rate of enterprise resources and employee job satisfaction, and promote knowledge sharing and business collaboration within the enterprise.
[0107] Therefore, the knowledge graph construction method based on a large model mentioned in this application, in actual scenarios, when faced with limited corpus such as few annotated samples, the large model can still extract relatively accurate knowledge information from a small amount of corpus through its pre-trained knowledge and ability to learn language patterns. This solves the problem of low extraction accuracy of traditional methods with a small amount of corpus, reduces dependence on a large amount of annotated data, and improves the adaptability and flexibility of knowledge extraction; it realizes full-process automation and efficient construction, shortens the construction cycle, and the automated assembly line technology organically integrates all aspects of knowledge graph construction, from data preparation to knowledge graph application and maintenance, realizing full-process automation processing, reducing the time cost and uncertainty caused by manual intervention, and making the entire construction process more efficient and orderly. Compared with the traditional manual or semi-automatic construction method, it significantly shortens the construction cycle of the knowledge graph and can provide knowledge services to enterprises more quickly; it improves the end-to-end construction framework and provides a complete end-to-end solution, with close collaboration between modules and smooth data flow. From data collection and cleaning to knowledge extraction, fusion, reasoning, evaluation and optimization, and finally application and maintenance, each link has a clear execution subject and time sequence arrangement. This perfect framework makes the construction process of the knowledge graph more standardized and standardized, improves the success rate and quality stability of the construction; improves the quality of knowledge fusion and the integrity of the knowledge graph, effectively solves the problem of entity ambiguity, and accurately judges the homonymy and synonymy of entities in different data sources through context information, attribute comparison and other methods in the knowledge fusion stage. For example, it can accurately distinguish "Zhang San" with the same name in different departments, or merge information with different expressions but actually referring to the same entity. This ensures the uniqueness and accuracy of entities in the knowledge graph, avoids knowledge confusion caused by entity ambiguity, and improves the quality of knowledge fusion; enhances the ability to reason and supplement knowledge, discovers potential knowledge relationships, and uses the reasoning ability of large models to mine potential relationships and knowledge based on existing knowledge graphs. For example, based on the employee's position information and project responsibilities, other related projects that the employee may participate in can be inferred, providing valuable reference for the company's project planning and personnel deployment. This knowledge reasoning function expands the depth and breadth of the knowledge graph, making the knowledge graph not only a static information storage, but also an intelligent system that can dynamically generate new knowledge; enrich the content of the knowledge graph, and continuously obtain more knowledge information and supplement it to the knowledge graph through interaction with external knowledge bases or domain experts. This enables the knowledge graph to keep up with the development of the company's business and changes in the external environment in a timely manner, and continuously enrich and improve its own content. Whether it is the latest industry standards, market trends, or new business processes and project experience within the company, they can be reflected in the knowledge graph in a timely manner, improving the timeliness and practicality of the knowledge graph.
[0108] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0109] Based on the same inventive concept, the embodiment of the present application also provides a large-model-based knowledge graph construction device for implementing the large-model-based knowledge graph construction method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more large-model-based knowledge graph construction device embodiments provided below can refer to the limitations of the large-model-based knowledge graph construction method above, and will not be repeated here.
[0110] In an exemplary embodiment, Figure 4 As shown, a knowledge graph construction device based on a large model is provided, including: an original data processing module 401, a target data processing module 402 and a knowledge graph construction module 403, wherein:
[0111] The original data processing module 401 is used to process the original data of different formats collected in the same archive system in advance to obtain target data of the same format.
[0112] The target data processing module 402 is used to determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data.
[0113] The knowledge graph construction module 403 is used to construct a preliminary knowledge graph structure based on entities, attributes and relationships; the preliminary knowledge graph structure is converted into a target knowledge graph by building a large model through a pre-trained graph, combining a database related to the field of the archive system and / or knowledge data provided by professionals; the target knowledge graph is continuously updated based on the real-time updated original data.
[0114] Furthermore, in one embodiment, the target data processing module 402 is also used to determine the construction requirements of the knowledge graph to be constructed, and determine the field and subject scope of the knowledge graph to be constructed based on the construction requirements; use the pre-trained graph to build a large model, combine the field and subject scope, and the database related to the field where the archive system is located and / or the knowledge data provided by professionals to encode and decode the target data.
[0115] Furthermore, in one embodiment, the target data processing module 402 is also used to compare the initial atlas data with the construction rules after the pre-trained atlas construction model extracts the initial atlas data, and to correct the initial atlas data based on the comparison result to obtain the target atlas data.
[0116] Furthermore, in one embodiment, the knowledge graph construction module 403 is also used to evaluate the entities, attributes and relationships contained in the target knowledge graph to obtain evaluation results; based on the evaluation results, adjust the parameters of the pre-trained graph construction model, as well as the structure and construction process of the target knowledge graph.
[0117] Furthermore, in one embodiment, the knowledge graph construction module 403 is also used to provide a feedback interface and receive user feedback through the feedback interface; and update the target knowledge graph based on the received user feedback.
[0118] Furthermore, in one embodiment, the raw data processing module 401 is also used to perform data cleaning, noise removal, and duplicate data removal on raw data of different formats collected in advance in the same archive system to obtain preprocessed data; and perform data conversion on the preprocessed data to obtain target data of the same format.
[0119] Each module in the above-mentioned knowledge graph construction device based on a large model can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0120] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store knowledge graph construction data based on a large model. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a knowledge graph construction method based on a large model is implemented.
[0121] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0122] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0123] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0124] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0125] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0126] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0127] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0128] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for constructing a knowledge graph based on a large model, characterized in that: The method comprises: Process the original data in different formats collected in the same archive system to obtain the target data in the same format; Determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data; Based on the entities, the attributes and the relationships, construct a preliminary knowledge graph structure; The preliminary knowledge graph structure is converted into a target knowledge graph by constructing a large model through the pre-trained graph and combining it with a database related to the field of the archive system and / or knowledge data provided by professionals; the target knowledge graph is continuously updated based on the original data updated in real time.
2. The method according to claim 1, characterized in that The determining of the construction requirements of the knowledge graph to be constructed, and using the pre-trained graph to construct a large model to encode and decode the target data, includes: Determine the construction requirements of the knowledge graph to be constructed, and determine the field and subject scope of the knowledge graph to be constructed based on the construction requirements; The target data is encoded and decoded by building a large model using pre-trained graphs, combining the field and the subject scope, as well as databases and / or knowledge data provided by professionals in the field of the archive system.
3. The method according to claim 1 or 2, characterized in that: The knowledge data provided by the professionals include the construction rules of the knowledge graph to be constructed; The method further comprises: After the pre-trained atlas construction model extracts the initial atlas data, the initial atlas data is compared with the construction rules, and the initial atlas data is corrected based on the comparison result to obtain the target atlas data.
4. The method according to claim 1, characterized in that: After converting the preliminary knowledge graph structure into a target knowledge graph, the method further includes: Evaluate the entities, attributes, and relationships contained in the target knowledge graph to obtain evaluation results; According to the evaluation results, the parameters of the pre-trained graph construction model, as well as the structure and construction process of the target knowledge graph are adjusted.
5. The method according to claim 1, characterized in that The method further comprises: Providing a feedback interface and receiving user feedback through the feedback interface; Based on the received user feedback, the target knowledge graph is updated.
6. The method according to claim 1, characterized in that The processing of the original data in different formats collected in the same archive system to obtain target data in the same format includes: Clean the raw data in different formats collected in the same archive system, remove noise, and remove duplicate data to obtain pre-processed data; The preprocessed data is converted to obtain target data in the same format.
7. A knowledge graph construction device based on a large model, characterized in that: The device comprises: The raw data processing module is used to process the raw data of different formats collected in the same archive system in advance to obtain the target data of the same format; The target data processing module is used to determine the construction requirements of the knowledge graph to be constructed, and use the pre-trained graph to build a large model to encode and decode the target data to obtain the entities, attributes and relationships of the target data; A knowledge graph construction module, used to construct a preliminary knowledge graph structure based on the entities, the attributes and the relationships; The knowledge graph construction module is also used to construct a large model through the pre-trained graph, combine the database related to the field of the archive system and / or the knowledge data provided by professionals, and convert the preliminary knowledge graph structure into a target knowledge graph; the target knowledge graph is continuously updated based on the real-time updated original data.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Intelligent knowledge architecture graph construction method and system based on large model
CN120144790A
Data management method and device, equipment and medium
CN120723758A
Knowledge graph construction method and device, electronic equipment and storage medium
CN121279419A