A labeling method for security management knowledge graph data
By slicing and augmenting the target document data of the industrial and mining safety management knowledge graph, combined with user audit, the problem of low labeling efficiency is solved, the accuracy and scope of data are improved, and the needs of real-time monitoring and early warning are met.
Patent Information
- Application Number
- CN202411909449.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The construction method of industrial and mining safety management knowledge graph has the problem of low labeling efficiency, which makes it difficult to meet the needs of real-time monitoring and early warning.
By slicing the target document data of the security management knowledge graph to be constructed, and inputting the slice data into the preset model for knowledge augmentation processing, the augmented target candidate data is generated. Combined with user audit and optimization operations, the candidate data is adjusted, and the triples corresponding to the target entity are finally identified and generated.
It improves the labeling efficiency and accuracy of industrial and mining safety management knowledge graph data, expands the data range required for building knowledge graphs, and meets the needs of real-time monitoring and early warning.
Smart Images

Figure CN119357411B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for labeling security management knowledge graph data. Background Art
[0002] With the rapid development of the industrial and mining industries, the safety production situation is becoming increasingly severe. Traditional safety management methods often rely on manual experience and post-processing, which is difficult to meet the needs of real-time monitoring and early warning.
[0003] Since the current construction of industrial and mining safety management knowledge graphs generally relies on manual labeling, if the knowledge graph approach is used to meet the above-mentioned real-time monitoring and early warning needs, there is a problem of low labeling efficiency.
[0004] Therefore, the current way of constructing the knowledge graph of industrial and mining safety management has the problem of low labeling efficiency. Summary of the invention
[0005] The present application provides a method for labeling safety management knowledge graph data to solve the problem of low labeling efficiency in the current construction method of industrial and mining safety management knowledge graph.
[0006] The first aspect of the present application provides a method for labeling security management knowledge graph data, including:
[0007] Slice the target document data for constructing the safety management knowledge graph to generate corresponding multiple target slice data; the target document data is composed of knowledge data related to industrial and mining safety;
[0008] Inputting each of the target slice data into a preset large model, so as to use the preset large model to perform knowledge augmentation processing based on preset prompt words and the target slice data, and generate corresponding target candidate data after augmentation;
[0009] In response to the user's review operation and optimization operation on the target candidate data, the target candidate data is adjusted accordingly to generate corresponding target knowledge data after adjustment;
[0010] A preset large model is used to identify a target entity in the target knowledge data and to generate a triple corresponding to the target entity.
[0011] Furthermore, in the above method, the target document data to be constructed for the security management knowledge graph is sliced to generate corresponding multiple target slice data, including:
[0012] Parsing the target document data using a natural language processing method to identify paragraphs of the target document data;
[0013] Dividing the target document data into corresponding multiple target slice data according to the paragraphs;
[0014] or
[0015] The target document data is sliced according to preset custom slice parameters to generate corresponding multiple target slice data; the preset custom slice parameters include slice length, slice separator, and slice overlap ratio.
[0016] Furthermore, in the method as described above, the preset prompt words include question generation prompt words and question answer prompt words;
[0017] The step of inputting each of the target slice data into a preset large model, using the preset large model to perform knowledge augmentation processing based on preset prompt words and the target slice data, and generating corresponding target candidate data after augmentation includes:
[0018] Inputting each of the target slice data into a preset macro model, so as to generate corresponding knowledge questions based on the question generation prompt words and the target slice data using the preset macro model;
[0019] Using a preset large model to generate corresponding knowledge answers based on the question answer prompt words and the knowledge questions;
[0020] The knowledge answer and each of the target slice data are determined as the corresponding target candidate data after augmentation.
[0021] Furthermore, in the above method, the step of using a preset large model to generate a corresponding knowledge answer based on the question answer prompt word and the knowledge question includes:
[0022] Adopting a preset large model based on the retrieval enhancement generation method RAG to retrieve corresponding question related information from a preset database based on the knowledge question;
[0023] Generate a corresponding knowledge answer based on the question answer prompt word and the question related information.
[0024] Furthermore, in the above method, the step of performing corresponding adjustment processing on the target candidate data to generate corresponding target knowledge data after adjustment includes:
[0025] If the user's review operation is the operation corresponding to the approval, the approved target candidate data will not be modified;
[0026] If the user's review operation is an operation corresponding to a failure in the review, then based on the optimization operation, the target candidate data that fails in the review is modified accordingly to generate corresponding optimization data;
[0027] The approved target candidate data and the optimized data are determined as the target knowledge data.
[0028] Furthermore, the method as described above, before slicing the target document data to be constructed with the security management knowledge graph to generate a plurality of corresponding target slice data, further includes:
[0029] In response to a user input operation on the initial knowledge data, storing the initial knowledge data in a preset database;
[0030] Obtain the initial document data of the security management knowledge graph to be constructed from the preset database;
[0031] The initial document data is format converted to generate the target document data.
[0032] A second aspect of the present application provides a security management knowledge graph data annotation device, including:
[0033] A slicing module is used to slice the target document data for constructing the safety management knowledge graph to generate corresponding multiple target slice data; the target document data is composed of knowledge data related to industrial and mining safety;
[0034] A knowledge augmentation module, used for inputting each of the target slice data into a preset large model, so as to perform knowledge augmentation processing based on preset prompt words and the target slice data using the preset large model, and generate corresponding target candidate data after augmentation;
[0035] A data optimization module, configured to perform corresponding adjustment processing on the target candidate data in response to the user's review operation and optimization operation on the target candidate data, and generate corresponding target knowledge data after adjustment;
[0036] A generation module is used to use a preset large model to identify the target entity in the target knowledge data and generate a triple corresponding to the target entity.
[0037] Furthermore, in the above device, the slicing module is specifically used for:
[0038] The target document data is parsed and processed by using a natural language processing method to identify paragraphs of the target document data; the target document data is divided into a plurality of corresponding target slice data according to the paragraphs;
[0039] or
[0040] The target document data is sliced according to preset custom slice parameters to generate corresponding multiple target slice data; the preset custom slice parameters include slice length, slice separator, and slice overlap ratio.
[0041] Furthermore, in the above device, the preset prompt words include question generation prompt words and question answer prompt words;
[0042] The knowledge augmentation module is specifically used for:
[0043] Input each of the target slice data into a preset large model, and use the preset large model to generate corresponding knowledge questions based on the question-generated prompt words and the target slice data; use the preset large model to generate corresponding knowledge answers based on the question-answered prompt words and the knowledge questions; and determine the knowledge answers and each of the target slice data as the corresponding target candidate data after augmentation.
[0044] Furthermore, in the above-mentioned device, when the knowledge augmentation module generates corresponding knowledge answers based on the question answer prompt words and the knowledge questions using a preset large model, it is specifically used to:
[0045] The preset large model is adopted to retrieve the corresponding question-related information from the preset database based on the knowledge question based on the retrieval enhancement generation method RAG; and the corresponding knowledge answer is generated based on the question answer prompt word and the question-related information.
[0046] Furthermore, in the above-mentioned device, when the data optimization module performs corresponding adjustment processing on the target candidate data to generate the adjusted corresponding target knowledge data, it is specifically used to:
[0047] If the user's review operation is an operation corresponding to a passed review, the target candidate data that has passed the review will not be modified; if the user's review operation is an operation corresponding to a failed review, the target candidate data that has failed the review will be modified accordingly based on the optimization operation to generate corresponding optimization data; the target candidate data that has passed the review and the optimization data will be determined as the target knowledge data.
[0048] Furthermore, the device as described above further comprises:
[0049] A data acquisition module is used to store the initial knowledge data in a preset database in response to the user's input operation on the initial knowledge data; obtain the initial document data of the security management knowledge graph to be constructed from the preset database; and perform format conversion processing on the initial document data to generate the target document data.
[0050] A third aspect of the present application provides an electronic device, including: a memory and a processor;
[0051] The memory stores computer-executable instructions;
[0052] The processor executes the computer-executable instructions stored in the memory to implement the method for labeling security management knowledge graph data as described in any one of the first aspects.
[0053] The fourth aspect of the present application provides a computer-readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed by a processor, they are used to implement the method for labeling security management knowledge graph data described in any one of the first aspects.
[0054] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the method for labeling security management knowledge graph data as described in any one of the first aspects.
[0055] The present application provides a method for annotating safety management knowledge graph data, the method comprising: slicing target document data for constructing a safety management knowledge graph to generate corresponding multiple target slice data; the target document data is composed of knowledge data related to industrial and mining safety; each of the target slice data is input into a preset large model, so as to adopt the preset large model to perform knowledge augmentation processing based on preset prompt words and target slice data, and generate corresponding target candidate data after augmentation; in response to the user's review operation and optimization operation on the target candidate data, the target candidate data is adjusted accordingly to generate corresponding target knowledge data after adjustment; the preset large model is used to identify the target entity in the target knowledge data and generate a triple corresponding to the target entity. The method for annotating safety management knowledge graph data of the present application expands the data used to construct the safety management knowledge graph by inputting each target slice data into a preset large model, so as to adopt the preset large model to perform knowledge augmentation processing based on preset prompt words and target slice data, and generate corresponding target candidate data after augmentation. At the same time, in response to the user's review and optimization operations on the target candidate data, the target candidate data is adjusted accordingly to generate the corresponding target knowledge data after adjustment, so as to further optimize the target candidate data in combination with the user review, thereby improving the accuracy of the data used to construct the security management knowledge graph. Secondly, a preset large model is used to identify the target entity in the target knowledge data and generate a triple corresponding to the target entity, thereby improving the efficiency of data annotation by combining the large model with user review. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0057] Figure 1 A scene diagram of a method for labeling security management knowledge graph data that can implement an embodiment of the present application;
[0058] Figure 2Schematic diagram of the process of labeling the security management knowledge graph data provided for this application Figure 1 ;
[0059] Figure 3 Schematic diagram of the process of labeling the security management knowledge graph data provided for this application Figure 2 ;
[0060] Figure 4 A schematic diagram of the structure of the device for labeling the security management knowledge graph data provided by this application;
[0061] Figure 5 A schematic diagram of the structure of the electronic device provided in this application.
[0062] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0063] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0064] It should be noted that the method for labeling security management knowledge graph data disclosed in the present invention can be used in the field of artificial intelligence technology. It can also be used in any field other than the field of artificial intelligence technology. The application field of the method for labeling security management knowledge graph data disclosed in the present invention is not limited.
[0065] The technical solution of the present application is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0066] In order to clearly understand the technical solution of this application, the concept of this solution is first introduced in detail. With the rapid development of the industrial and mining industry, the safety production situation is becoming increasingly severe. Traditional safety management methods often rely on manual experience and post-processing, which is difficult to meet the needs of real-time monitoring and early warning.
[0067] At present, the construction of industrial and mining safety management knowledge graphs mainly relies on manual annotation, which has problems such as slow annotation speed, high cost, and easy errors. At the same time, the commonly used automatic annotation methods often cannot accurately understand the professional terms and contextual relationships in the field of safety management, resulting in poor annotation results.
[0068] Therefore, the current way of constructing the knowledge graph of industrial and mining safety management has the problem of low labeling efficiency and needs further optimization.
[0069] Therefore, in response to the problems in the prior art, the inventors discovered in their research that in order to solve the problem, a large model and manual review can be combined to utilize the higher accuracy of manual review and the higher efficiency of the large model to improve the labeling efficiency. At the same time, the standard effect can also be better.
[0070] Specifically, the target document data for constructing the security management knowledge graph is sliced to generate corresponding multiple target slice data. Each target slice data is input into a preset large model, so that the preset large model is used to perform knowledge augmentation processing based on preset prompt words and target slice data, and the corresponding target candidate data after augmentation is generated, thereby expanding the data used to construct the security management knowledge graph. At the same time, in response to the user's review and optimization operations on the target candidate data, the target candidate data is adjusted accordingly to generate the corresponding target knowledge data after adjustment, so as to further optimize the target candidate data in combination with the user review, thereby improving the accuracy of the data used to construct the security management knowledge graph. Secondly, the preset large model is used to identify the target entity in the target knowledge data and generate the triple corresponding to the target entity, thereby improving the efficiency of data annotation by combining the large model with user review.
[0071] Based on the above creative findings, the inventor proposed the technical solution of the present application.
[0072] The following is an introduction to the application scenarios of the method for labeling security management knowledge graph data provided by the embodiment of the present application. Figure 1 As shown, 1 is a first electronic device, 2 is a second electronic device, and 3 is a user. The first electronic device 1 can be a computer, a tablet computer, etc. The second electronic device 2 can be various database servers, knowledge platforms, etc., storing data required for building a safety management knowledge graph. The database server can be a database related to industrial and mining safety management. User 3 can be an expert in the field of industrial and mining safety.
[0073] For example, when it is necessary to construct a safety management knowledge graph, the first electronic device 1 obtains the target document data required for constructing the safety management knowledge graph from the second electronic device 2, wherein the target document data is composed of knowledge data related to industrial and mining safety. That is, ① the second electronic device 2 sends the target document data to the first electronic device 1, and at the same time, the user 3 participates in the data annotation process, that is, ② reviewing the data. At this time, the first electronic device 1 performs the following processing:
[0074] ③ Combined with the user's review, the target document data is annotated as follows:
[0075] The target document data for constructing the security management knowledge graph is sliced to generate corresponding multiple target slice data.
[0076] Each target slice data is input into a preset large model, so as to use the preset large model to perform knowledge augmentation processing based on preset prompt words and target slice data, and generate corresponding target candidate data after augmentation.
[0077] In response to the user's review operation and optimization operation on the target candidate data, the target candidate data is adjusted accordingly to generate corresponding target knowledge data after adjustment.
[0078] A preset large model is used to identify target entities in target knowledge data and generate triples corresponding to the target entities.
[0079] After the triples are generated, a safety management knowledge graph can be constructed based on the triples. At the same time, the safety management knowledge graph can be further audited and verified by expert review to improve the accuracy of the safety management knowledge graph.
[0080] The embodiments of the present application are introduced below in conjunction with the drawings in the specification.
[0081] Figure 2 Schematic diagram of the process of labeling the security management knowledge graph data provided for this application Figure 1 ,like Figure 2 As shown, in this embodiment, the execution subject of the embodiment of the present application is a labeling device for security management knowledge graph data, and the labeling device for security management knowledge graph data can be integrated into electronic devices, such as tablet computers, computers, etc. The labeling method for security management knowledge graph data provided in this embodiment includes the following steps:
[0082] S101, slicing the target document data for building the safety management knowledge graph to generate a plurality of corresponding target slice data. The target document data is composed of knowledge data related to industrial and mining safety.
[0083] In this embodiment, the slicing process may adopt a custom slicing method or a natural language processing method, which is not limited in this embodiment.
[0084] The purpose of slicing is to divide text data into smaller units (such as words or sentences) and convert these units into vector representations so that subsequent large models can understand and process the text data.
[0085] In this embodiment, the target document data is document data related to industrial and mining safety, and the target document data can be generated from industrial and mining safety-related data stored in a preset database after preprocessing, including data conversion, cleaning, format unification, and the like.
[0086] Optionally, the document data related to industrial and mining safety include: operating procedure data such as equipment operation, coal preparation process data, safety management system within the coal preparation plant, statistics on washing and processing accidents, safety measures of the coal preparation plant, etc.
[0087] S102, inputting each target slice data into a preset large model, so as to use the preset large model to perform knowledge augmentation processing based on preset prompt words and the target slice data, and generate corresponding target candidate data after augmentation.
[0088] In this embodiment, the preset big model can adopt a language big model, and the language big model can identify the target slice data through natural language processing (NLP) technology. At the same time, the target slice data is augmented with knowledge through RAG (Retrieval-Augmented Generation) to generate corresponding target candidate data after augmentation.
[0089] In this embodiment, by performing knowledge augmentation processing on the target slice data through a preset large model, the data used for the subsequent construction of the safety management knowledge graph can be expanded. For example, the target slice data includes relevant data on the safe use of equipment in factories and mines, among which the commonly used controls and switches of the equipment may not involve data on potential safety hazards or safety accidents in the past. For example, if the control gear of a certain device is adjusted to a larger gear in a certain environment, a small probability safety accident may occur. At this time, the preset large model can be used to perform knowledge augmentation processing based on the preset prompt words and the target slice data, such as adding corresponding data in the form of questions and answers, and enriching the data range used to construct the safety management knowledge graph.
[0090] Optionally, a large model may be used for augmentation processing, or synonyms or antonyms may be augmented to further enrich the data scope used to construct the security management knowledge graph.
[0091] Optionally, the preset prompt words can generate relevant prompt words for the question, prompt words related to the question answer, prompt words for synonym expansion, etc.
[0092] Optionally, other deep learning models such as convolutional neural network (CNN), recurrent neural network (RNN), etc. can be combined to perform knowledge augmentation by extracting features from target slice data to adapt to more data characteristics and application scenarios.
[0093] S103, in response to the user's review operation and optimization operation on the target candidate data, corresponding adjustment processing is performed on the target candidate data to generate corresponding target knowledge data after adjustment.
[0094] In this embodiment, the user may be an expert in the field of industrial and mining safety. When the user reviews the target candidate data, for example, if the target candidate data is determined to be correct, the review result may be determined to be approved. If the target candidate data is determined to need further modification, the review result may be determined to be unapproved and needs further modification. At the same time, the user may adjust the target candidate data through optimization operations, such as inputting modified data.
[0095] In this embodiment, the corresponding target knowledge data after adjustment has been reviewed and optimized by the user, thereby improving the accuracy of the target knowledge data and being more suitable for relevant scenarios in the field of industrial and mining safety.
[0096] S104, using a preset large model to identify a target entity in the target knowledge data and generate a triple corresponding to the target entity.
[0097] In this embodiment, a triple is a basic structure for representing knowledge. A triple consists of three parts: subject, predicate, and object. This structure is usually used to describe entities and their relationships, similar to a simple sentence in natural language. Generating a triple corresponding to the target data can provide a basis for subsequently determining the relationship between entities.
[0098] In this embodiment, natural language processing can be used to identify entities. At the same time, a basic framework schema can be created based on the target entities and the associations between the target entities. At the same time, a security management knowledge graph can be constructed based on the schema and the entities.
[0099] Optionally, a natural language processing (NLP) model can be used for entity recognition and relationship extraction to improve recognition accuracy and efficiency.
[0100] An embodiment of the present application provides a method for annotating safety management knowledge graph data, the method comprising: slicing target document data for constructing a safety management knowledge graph to generate corresponding multiple target slice data. The target document data is composed of knowledge data related to industrial and mining safety. Each target slice data is input into a preset large model, so that the preset large model is used to perform knowledge augmentation processing based on preset prompt words and target slice data to generate corresponding target candidate data after augmentation. In response to the user's review operation and optimization operation on the target candidate data, the target candidate data is adjusted accordingly to generate the corresponding target knowledge data after adjustment. The preset large model is used to identify the target entity in the target knowledge data and generate a triple corresponding to the target entity.
[0101] The method for annotating the security management knowledge graph data of the present application, by inputting each target slice data into a preset large model, so as to use the preset large model to perform knowledge augmentation processing based on preset prompt words and target slice data, and generate the corresponding target candidate data after augmentation, thereby expanding the data used to construct the security management knowledge graph. At the same time, in response to the user's review and optimization operations on the target candidate data, the target candidate data is adjusted accordingly to generate the corresponding target knowledge data after adjustment, so as to further optimize the target candidate data in combination with the user review, thereby improving the accuracy of the data used to construct the security management knowledge graph. Secondly, the preset large model is used to identify the target entity in the target knowledge data and generate the triple corresponding to the target entity, thereby improving the efficiency of data annotation by combining the large model with user review.
[0102] Figure 3 Schematic diagram of the process of labeling the security management knowledge graph data provided for this application Figure 2 ,like Figure 3 As shown, the method for labeling security management knowledge graph data provided in this embodiment is a further refinement of the method for labeling security management knowledge graph data provided in the previous embodiment of this application. The method for labeling security management knowledge graph data provided in this embodiment includes the following steps.
[0103] S201, slicing the target document data for constructing the security management knowledge graph to generate corresponding multiple target slice data.
[0104] Optionally, in this embodiment, S201 can be performed in two different ways, the first of which is as follows:
[0105] The target document data is parsed using natural language processing to identify paragraphs of the target document data.
[0106] The target document data is divided into a plurality of corresponding target slice data according to the paragraphs.
[0107] The second is as follows:
[0108] The target document data is sliced according to the preset custom slice parameters to generate corresponding multiple target slice data. The preset custom slice parameters include slice length, slice separator, and slice overlap ratio.
[0109] In this embodiment, the slice overlap ratio refers to the ratio of the overlapping portion between adjacent slices to the slice length when the data is divided into multiple slices (or windows). According to the needs of the actual application, it can be set to 5%, 10%, etc.
[0110] By parsing the document content through natural language processing, identifying the paragraphs of the document, and dividing the document into different slices, the efficiency of generating slices is higher, which is more convenient and more efficient than the second method. By slicing the target document data by presetting custom slice parameters and generating corresponding multiple target slice data, more custom content can be added, making the generated target slice data more accurate and more suitable for application scenarios in the field of industrial and mining safety.
[0111] Optionally, in this embodiment, a data preprocessing process is also included before S201, which is as follows:
[0112] In response to a user's input operation on the initial knowledge data, the initial knowledge data is stored in a preset database.
[0113] Obtain the initial document data for the security management knowledge graph to be constructed from the preset database.
[0114] The initial document data is formatted and processed to generate target document data.
[0115] In this embodiment, the human being may be an expert in the field of industrial and mining safety management, and inputs the collected knowledge in the field into a preset database, wherein the preset database stores the knowledge in the field of industrial and mining safety management.
[0116] The preset database may be a combination of an accident case database and an operating procedure database, or any one of the databases. The accident case database is used to store case data related to safety occurrences, and the operating procedure database is used to store operating procedures related to safety.
[0117] In this embodiment, since the large model may have certain requirements for input data, the initial document data may be format converted in advance to generate target document data, wherein the format of the target document data complies with the format requirements of the preset large model.
[0118] Optionally, the format of the initial document data may be txt format, Markdown (MD) format, doc format, docx format, ppt format, pptx format, pdf format, png format, jpg format, jpeg format, xls format, xlsx format, csv format, json format, etc. The format of the target document data may be Markdown text format.
[0119] It should be noted that the preset prompt words include question generation prompt words and question answer prompt words.
[0120] S202, input each target slice data into a preset large model, so as to use the preset large model to generate corresponding knowledge questions based on question generation prompt words and target slice data.
[0121] In this embodiment, the question generation prompt word is a prompt word used to generate a question, and the prompt word can be used according to actual applications.
[0122] Exemplarily, the prompt words for question generation are as follows: "#Role Expert in coal preparation plant safety questions. #Task You are good at knowledge questions and answers in the field of safety in coal preparation plants. Please generate 8 questions based on the content. Read and understand the text content related to the safety field in the original article. Extract specific questions based on the different aspects of safety issues in coal preparation plants. Prioritize questions based on the main issues described in the article. Organize and present a list of questions. In the content, if there is a closer problem description, please try to output relevant questions in accordance with the original text."
[0123] S203, using a preset large model to generate corresponding knowledge answers based on question answer prompt words and knowledge questions.
[0124] In this embodiment, the question answer prompt words and the question generation prompt words can also be set according to actual applications, which will not be described in detail here.
[0125] Optionally, in this embodiment, S203 may be specifically as follows:
[0126] The preset large model is used to retrieve the corresponding question-related information from the preset database based on the retrieval enhanced generation method RAG based on the knowledge question.
[0127] Generate corresponding knowledge answers based on question answer prompts and question-related information.
[0128] In this embodiment, RAG is a technology that combines information retrieval and generation models, and question-related information refers to information that is related to knowledge questions. By using RAG, corresponding question-related information can be retrieved from a preset database based on knowledge questions, thereby generating more accurate knowledge answers to improve the accuracy of the augmented target candidate data.
[0129] Optionally, when generating questions such as S202, RAG may be used in combination with question generation prompt words to generate corresponding knowledge questions to improve the accuracy of the knowledge questions.
[0130] S204, determining the knowledge answer and each target slice data as the corresponding target candidate data after augmentation.
[0131] In this embodiment, by combining a large model with knowledge questions and answers, more related knowledge can be added, such as knowledge data beyond synonyms, thereby further enriching the data scope used to construct the security management knowledge graph.
[0132] S205 , in response to the user's review operation and optimization operation on the target candidate data, corresponding adjustment processing is performed on the target candidate data to generate corresponding target knowledge data after adjustment.
[0133] Optionally, in this embodiment, S205 may be specifically as follows:
[0134] If the user's review operation is the operation corresponding to the approval, the target candidate data that has passed the approval will not be modified.
[0135] If the user's audit operation is an operation corresponding to a failure in the audit, then the target candidate data that fails the audit is modified accordingly based on the optimization operation to generate corresponding optimized data.
[0136] The target candidate data and optimized data that have passed the review are determined as target knowledge data.
[0137] In this embodiment, the audit operation may be a user performing a corresponding operation on an audit-related control in a preset data display interface, such as clicking an audit-approved control or a audit-failed control.
[0138] If you click on the control that does not pass the review, you can further enter the optimized data in the preset data input box so that the optimized data directly replaces the original target candidate data.
[0139] In addition, the data that failed the review can also be modified in the preset data display interface to directly generate corresponding optimization data.
[0140] The corresponding target knowledge data after adjustment has been reviewed and optimized by users, thereby improving the accuracy of the target knowledge data, ensuring the quality of knowledge, and being more suitable for relevant scenarios in the field of industrial and mining safety.
[0141] S206, using a preset large model to identify a target entity in the target knowledge data and generate a triple corresponding to the target entity.
[0142] In order to further illustrate the annotation method of the security management knowledge graph data of this application, the overall annotation process will be further explained below.
[0143] The details are as follows:
[0144] Complete knowledge collection, review, extraction and other tasks in one stop on the knowledge collection platform. The knowledge collection process is: upload basic information and convert it -> introduce knowledge base -> expand large model knowledge points -> review knowledge points -> create large model Q&A tasks -> review Q&A results -> create large model extraction tasks -> review extraction results -> end.
[0145] First, the collected industrial and mining safety-related data is preprocessed, including data conversion, cleaning, format unification, slicing, etc. Natural language processing technology is used to expand the knowledge points in the field of safety management knowledge of text data, and then the knowledge scope is continuously expanded through a combination of manual review and machine learning. Then, advanced natural language processing (NLP) technology is introduced to conduct knowledge Q&A. After the Q&A results are reviewed by experts in the safety field, large models are used for entity recognition, relationship extraction, and attribute labeling. Finally, the labeled data is stored in the knowledge graph database to provide support for subsequent query, reasoning, and analysis.
[0146] The system part of this embodiment mainly includes data preprocessing module, data slicing module, knowledge point augmentation module, knowledge question and answer module, entity recognition and relationship extraction module, etc. The modules communicate and collaborate through data interfaces to jointly complete the tasks of data annotation and graph construction.
[0147] Data preprocessing module: supports uploading various types of document data in the aforementioned embodiments, and also supports directly pasting text information. Convert the uploaded basic data, convert unrecognizable documents into Markdown text format recognizable by the program, and slice and annotate the documents after conversion to improve the accuracy of subsequent slicing. A large amount of basic data in the field of industrial and mining safety, including safety management systems in coal preparation plants, statistics on washing and processing accidents, and safety measures in coal preparation plants.
[0148] Data slicing module: supports automatic slicing or custom slicing methods. The main purpose is to convert queries into vectors to meet the processing requirements of large models. Automatic slicing uses natural language processing technology to parse document content, identify document paragraphs, and divide documents into different slices. Custom slicing is to set parameters such as slice length, slice delimiter, slice overlap ratio, etc. The slicing algorithm slices according to the parameters.
[0149] Knowledge point expansion module: Using natural language processing technology and prompt word engineering, the text data is expanded to expand the knowledge points in the field of security management knowledge and expand the coverage of security knowledge. Security experts review the expanded knowledge points and retain high-quality knowledge points.
[0150] Knowledge question answering module: Use RAG to retrieve information related to the question in the knowledge base, and use the retrieved information as context input for the generative model (i.e., large language model) to enhance the model's ability to understand and answer specific questions, and combine the language large model (LLM) to generate answers that meet user needs. Security experts score and evaluate the answers of the large model, and modify problematic answers to ensure the quality of knowledge.
[0151] Entity recognition and relationship extraction module: Through safety specifications and other knowledge elements, the entity type relationship is designed from top to bottom. For example, for the use scenario of the hazard source knowledge graph, triples such as "hazard source", "control method", and "risk control method of hazard source" are identified; for the use scenario of the safety inspection knowledge graph, triples such as "safety inspection items", "standards", and "specific standards for inspection" are identified; for the use scenario of the occupational health knowledge graph, triples such as "occupational disease", "measures", and "occupational disease protection measures" are identified. Entity recognition and relationship extraction are performed in the field of industrial and mining safety management knowledge for high-quality knowledge answers. After the extraction is completed, safety experts will score and evaluate the extraction results, and modify the problematic triples to ensure the quality of knowledge. At the same time, the triples that finally pass the review will be stored in the corresponding database.
[0152] The method for labeling the safety management knowledge graph data in this embodiment reduces a large amount of manual input workload through an automated knowledge point expansion mechanism, while expanding the coverage of knowledge. At the same time, combined with expert review, the quality of knowledge can be guaranteed, the accuracy of data required to construct the safety management knowledge graph can be improved, and the quality of the safety management knowledge graph constructed subsequently can be improved.
[0153] Figure 4 A schematic diagram of the structure of the annotation device for the security management knowledge graph data provided by this application, such as Figure 4As shown, in this embodiment, the security management knowledge graph data labeling device 300 can be set in an electronic device, and the security management knowledge graph data labeling device 300 includes:
[0154] The slicing module 301 is used to slice the target document data for building the safety management knowledge graph to generate corresponding multiple target slice data. The target document data is composed of knowledge data related to industrial and mining safety.
[0155] The knowledge augmentation module 302 is used to input each target slice data into a preset large model, so as to use the preset large model to perform knowledge augmentation processing based on preset prompt words and target slice data, and generate corresponding target candidate data after augmentation.
[0156] The data optimization module 303 is used to respond to the user's review operation and optimization operation on the target candidate data, perform corresponding adjustment processing on the target candidate data, and generate corresponding target knowledge data after adjustment.
[0157] The generation module 304 is used to identify the target entity in the target knowledge data using a preset large model and generate a triple corresponding to the target entity.
[0158] The security management knowledge graph data annotation device provided in this embodiment can be executed Figure 2 The technical solution of the method embodiment shown in the figure has the same implementation principle and technical effect as Figure 2 The method embodiments shown are similar and will not be described in detail here.
[0159] The security management knowledge graph data labeling device provided in this application is based on the security management knowledge graph data labeling device provided in the previous embodiment, and the security management knowledge graph data labeling device is further refined. The security management knowledge graph data labeling device 300 includes:
[0160] Optionally, in this embodiment, the slicing module 301 is specifically used for:
[0161] The target document data is parsed and processed by natural language processing to identify the paragraphs of the target document data. The target document data is divided into a plurality of corresponding target slice data according to the paragraphs.
[0162] or
[0163] The target document data is sliced according to the preset custom slice parameters to generate corresponding multiple target slice data. The preset custom slice parameters include slice length, slice separator, and slice overlap ratio.
[0164] Optionally, in this embodiment, the preset prompt words include question generation prompt words and question answer prompt words.
[0165] The knowledge augmentation module 302 is specifically used for:
[0166] Each target slice data is input into a preset large model, so as to generate corresponding knowledge questions based on the question generation prompt words and the target slice data using the preset large model. The preset large model is used to generate corresponding knowledge answers based on the question answer prompt words and the knowledge questions. The knowledge answers and each target slice data are determined as the corresponding target candidate data after augmentation.
[0167] Optionally, in this embodiment, when the knowledge augmentation module 302 generates corresponding knowledge answers based on the question answer prompt words and the knowledge questions using the preset large model, it is specifically used to:
[0168] The preset large model is used to retrieve the corresponding question-related information from the preset database based on the knowledge question based on the retrieval enhancement generation method RAG. The corresponding knowledge answer is generated based on the question answer prompt word and question-related information.
[0169] Optionally, in this embodiment, when the data optimization module 303 performs corresponding adjustment processing on the target candidate data to generate the adjusted corresponding target knowledge data, it is specifically used to:
[0170] If the user's audit operation is an operation corresponding to the audit passing, the target candidate data that passed the audit will not be modified. If the user's audit operation is an operation corresponding to the audit failing, the target candidate data that failed the audit will be modified accordingly based on the optimization operation to generate corresponding optimized data. The target candidate data that passed the audit and the optimized data are determined as the target knowledge data.
[0171] Optionally, in this embodiment, the security management knowledge graph data labeling device 300 further includes:
[0172] The data acquisition module is used to store the initial knowledge data in a preset database in response to the user's input operation on the initial knowledge data. The initial document data of the security management knowledge graph to be constructed is acquired from the preset database. The initial document data is formatted and processed to generate target document data.
[0173] The security management knowledge graph data annotation device provided in this embodiment can be executed Figure 2-Figure 3 The technical solution of the method embodiment shown in the figure has the same implementation principle and technical effect as Figure 2-Figure 3 The method embodiments shown are similar and will not be described in detail here.
[0174] According to an embodiment of the present application, the present application also provides an electronic device, a computer-readable storage medium and a computer program product.
[0175] like Figure 5 As shown, Figure 5is a schematic diagram of the structure of the electronic device provided by the present application. The electronic device is intended to be electronic devices of various forms, such as a tablet computer, a computer, etc. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0176] like Figure 5 As shown, the electronic device includes: a processor 401 and a memory 402. The various components are connected to each other using different buses and can be installed on a common mainboard or in other ways as required. The processor can process instructions executed in the electronic device.
[0177] The memory 402 is a non-transitory computer-readable storage medium provided in this application. The memory stores instructions executable by at least one processor to enable at least one processor to perform the method for labeling security management knowledge graph data provided in this application. The non-transitory computer-readable storage medium of this application stores computer instructions, which are used to enable a computer to execute the method for labeling security management knowledge graph data provided in this application.
[0178] The memory 402 is a non-transient computer-readable storage medium that can be used to store non-transient software programs, non-transient computer executable programs and modules, such as the program instructions / modules corresponding to the method for labeling security management knowledge graph data in the embodiment of the present application (for example, the attached Figure 4 The processor 401 executes various functional applications and data processing of the electronic device by running the non-transient software programs, instructions and modules stored in the memory 402, that is, the annotation method of the security management knowledge graph data in the above method embodiment is implemented.
[0179] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.
[0180] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.
[0181] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0182] At the same time, this embodiment also provides a computer product. When the instructions in the computer product are executed by the processor of an electronic device, the electronic device can execute the method for labeling security management knowledge graph data of the above embodiment.
[0183] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0184] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0185] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0186] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0188] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0189] Those skilled in the art will appreciate that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed. The aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various media that can store program codes.
[0190] Those skilled in the art will readily come up with other implementations of the embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modifications, uses or adaptations of the embodiments of the present application, which follow the general principles of the embodiments of the present application and include common knowledge or customary technical means in the art that are not disclosed in the embodiments of the present application.
[0191] It should be understood that the embodiments of the present application are not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the embodiments of the present application is limited only by the appended claims.
Claims
1. A method for labeling security management knowledge graph data, characterized in that: include: Slice the target document data for building the security management knowledge graph to generate corresponding multiple target slice data; The target document data is composed of knowledge data related to industrial and mining safety; The preset prompt words include question generation prompt words and question answer prompt words; Inputting each of the target slice data into a preset macro model, so as to generate corresponding knowledge questions based on the question generation prompt words and the target slice data using the preset macro model; Using a preset large model to generate corresponding knowledge answers based on the question answer prompt words and the knowledge questions; Determine the knowledge answer and each of the target slice data as corresponding target candidate data after augmentation; The method of using a preset large model to generate a corresponding knowledge answer based on the question answer prompt word and the knowledge question includes: Adopting a preset large model based on the retrieval enhancement generation method RAG to retrieve corresponding question related information from a preset database based on the knowledge question; Generate a corresponding knowledge answer based on the question answer prompt word and the question related information; In response to the user's review operation and optimization operation on the target candidate data, the target candidate data is adjusted accordingly to generate corresponding target knowledge data after adjustment; A preset large model is used to identify a target entity in the target knowledge data and to generate a triple corresponding to the target entity.
2. The method according to claim 1, characterized in that The target document data to be constructed for the security management knowledge graph is sliced to generate corresponding multiple target slice data, including: Parsing the target document data using a natural language processing method to identify paragraphs of the target document data; Dividing the target document data into corresponding multiple target slice data according to the paragraphs; or The target document data is sliced according to preset custom slice parameters to generate corresponding multiple target slice data; the preset custom slice parameters include slice length, slice separator, and slice overlap ratio.
3. The method according to claim 1, characterized in that: The step of performing corresponding adjustment processing on the target candidate data to generate corresponding target knowledge data after adjustment includes: If the user's review operation is the operation corresponding to the approval, the approved target candidate data will not be modified; If the user's review operation is an operation corresponding to a failure in the review, then based on the optimization operation, the target candidate data that fails in the review is modified accordingly to generate corresponding optimization data; The approved target candidate data and the optimized data are determined as the target knowledge data.
4. The method according to any one of claims 1 to 3, characterized in that: Before slicing the target document data for constructing the security management knowledge graph to generate a plurality of corresponding target slice data, the method further includes: In response to a user input operation on the initial knowledge data, storing the initial knowledge data in a preset database; Obtain the initial document data of the security management knowledge graph to be constructed from the preset database; The initial document data is format converted to generate the target document data.
5. A device for labeling security management knowledge graph data, characterized in that: include: The slicing module is used to slice the target document data for building the security management knowledge graph and generate corresponding multiple target slice data; The target document data is composed of knowledge data related to industrial and mining safety; The preset prompt words include question generation prompt words and question answer prompt words; A knowledge augmentation module, used for inputting each of the target slice data into a preset large model, so as to generate corresponding knowledge questions based on the question generation prompt words and the target slice data using the preset large model; Generate corresponding knowledge answers based on the question answer prompt words and the knowledge questions using a preset large model; determine the knowledge answers and each of the target slice data as corresponding target candidate data after augmentation; When the knowledge augmentation module generates corresponding knowledge answers based on the question answer prompt words and the knowledge questions using a preset large model, it is specifically used to: Adopting a preset large model based on the retrieval enhancement generation method RAG to retrieve corresponding question related information from a preset database based on the knowledge question; generating a corresponding knowledge answer based on the question answer prompt word and the question related information; A data optimization module, configured to perform corresponding adjustment processing on the target candidate data in response to the user's review operation and optimization operation on the target candidate data, and generate corresponding target knowledge data after adjustment; A generation module is used to use a preset large model to identify the target entity in the target knowledge data and generate a triple corresponding to the target entity.
6. An electronic device, characterized in that: include: Memory and processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method for labeling security management knowledge graph data as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method for labeling security management knowledge graph data as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the method for labeling security management knowledge graph data as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Knowledge graph construction method and device, equipment and storage medium
CN117910563A
Method for constructing vertical domain knowledge graph and related device
CN118504679A
Knowledge extraction method, apparatus, electronic device, and storage medium
WO2021212682A1