Industrial Knowledge Extraction and Editing System and Method Based on Human-Machine Collaboration

By building a knowledge extraction model and a human-computer collaboration system based on deep learning algorithms, the problems of insufficient semantic understanding ability and difficulty in taking into account both efficiency and accuracy in the industrial knowledge management system are solved, and efficient and accurate industrial knowledge extraction and editing are achieved.

CN119830886BActive Publication Date: 2025-06-13ANHUI HIGH QUALITY MINING TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510308925.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing industrial knowledge management system has problems such as insufficient semantic understanding ability, difficulty in taking into account the efficiency and accuracy of extraction, and the single human-computer collaboration function.

Method used

A knowledge extraction model based on deep learning algorithms is adopted, combined with change analysis, audit decision-making and user interaction modules, an industrial knowledge extraction and editing system based on human-computer collaboration is built. The system improves semantic understanding ability through model training strategies of preprocessing and entity annotation, pre-training and supervised training, and improves the accuracy and efficiency of knowledge extraction through expert review and change feature analysis.

Benefits of technology

It significantly improves the accuracy and extraction efficiency of semantic understanding of industrial knowledge, ensures high quality and real-time updates of knowledge, solves the problem of difficulty in taking into account efficiency and accuracy, and reduces resource costs through human-machine collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830886B_ABST
    Figure CN119830886B_ABST
Patent Text Reader

Abstract

The present invention discloses an industrial knowledge extraction and editing system and method based on human-machine collaboration. The method includes: preprocessing the original industrial text and inputting it into a knowledge extraction model to obtain a model output, and obtaining a number of structured knowledge and initial credibility based on the model output; collecting historical modification records of a number of structured knowledge, and obtaining change features according to the historical modification records; calculating expert credibility based on the historical review data of a number of experts, and obtaining a review priority according to the expert credibility and change indicators; displaying and modifying the structured knowledge according to the review priority, and updating the structured knowledge credibility and expert credibility according to the expert modification records. The present invention relates to the technical field of industrial knowledge management, and solves the technical problems of insufficient semantic understanding ability, difficulty in balancing extraction efficiency and accuracy, and single human-machine collaboration function existing in the existing industrial support extraction and editing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial knowledge management, involves deep learning technology, and specifically relates to an industrial knowledge extraction and editing system and method based on human-machine collaboration. Background Art

[0002] With the in-depth development of industrial digital transformation, the industrial knowledge management system, as the core infrastructure of intelligent manufacturing, undertakes the important function of structurally storing knowledge such as equipment parameters, process standards, and fault cases, and provides key decision-making support for production optimization, equipment maintenance, and supply chain management. However, the multi-source heterogeneous characteristics of industrial knowledge (covering ERP master data, technical documents, maintenance logs, etc.), the high density of professional terms, and the frequent dynamic updates make traditional technical solutions face significant challenges in the knowledge extraction and maintenance links.

[0003] The current mainstream technologies mainly rely on two methods: rule-based text mining and expert manual maintenance. In the rule-based text mining technology, although natural language processing models can be used to extract entities and relationships from unstructured text, their semantic understanding ability has obvious limitations. The general pre-trained model has insufficient context representation of industrial domain professional terms, which easily leads to entity recognition deviation, such as confusing equipment models with irrelevant codes. In addition, the extraction results often have the problem of knowledge fragmentation due to the lack of cross-document association verification, and automatic resolution cannot be achieved when parameter conflicts occur in different data sources. The expert manual annotation and maintenance mode can ensure the accuracy of knowledge, but it faces the problems of efficiency bottlenecks and collaboration costs. The manual processing speed is difficult to match the real-time update scale of industrial data, and cross-departmental knowledge collaboration requires complex approval processes, resulting in a lag in the update of key information.

[0004] The above defects jointly trigger the "quality-efficiency" paradox of the industrial knowledge management system: if the wrong parameters extracted automatically are not corrected in time, the credibility of the knowledge base will continue to decline; while high-precision manual maintenance is accompanied by a sharp increase in resource costs. In addition, low-quality knowledge may lead to intelligent decision-making errors, causing major risks such as production interruptions. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes an industrial knowledge extraction and editing system and method based on human-machine collaboration, which is used to solve the technical problems of insufficient semantic understanding ability, difficult to balance extraction efficiency and accuracy, and single human-machine collaboration function existing in the existing industrial knowledge management system.

[0006] To achieve the above object, the first aspect of the present invention provides an industrial knowledge extraction and editing system based on human-machine collaboration, including:

[0007] Knowledge extraction module: It is used to preprocess the original industrial text and input it into the knowledge extraction model to obtain the model output, and obtain a number of structured knowledge and initial credibility based on the model output; among them, the knowledge extraction model is constructed based on deep learning algorithms;

[0008] Change analysis module: It is used to collect the historical modification records of a number of structured knowledge and obtain change features according to the historical modification records; among them, the change features include knowledge change indicators and knowledge vulnerability scores;

[0009] Review decision-making module: It is used to calculate the expert credibility according to the historical review data of a number of experts, and obtain the review priority according to the expert credibility and change features;

[0010] User interaction module: It is used to provide an interactive page for expert review, display and modify structured knowledge according to the review priority, and update the structured knowledge credibility and expert credibility according to the expert modification records.

[0011] It should be noted that the knowledge extraction module, change analysis module, review decision-making module and user interaction module of the present invention are in communication connection.

[0012] Furthermore, the knowledge extraction model is constructed based on deep learning algorithms, including:

[0013] A1-1, adding a Bi-LSTM layer and a CRF layer in sequence to the BERT output layer in the BERT-Base-Chinese model to construct an initial model;

[0014] A1-2, collecting a number of industrial texts, preprocessing and entity annotating the number of industrial texts to obtain an unannotated text set and an annotated data set; among them, the annotated data set includes a text set and a corresponding annotation set;

[0015] A1-3, using the unannotated text set to pre-train the initial model to obtain a pre-trained model;

[0016] A1-4, inputting the annotated data set into the pre-trained model for supervised training to obtain a knowledge extraction model; among them, the input of the knowledge extraction model is industrial text, and the output is industrial labels and label prediction probabilities, and the industrial labels include material names, brands, models, specification parameters, and others.

[0017] In the construction of the knowledge extraction model, the combination of the context understanding ability of BERT and the sequence modeling ability of Bi-LSTM can effectively handle the long-distance dependence relationships in complex industrial texts and improve the accuracy of feature extraction. At the same time, the model training strategy that combines pre-training and supervised training makes full use of the domain-general features of unlabeled data, reduces the dependence on labeled data, and significantly improves the generalization ability of the model in industrial scenarios, ensuring its adaptability to diverse industrial text inputs.

[0018] Furthermore, obtaining several structured knowledges and initial credibility based on the model output includes:

[0019] A2-1, defining the necessary attributes, specification attributes, and feature attributes of the standard knowledge template to obtain the standard knowledge template;

[0020] A2-2, unifying the format of industrial labels in the model output according to the standard knowledge template to obtain several structured knowledges;

[0021] A2-3, calculating the text features of the structured knowledge according to the standard knowledge template; where the text features include integrity C, normality S, consistency K, source level L, and timeliness score F;

[0022] A2-4, according to the text quality calculation formula: TQS = α 1 ×C + α 2 ×S + α 3 ×K to calculate the text quality score TQS of the original industrial text;

[0023] A2-5, according to the data source score calculation formula: SCS = β 1 ×L + β 2 ×F to calculate the data source score SCS of the original industrial text;

[0024] A2-6, according to the formula MCS = 1 / n∑P i to calculate the model confidence score MCS; where n represents the number of industrial labels, and P i represents the prediction probability of the i-th industrial label;

[0025] A2-7, using the weighted summation method, calculating the initial credibility IC of several structured knowledges according to the text quality score TQS, data source score, and model confidence score MCS; where α 1 、α 2 、α 3 and β 1 、β 2 represent the weight coefficients corresponding to each formula respectively, and the sum of the corresponding weight coefficients is 1.

[0026] To solve the problems of chaotic industrial text formats and fragmented knowledge, the text format is unified by defining a standard knowledge template, and the text quality is quantified from dimensions such as integrity, standardization, and consistency. At the same time, the reliability of data is evaluated by combining the data source level and timeliness, and the stability of the prediction result is reflected by the model confidence. Finally, the initial credibility is calculated by comprehensively considering multiple factors through the weighted summation method to comprehensively reflect the reliability of knowledge, providing clear quantitative evaluation indicators for subsequent review and application.

[0027] Furthermore, the text features of the structured knowledge calculated according to the standard knowledge template include:

[0028] The ratio of the total number of attributes N actually filled in the structured knowledge total to the number of necessary attributes in the standard knowledge template is marked as the integrity C;

[0029] The ratio of the number of attributes in the structured knowledge that conform to the standard knowledge template to the total number of attributes N total is marked as the standardization S;

[0030] The number of attributes in the structured knowledge that do not conform to the standard knowledge template is marked as the number of contradictory attributes N C , and the consistency K is calculated according to the formula K = 1 - N C / N total ;

[0031] The structured knowledge is divided according to the industrial text data source and the source level is set to obtain the source level L;

[0032] And the timeliness score F is calculated according to the time interval between the generation time of the industrial text data and the current time , and the formula is .

[0033] Furthermore, the change features obtained according to the historical modification records include:

[0034] B1-1. The modification frequency score FS is calculated according to the number of modifications of the structured knowledge in the historical modification records, and the formula is: ; where represents the modification frequency within a preset short-term time, represents the modification frequency within a preset medium-term time, represents the modification frequency within a preset long-term time, , , represent the weight coefficients;

[0035] B1-2. The structured knowledge is divided into numerical knowledge and categorical knowledge according to the data type, and the structured knowledge V before modification in the historical modification records is collected oldand the modified structured knowledge V new ;

[0036] B1-3. For numerical knowledge, the modification amplitude score AS is calculated according to the formula ; for categorical knowledge, the modification amplitude score AS is calculated according to the formula ; where represents the preset reference value, represents the structured knowledge vector before modification, represents the structured knowledge vector after modification;

[0037] B1-4. Establish the knowledge graph association relationship based on a number of structured knowledge, and determine the direct associated node ND and the indirect associated node NI of each node; among them, in the knowledge graph association relationship, the structured knowledge entries are used as nodes;

[0038] B1-5. Calculate the influence range score IS according to the formula ; where represents the node importance of the direct associated node, represents the node importance of the indirect associated node, represents the level of the indirect associated node, i represents the structured knowledge index, I max represents the maximum value of the node importance among a number of associated nodes;

[0039] B1-6. Perform a weighted sum of the knowledge change indicators to obtain the knowledge vulnerability score VS; among them, the knowledge change indicators include the modification frequency score FS, the modification amplitude score AS, and the influence range score IS.

[0040] The change characteristics evaluate the stability of knowledge from three dimensions: modification frequency, amplitude, and influence range by analyzing historical modification records. Among them, the knowledge graph association relationship calculates the influence range through node importance, quantifying the chain effect of knowledge changes; the differential processing of numerical and categorical knowledge ensures the accurate calculation of the modification amplitude; finally, the obtained knowledge vulnerability score provides a basis for the review priority, helping the system dynamically track the change risks of knowledge, identify frequently modified or highly influential knowledge, and prioritize the handling of potential problems.

[0041] Furthermore, calculating the expert credibility based on the historical review data of several experts includes:

[0042] C1-1. Collect the historical review data of several experts, including the number of tasks dispatched N A 、the number of tasks accepted N R 、review content, review results, review records, the number of reviews N T ; among them, the historical review data is stored in the review database;

[0043] C1-2. Calculate the accuracy score EA of the current expert based on the review result, the number of reviews, and the time difference between the review time in the review record and the current time. The formula is as follows: , where the formula is: ; among them, represents the review result, and represents the correctly reviewed structured knowledge, represents the incorrectly reviewed structured knowledge, represents the total number of reviews, and i represents the structured knowledge index;

[0044] C1-3. Use the cosine similarity calculation formula to count the similar review records with a similarity greater than the preset similarity threshold to the review content of the current expert, and obtain the total number of similar cases , the number of cases with the same review result and the number of cases with the same review result as that of the highly credible experts ; among them, the highly credible experts include system-defined experts and experts with an expert credibility greater than the preset credibility threshold obtained through the update of expert credibility;

[0045] C1-4. Calculate the consistency score EC of the current expert according to the formula ;

[0046] C1-5. Calculate the activity score EAct of the current expert according to the formula ; among them, represents the number of tasks with detailed descriptions in the review result;

[0047] C1-6. Use the weighted summation method to calculate the expert credibility E based on the accuracy score EA, the consistency score EC, and the activity score EAct of the current expert. Among them, , and , , respectively represent the weight coefficients of the corresponding formulas.

[0048] In the formula for calculating expert credibility, the accuracy score reflects the correctness of the expert's review result; the consistency score considers the consistency of the expert's review result with that of other experts (especially highly credible experts); the activity score combines the timely response rate , the task completion rate and the detailed description rate , reflecting the enthusiasm and dedication of the expert in participating in the review work, and can more comprehensively and accurately reflect the actual ability and work performance of the expert.

[0049] Further, obtaining the review priority based on the expert credibility and change characteristics includes:

[0050] C2-1, performing a weighted sum of the modification frequency score FS, the impact scope score IS, and the knowledge vulnerability score VS in the change metrics to obtain the change risk CR;

[0051] C2-2, according to the formula calculate to obtain the review priority Pr of the structured knowledge; where, 、 represent the weight coefficients, and their sum is 1.

[0052] In the process of determining the review priority, by integrating indicators such as knowledge vulnerability and modification frequency in the change risk, it ensures the priority processing of high-risk knowledge; and combines the expert credibility to ensure that high-capability experts handle key tasks, balancing the efficiency and accuracy of the review, achieving an efficient allocation of review resources to improve the reliability of the review results.

[0053] Further, the display and modification of the structured knowledge according to the review priority includes:

[0054] Sort and display the structured knowledge according to the size of the review priority, and obtain several assigned experts for the structured knowledge according to the preset expert assignment rules;

[0055] Multiply the expert credibility E of the assigned expert by the domain weight to obtain the expert weight ; where j represents the expert index;

[0056] For numerical knowledge, collect the proposed modification values Vp of several assigned experts and the expert approval values Ve of high-credible experts, and according to the formula Eo j =1 - Vp - Ve / Vp calculate to obtain the expert opinion value Eo of the jth expert j ;

[0057] For categorical knowledge, when the assigned expert agrees to modify, the expert opinion value Eo j =1; when the assigned expert disagrees to modify, the expert opinion value Eo j =0;

[0058] According to the formula calculate to obtain the fusion result value FinalScore;

[0059] According to the formula calculate to obtain the opinion divergence degree DS;

[0060] Determine whether the opinion divergence degree DS is greater than the preset divergence threshold; if so, modify the structured knowledge with the expert opinion value of the highest expert weight; if not, modify the structured knowledge with the fusion result value.

[0061] Further, updating the credibility of the structured knowledge and the credibility of the expert according to the expert modification record includes:

[0062] According to the formula IC new_i =IC i ×(1 + 0.1×(1 - DS)) to calculate the updated credibility IC of the structured knowledge new_i ; where IC i represents the initial credibility of the i-th structured knowledge;

[0063] According to the formula E new_j =E j ×(1 + 0.05×(1 - Eo j - FinalScore )) to calculate the updated credibility E of the j-th expert new_j ; where E j represents the credibility of the j-th expert.

[0064] The smaller the opinion divergence degree, the greater the improvement in the credibility of the structured knowledge, which encourages experts to reach a consensus; the smaller the opinion deviation value, the greater the improvement in the credibility of the expert, which encourages experts to provide accurate opinions; through this dynamic update mechanism, the knowledge quality and expert ability are continuously optimized, forming a closed-loop feedback system to reduce the need for repeated reviews and further improve the review quality, promoting the continuous improvement of the overall performance of the system.

[0065] The second aspect of the present invention provides an industrial knowledge extraction and editing method based on human-machine collaboration, including:

[0066] Preprocess the original industrial text and input it into the knowledge extraction model to obtain the model output, and obtain a number of structured knowledge and initial credibility based on the model output; where the knowledge extraction model is constructed based on a deep learning algorithm;

[0067] Collect the historical modification records of a number of structured knowledge, and obtain the change characteristics according to the historical modification records; where the change characteristics include knowledge change indicators and knowledge vulnerability scores;

[0068] Calculate the credibility of the experts according to the historical review data of a number of experts, and obtain the review priority according to the credibility of the experts and the change characteristics;

[0069] Display and modify the structured knowledge according to the review priority, and update the credibility of the structured knowledge and the credibility of the experts according to the expert modification records.

[0070] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0071] By constructing a knowledge extraction model based on deep learning algorithms (successively adding a Bi-LSTM layer and a CRF layer to the BERT output layer in the BERT-Base-Chinese model), the present invention gives full play to the powerful context understanding ability of BERT and the advantage of Bi-LSTM in capturing long-distance dependencies, and can deeply understand professional terms and complex concepts in industrial texts. After pre-training and supervised training, the model can more accurately identify entities and relationships in industrial texts, significantly improving the accuracy of semantic understanding and providing a reliable basis for subsequent knowledge extraction and application;

[0072] The present invention adopts a human-machine collaboration method. The knowledge extraction module first uses the knowledge extraction model to automatically preprocess and extract knowledge from the original industrial text, obtaining structured knowledge and initial credibility, realizing the efficient extraction of most knowledge. At the same time, the change analysis module analyzes the change characteristics of knowledge, the review decision module determines the review priority according to expert credibility and change indicators, and the user interaction module allows experts to review and modify the structured knowledge according to the review priority. This method combining the efficiency of automatic extraction and the accuracy of expert review not only improves the extraction efficiency but also ensures the extraction quality, effectively solving the problem that it is difficult to balance efficiency and accuracy in the prior art;

[0073] The user interaction module of the present invention provides an interactive page for expert review. Experts can display and modify the structured knowledge according to the review priority. During the modification process, different processing methods are used for numerical knowledge and categorical knowledge, and by calculating indicators such as expert opinion values, fusion result values, and opinion divergence degrees, the opinions of experts are comprehensively considered to determine the final modification plan. This interactive knowledge editing method provides a powerful editing function for experts, enabling experts to easily refine and supplement the extracted knowledge, better meeting the needs of experts for knowledge optimization, thus making the knowledge system more perfect and accurate and improving the application value of knowledge. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0075] Figure 1 It is a framework diagram of an industrial knowledge extraction and editing system based on human-machine collaboration provided by the present invention;

[0076] Figure 2 It is a schematic flowchart of the industrial knowledge extraction and editing method based on human-machine collaboration provided by the present invention;

[0077] Figure 3 It is a schematic flowchart of the construction process of the knowledge extraction model provided by the present invention;

[0078] Figure 4 It is a schematic flowchart of the working process of the change analysis module provided by the present invention. Detailed implementation manners

[0079] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0080] Please refer to Figures 1-4 , the first aspect embodiment of the present invention provides an industrial knowledge extraction and editing system based on human-machine collaboration, including:

[0081] Knowledge extraction module: used to preprocess the original industrial text and input it into the knowledge extraction model to obtain the model output, and obtain a number of structured knowledges and initial credibility based on the model output; wherein, the knowledge extraction model is constructed based on deep learning algorithms;

[0082] Change analysis module: used to collect the historical modification records of a number of structured knowledges, and obtain change features according to the historical modification records; wherein, the change features include knowledge change indicators and knowledge vulnerability scores;

[0083] Review decision module: used to calculate the expert credibility according to the historical review data of a number of experts, and obtain the review priority according to the expert credibility and change features;

[0084] User interaction module: used to provide an interactive page for expert review, display and modify the structured knowledge according to the review priority, and update the structured knowledge credibility and expert credibility according to the expert modification records.

[0085] Based on the above technical modules, the present invention realizes the efficient automatic extraction and dynamic optimization management of industrial knowledge; specifically, the present invention significantly improves the accuracy and domain adaptability of knowledge extraction through the closed-loop collaboration mechanism of the deep learning model and expert review; the system innovatively constructs a multi-dimensional credibility evaluation system, dynamically associates the knowledge vulnerability score with the expert credibility, and realizes the intelligent scheduling of review resources and priority decision-making; through the in-depth mining of the knowledge evolution law by the change feature analysis module, the weak links in the knowledge system are effectively identified, providing data support for the continuous optimization of the industrial knowledge base; its human-machine collaboration mechanism not only reduces the manual review cost, but also real-time corrects the knowledge credibility weight through expert feedback, forming a positive cycle of iterative improvement of knowledge quality, and finally constructs an industrial knowledge extraction and editing system with self-evolution ability:

[0086] Knowledge extraction module: used to preprocess the original industrial text and input it into the knowledge extraction model to obtain the model output, and obtain a number of structured knowledge and initial credibility based on the model output;

[0087] To achieve the accurate conversion of industrial knowledge from unstructured text to standardized semantic representation, in the knowledge extraction module, a knowledge extraction model constructed based on deep learning algorithms is used to intelligently parse multi-source industrial texts, and then the model output is converted into structured knowledge of industrial texts based on standard knowledge templates, and the initial credibility of the structured knowledge is obtained according to the text features of the structured knowledge;

[0088] In one implementation, the construction process of the knowledge extraction model in the knowledge extraction module may include the following steps:

[0089] First, prepare and preprocess the data:

[0090] To achieve the comprehensive coverage of industrial knowledge, the system constructs a multi-source heterogeneous data fusion mechanism and integrates the following core data sources:

[0091] Large-scale industrial material text corpus (covering fields such as machinery, electricity, and automation);

[0092] ERP system material master data (including more than 150,000 material codes and their classification systems);

[0093] Structured data of purchase orders (parsing supplier delivery parameters in XML / JSON format);

[0094] Material technical specifications (including PDF / scanned documents processed by OCR);

[0095] Equipment maintenance manuals (including fault code mapping tables and maintenance record time series data);

[0096] And perform the following preprocessing process on multi-source industrial text data:

[0097] ① Text cleaning: Use regular expressions to remove special symbols (such as "■", "※"), develop a standardization converter for measurement units (such as unifying "380 volts → 380V", "50 hertz → 50Hz", "3.7 kilowatts → 3.7kW"), and construct a standardized industrial text library;

[0098] ② Domain word segmentation: Improve the Jieba word segmentation tool based on an industrial dictionary, and establish a professional term protection list (such as the whole recognition of "harmonic reducer" to avoid incorrect splitting into "harmonic / reducer");

[0099] ③ Annotation data construction: Use the BIO annotation scheme to annotate entity boundaries and types, and obtain the following example of an annotation set: {

[0100] B-NAME, I-NAME, / / Material name

[0101] B-BRAND, I-BRAND, / / Brand

[0102] B-MODEL, I-MODEL, / / Model

[0103] B-SPEC, I-SPEC, / / Specification parameters

[0104] O / / Others

[0105] }

[0106] This preprocessing process forms a closed loop with the subsequent model architecture - the cleaned data is directly input into the domain pre-training module, while the annotation data serves as a supervision signal to guide the fine-tuning stage.

[0107] Next, design the model architecture:

[0108] Based on the particularity of the industrial scenario, adopt a progressive domain adaptation architecture, and its technical evolution path is as follows:

[0109] A [General language understanding] --> B [Industrial semantic enhancement] --> C [Sequence annotation optimization];

[0110] Based on the above technical evolution path, in this embodiment, starting from BERT-Base-Chinese, perform domain enhancement improvements, mainly including:

[0111] ① Vocabulary expansion: Extract the top 5% of high-frequency professional terms from several specifications through the TF-IDF algorithm, and add several professional entries (such as "Modbus-TCP protocol", "IP protection level", etc.);

[0112] ② Domain pre-training: Continuously pre-train on several industrial texts and introduce a parameter contrast loss: ; where represents a pair of input text samples, which are different description segments of the same device, such as the "rated voltage" field in a technical manual and the "input voltage" field in a purchase order); P is a set of positive sample pairs, representing text pairs from the same device model and the same parameter category, such as "380V" and "380 volts" both describe voltage values; N is a set of negative sample pairs, representing texts from different devices or different parameter categories, such as taking the voltage parameter "380V" and the power parameter "3.7KW" as a negative pair; represents the exponential similarity value, usually measuring the semantic similarity of two text descriptions with cosine similarity, represents the model output; among which the positive sample pair are text descriptions from the same device model, and the negative sample N is randomly sampled, forcing the model to learn the deep semantic associations of device parameters;

[0113] Then, aiming at the common multi-value juxtaposition and cross-sentence association characteristics in industrial texts, a hierarchical decoding architecture is designed:

[0114] Bi-LSTM layer (hidden_size = 256, layers = 2): Capture long-distance dependencies, for example, establish the association between temperature values and condition descriptions in "operating temperature range: -20 ~ +60 (under normal pressure conditions)";

[0115] CRF layer: Constrain the label sequence through a transition matrix, prohibiting illegal combinations such as "I-BRAND" appearing after the "O" label;

[0116] Then, a phased progressive training framework is adopted to address the dual challenges of scarce labeled data and semantic complexity in the industrial field, mainly including two stages: domain adaptation pre-training and entity recognition fine-tuning:

[0117] Stage 1: Use unlabeled data for domain adaptation pre-training; in addition to the standard MLM task, an entity boundary prediction task is added. By randomly masking 15% of the entities in unlabeled industrial texts (such as "[MASK] frequency converter"), the model is required to predict the types of the masked entities, enabling the model to learn the general features of the industrial field and the related characteristics of entities in a large amount of unlabeled industrial texts;

[0118] Stage 2: Use labeled data for entity recognition fine-tuning; and use the dynamic loss function for the optimization training of the model, where represents the weight coefficient of the loss function, which can gradually decay from 0.8 to 0.3 as the number of training rounds progresses, gradually strengthening the sequence constraint, represents the cross-entropy loss function, represents the conditional random field loss function; finally, through a two-stage model training, the knowledge extraction model of this embodiment is obtained.

[0119] In one implementation, based on the model output, a number of structured knowledge and initial credibility are obtained, including the following steps:

[0120] A2-1, Define the necessary attributes, specification attributes, and feature attributes of the standard knowledge template to obtain the standard knowledge template;

[0121] A2-2, Unify the format of the industrial labels in the model output according to the standard knowledge template to obtain a number of structured knowledge;

[0122] A2-3, Calculate the text features of the structured knowledge according to the standard knowledge template; among them, the text features include integrity C, normality S, consistency K, source level L, and timeliness score F;

[0123] A2-4, According to the text quality calculation formula: TQS = α 1 ×C + α 2 ×S + α 3 ×K to calculate the text quality score TQS of the original industrial text; where α 1 , α 2 , α 3 are weight coefficients, and the default values are: α 1 = 0.4, α 2 = 0.3, α 3 = 0.3;

[0124] A2-5, According to the data source score calculation formula: SCS = β 1 ×L + β 2 ×F to calculate the data source score SCS of the original industrial text; where β 1 , β 2 are weight coefficients, and the default values are: β 1 = 0.6, β 2 = 0.4;

[0125] A2-6, According to the formula MCS = 1 / n∑P i to calculate the model confidence MCS; where n represents the number of industrial labels, and P i represents the prediction probability of the i-th industrial entity label;

[0126] A2-7, Using the weighted summation method, calculate the initial credibility IC of a number of structured knowledge according to the text quality score TQS, the data source score SCS, and the model confidence MCS, and the default values of the weight coefficients are: WTQS = 0.3, W SCS = 0.4, W MCS = 0.3;

[0127] Among them, the text features of structured knowledge include:

[0128] The ratio of the total number of attributes N total actually filled in the structured knowledge to the number of necessary attributes of the standard knowledge template is marked as the integrity C;

[0129] The ratio of the number of attributes in the structured knowledge that conform to the standard knowledge template to the total number of attributes Ntotal is marked as the standardization S;

[0130] The number of attributes in the structured knowledge that do not conform to the standard knowledge template is marked as the number of contradictory attributes N C , and according to the formula K = 1 - N C / N total calculate the consistency K;

[0131] The structured knowledge is divided according to the industrial text data source and the source level is set, specifically including:

[0132] Level1 (L = 1.0): Original factory technical documents; Level2 (L = 0.8): ERP master data;

[0133] Level3 (L = 0.6): Purchase order data; Level4 (L = 0.4): Manually entered data;

[0134] Obtain the source level L;

[0135] And calculate the timeliness score F according to the time interval between the generation time of the industrial text data and the current time The formula is .

[0136] In one implementation, the definition of the standard knowledge template in step A2-1 can be expressed as follows:

[0137] {

[0138] "template_version": "1.0",

[0139] "material_type": { # Material type

[0140] "electrical_equipment": { # Electrical equipment

[0141] "required_fields": ["name", "brand", "model"], # Required attributes

[0142] "spec_fields": { # Specification attributes

[0143] "electrical": ["voltage", "current", "power"],

[0144] "physical": ["length", "width", "height", "weight"],

[0145] "environmental": ["temperature", "humidity"]

[0146] },

[0147] "feature_fields": ["functions", "certificates", "accessories"] # Feature attributes

[0148] },

[0149] "mechanical_parts": { # Mechanical parts

[0150] "required_fields": ["name", "brand", "model"], # Required attributes

[0151] "spec_fields": { # Specification attributes

[0152] "dimensional": ["size", "thickness", "diameter"],

[0153] "material": ["material_type", "hardness", "finish"]

[0154] },

[0155] "feature_fields": ["usage", "compatibility"] # Feature attributes

[0156] }

[0157] / / Other material type definitions

[0158] }

[0159] }

[0160] In one implementation, the structured knowledge after format unification by the standard knowledge template in step A2-2 can be as follows:

[0161] { # Basic information

[0162] "material_id": "EE20231201001",

[0163] "material_type": "electrical_equipment",

[0164] "base_info": { # Necessary attributes

[0165] "name": "Frequency converter",

[0166] "brand": "Yaskawa", # "Yaskawa" represents the brand name

[0167] "model": "CIMR-HB4A0009FBA",

[0168] "category_code": "EE-VFD-001"

[0169] },

[0170] "specifications": { # Specification attributes

[0171] "electrical": {

[0172] "voltage": {

[0173] "value": 380,

[0174] "unit": "V",

[0175] "standard_unit": "V",

[0176] "confidence": 0.98

[0177] },

[0178] "power": {

[0179] "value": 3.7,

[0180] "unit": "KW",

[0181] "standard_unit": "KW",

[0182] "confidence": 0.95

[0183] }

[0184] }

[0185] },

[0186] "features": { # Feature attributes

[0187] "functions": ["Braking unit"],

[0188] "certificates": [],

[0189] "accessories": []

[0190] },

[0191] "quality_metrics": { # Quality indicators

[0192] "completeness": 0.92,

[0193] "consistency": 0.95,

[0194] "standardization": 0.90

[0195] },

[0196] "meta": { # Metadata information

[0197] "process_time": "2023-12-01 10:30:00",

[0198] "template_version": "1.0",

[0199] "source": "text_extraction"

[0200] }

[0201] }

[0202] In this embodiment, the knowledge extraction module integrates multiple types of data sources such as ERP master data and technical specifications by constructing a multi-source heterogeneous data fusion mechanism. After preprocessing such as text cleaning, domain word segmentation, and BIO annotation, the original text is transformed into high-quality input. Based on a progressive domain adaptation architecture, a high-precision knowledge extraction model is constructed through vocabulary expansion, domain pre-training, and hierarchical decoding (Bi-LSTM+CRF to capture long-distance dependencies). Then, the industrial labels output by the model are mapped to structured data through a standard knowledge template, and the initial credibility is calculated through weighted calculation in three dimensions of text features, data sources, and model confidence, providing a basis for subsequent review priority ranking.

[0203] Change analysis module: used to collect historical modification records of a number of structured knowledge, and obtain change characteristics according to the historical modification records; among them, the change characteristics include knowledge change indicators and knowledge vulnerability scores.

[0204] In order to achieve precise identification and risk assessment of industrial knowledge changes, in the change analysis module, the system first collects the historical modification records of the extracted structured knowledge (including modification time, modification content, modifier, etc.), and then analyzes its change characteristics through multi-dimensional quantification, so as to locate high-risk knowledge items and generate vulnerability scores, providing a basis for review priority ranking.

[0205] Specifically, the workflow of the change analysis module includes:

[0206] 1. Calculate the modification frequency score:

[0207] Extract the number of modifications of each structured knowledge from the historical modification records, and divide it into: the number of modifications N within the short term (within 7 days) short ; the number of modifications N within the medium term (within 30 days) mid and the number of modifications N within the long term (within 180 days) long ;

[0208] Then calculate the ratio with each time window to obtain the modification frequency within the preset short-term time , the modification frequency within the preset medium-term time , and the modification frequency within the preset long-term time ;

[0209] Calculate the modification frequency score FS according to the formula ; where , , represent weight coefficients;

[0210] 2. Calculate the modification amplitude score:

[0211] Divide the structured knowledge into numerical knowledge and categorical knowledge according to the data type, and collect the structured knowledge V before modification in the historical modification records old and the structured knowledge V after modification new ;

[0212] For numerical knowledge, according to the formula calculate the modification amplitude score AS;

[0213] For categorical knowledge, according to the formula calculate the modification amplitude score AS; where represents a preset reference value, obtained based on attribute rating values, industry standards, or historical statistical data represents the structured knowledge vector before modification represents the structured knowledge vector after modification;

[0214] 3. Calculate the influence range score:

[0215] Taking the structured knowledge entries as nodes, establish a knowledge graph through attribute association, and determine the direct associated nodes ND and indirect associated nodes NI of each node; among them, the nodes are mainly divided into three levels, and the corresponding importance degrees are as follows:

[0216] Core nodes (such as device models): 3; Important nodes (such as key parameters): 2; Ordinary nodes (such as accessories): 1;

[0217] Then according to the formula calculate the influence range score IS; where represents the node importance degree of the direct associated nodes represents the node importance degree of the indirect associated nodes represents the level of the indirect associated nodes, i represents the structured knowledge index, I max represents the maximum value of the node importance degrees among several associated nodes, that is, in this embodiment, I max = 3 (core nodes);

[0218] 4. Calculate the knowledge vulnerability score

[0219] According to the weighted summation formula VS = W FS × FS + W AS × AS + W IS × IS calculate the knowledge vulnerability score VS; where, the default values of the weight coefficients are respectively: W FS = 0.4, W AS = 0.3, W IS = 0.3;

[0220] And a grade division is made according to the knowledge vulnerability score:

[0221] High vulnerability: VS >= 0.7; Medium vulnerability: 0.4 <= VS < 0.7; Low vulnerability: VS < 0.4;

[0222] Based on the above work process, in the change analysis module, the system realizes the accurate identification and hierarchical management of industrial knowledge change risks by quantitatively analyzing historical modification records. Specifically, the change analysis module identifies the change frequencies of knowledge items in the short term, medium term, and long term by calculating the modification frequency score; quantifies the severity of numerical or categorical changes through the modification magnitude score; evaluates the chain effect of knowledge changes in the knowledge graph through the influence scope score; and finally obtains the knowledge vulnerability score through the weighted summation method. It not only realizes the traceability of knowledge changes but also provides a scientific basis for review decisions through quantitative indicators, ensuring that high-risk knowledge (such as key parameters with high-frequency modifications) is preferentially reviewed by experts, thereby enhancing the stability and reliability of the industrial knowledge base.

[0223] Review decision-making module: Used to calculate the expert credibility based on the historical review data of several experts and obtain the review priority according to the expert credibility and change characteristics;

[0224] In the review decision-making module, in order to achieve the intelligent matching of expert capabilities and knowledge risks, the system dynamically generates the review priority by quantitatively analyzing the expert credibility in multiple dimensions and combining knowledge change indicators, ensuring that high-risk knowledge is preferentially processed by experts with high credibility;

[0225] In one implementation, the calculation process of the expert credibility in the review decision-making module may include the following steps:

[0226] C1-1, Collect the historical review data of several experts, including the number of task assignments N A and the number of task acceptances N R , review content, review results, review records, and the number of reviews N T ; among them, the historical review data is stored in the review database;

[0227] C1-2, According to the review results, the number of reviews, and the time difference between the review time in the review records and the current time , calculate the accuracy score EA of the current expert, and the formula is: ; among them, represents the review result, and represents the structured knowledge with correct review, represents the structured knowledge with incorrect review, represents the total number of reviews, and i represents the structured knowledge index;

[0228] C1-3. Using the cosine similarity calculation formula, count the similar review records with a similarity greater than the preset similarity threshold to the review content of the current expert, and obtain the total number of similar cases and the number of cases with the same review result as well as the number of cases with the same review result as that of the highly credible expert wherein, the highly credible expert refers to an expert whose expert credibility is greater than the preset credibility threshold;

[0229] C1-4. According to the formula calculate the consistency score EC of the current expert; wherein, the default values of the weight coefficients are respectively: = 0.6, = 0.4;

[0230] C1-5. According to the formula calculate the activity score EAct of the current expert; wherein, represents the number of tasks with detailed descriptions in the review results, and the default values of the weight coefficients are respectively: = 0.4, = 0.3, = 0.3;

[0231] C1-6. Using the weighted summation method, calculate the expert credibility E according to the accuracy score EA, consistency score EC and activity score EAct of the current expert; and the default values of the weight coefficients of each index are respectively: W EA = 0.8, W EC = 0.15, W EAct = 0.05.

[0232] In the expert credibility calculation formula, the accuracy score combines the time decay factor, giving priority to considering the recent review performance, avoiding the influence of outdated historical data on the evaluation, and truly reflecting the current ability; the consistency score matches similar cases through cosine similarity and compares the consistency with highly credible experts to ensure the unity of the review criteria and enhance the authority of the review results; in addition, the activity score comprehensively considers the response speed, task completion rate, and detailed description rate, motivating experts to actively participate in the review and promoting the timely completion of high-quality reviews.

[0233] In one implementation, the review priority in the review decision module can be obtained through the following steps:

[0234] C2-1. According to the formula CR = W FS × FS + W IS × IS + W VS × VS calculate the change risk CR; wherein, the default values of the weight coefficients are respectively: W FS = 0.4, W IS = 0.3, WVS = 0.3

[0235] C2-2, calculate the review priority Pr of the structured knowledge according to the formula ; where 、 represent the weight coefficients, and the default values are respectively: = 0.6, = 0.4.

[0236] The review decision module realizes the intelligent scheduling of review tasks by quantifying expert capabilities and knowledge risks. For example, high-priority tasks are automatically matched with experts with high credibility (E≥0.8) to ensure the review quality of key knowledge; while low-priority tasks can be assigned to novice experts for capacity training.

[0237] User interaction module: used to provide an interactive page for expert review, display and modify structured knowledge according to the review priority, and update the credibility of structured knowledge and expert credibility according to the expert modification record.

[0238] Finally, in the user interaction module, in order to ensure the accuracy and efficiency of industrial knowledge review and the dynamic update of knowledge and expert credibility, by providing an interactive page for expert review, displaying and modifying structured knowledge according to the review priority, and updating the relevant credibility according to the expert modification record, the intelligent and scientific knowledge review process is realized.

[0239] In one implementation, the display and modification of structured knowledge according to the review priority may include the following steps:

[0240] First, sort and display the structured knowledge according to the review priority so that experts can clearly see the knowledge items with different priorities. Then, according to the preset expert assignment rules, determine several assigned experts for each structured knowledge, for example:

[0241] High risk (Pr>=0.8): 3 experts are required; Medium risk (0.5<=Pr<0.8): 2 experts are required; Low risk (Pr<0.5): 1 expert is required;

[0242] For each assigned expert, multiply their expert credibility E by the domain weight to obtain the expert weight (j represents the expert index). Among them, the domain weight reflects the professional degree of the expert in a specific field, can comprehensively consider the overall credibility of the expert and the authority in this field, and is mainly assigned through the education and work background of the expert (including educational background, work experience, professional qualifications, etc.);

[0243] For numerical knowledge, collect the proposed modification values Vp of several distribution experts and the expert approval values Ve of highly credible experts, and according to the formula Eo j = 1 - Vp - Ve / Vp to calculate the expert opinion value Eo j of the j-th expert, to measure the degree of difference between the proposed modification value of the current expert and the expert approval value of the highly credible expert. The smaller the difference, the higher the opinion value of the current expert;

[0244] For categorical knowledge, when the distribution expert agrees to modify, the expert opinion value Eo j = 1; when the distribution expert disagrees to modify, the expert opinion value Eo j = 0;

[0245] Then calculate the fusion result value and the degree of opinion divergence: According to the formula and the formula calculate the fusion result value FinalScore and the degree of opinion divergence DS respectively;

[0246] Finally, judge whether the degree of opinion divergence DS is greater than the preset divergence threshold; if so, use the expert opinion value with the highest expert weight to modify the structured knowledge to ensure that when there are large differences in opinions, the most authoritative expert will lead the modification; if not, use the fusion result value to modify the structured knowledge to reflect the synthesis of the opinions of the majority of experts.

[0247] In one implementation, update the credibility of the structured knowledge and the credibility of the experts according to the expert modification records, which may include the following operation steps:

[0248] According to the formula IC new_i = IC i × (1 + 0.1×(1 - DS)) to calculate the updated credibility IC new_i of the structured knowledge; where IC i represents the initial credibility of the i-th structured knowledge;

[0249] According to the formula E new_j = E j × (1 + 0.05×(1 - Eo j - FinalScore )) to calculate the updated credibility E new_j of the j-th expert; where E j represents the credibility of the j-th expert.

[0250] In the user interaction module, by sorting and displaying structured knowledge according to the review priority, experts can quickly focus on high-risk knowledge, improving the review efficiency; calculating expert weights, expert opinion values, fusion result values, and opinion divergence degrees can comprehensively consider the professional capabilities and opinions of different experts to make more reasonable modification decisions; while updating the credibility of structured knowledge and expert credibility enables the evaluation of knowledge and experts to be continuously adjusted and optimized during the review process, constructing an ever-updating review system and enhancing the quality and level of industrial knowledge management.

[0251] The second aspect embodiment of the present invention provides an industrial knowledge extraction and editing method based on human-machine collaboration, including:

[0252] Preprocess the original industrial text and input it into the knowledge extraction model to obtain the model output, and obtain a number of structured knowledge and initial credibility based on the model output; among them, the knowledge extraction model is constructed based on a deep learning algorithm;

[0253] Collect the historical modification records of a number of structured knowledge, and obtain change features according to the historical modification records; among them, the change features include knowledge change indicators and knowledge vulnerability scores;

[0254] Calculate the expert credibility according to the historical review data of a number of experts, and obtain the review priority according to the expert credibility and change features;

[0255] Display and modify the structured knowledge according to the review priority, and update the credibility of the structured knowledge and the expert credibility according to the expert modification records.

[0256] Some of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is obtained by software simulation of a large amount of collected data to get a formula closest to the actual situation; the preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained by simulating a large amount of data.

[0257] The working principle of the present invention:

[0258] The knowledge extraction module uses a deep learning model to preprocess industrial text, perform entity recognition and structured conversion, and generate structured knowledge with initial credibility;

[0259] The change analysis module calculates the modification frequency, amplitude, and influence range of knowledge by analyzing historical modification records, generates a knowledge vulnerability score, and evaluates the knowledge stability;

[0260] The review decision module dynamically calculates the expert credibility based on the expert historical review data (accuracy, consistency, activity), and determines the review priority in combination with the knowledge vulnerability score;

[0261] Finally, the user interaction module displays knowledge items according to the priority. Experts review and modify them through the interactive interface, and then, by integrating the opinions of experts, the knowledge credibility and expert credibility are dynamically updated to form a cyclic update and optimization.

[0262] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. Industrial knowledge extraction and editing system based on human-machine collaboration, characterized by: include: Knowledge extraction module: used to pre-process the original industrial text and input it into the knowledge extraction model to obtain the model output, and obtain some structured knowledge and initial credibility based on the model output; the knowledge extraction model is built based on the deep learning algorithm; Change analysis module: used to collect historical modification records of some structured knowledge and obtain change features based on the historical modification records; the change features include knowledge change indicators and knowledge vulnerability scores; Audit decision module: used to calculate the expert credibility based on the historical audit data of several experts, and obtain the audit priority based on the expert credibility and change characteristics; User interaction module: used to provide an interactive page for expert review, display and modify structured knowledge according to the review priority, and update the credibility of structured knowledge and expert credibility according to the expert modification records.

2. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 1 is characterized in that: The knowledge extraction model is built based on a deep learning algorithm, including: A1-1, add Bi-LSTM layer and CRF layer to the BERT output layer in the BERT-Base-Chinese model to build the initial model; A1-2, collect a number of industrial texts, preprocess and entity annotate the industrial texts, and obtain an unannotated text set and an annotated data set; wherein the annotated data set includes a text set and a corresponding annotation set; A1-3, pre-train the initial model using the unlabeled text set to obtain a pre-trained model; A1-4, input the labeled data set into the pre-training model for supervised training to obtain a knowledge extraction model; wherein the input of the knowledge extraction model is industrial text, and the output is industrial labels and label prediction probabilities, and the industrial labels include material name, brand, model, specification parameters and others.

3. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 1 is characterized in that: The model-based output obtains several structured knowledge and initial credibility, including: A2-1, define the necessary attributes, specification attributes and characteristic attributes of the standard knowledge template to obtain the standard knowledge template; A2-2, unify the format of industrial labels in the model output according to the standard knowledge template to obtain some structured knowledge; A2-3, calculate the text features of structured knowledge based on the standard knowledge template; the text features include completeness C, standardization S, consistency K, source level L and timeliness score F; A2-4, calculate the text quality score TQS of the original industrial text according to the text quality calculation formula: TQS=α1×C+α2×S+α3×K; A2-5, calculate the data source score SCS of the original industrial text according to the data source score calculation formula: SCS = β1 × L + β2 × F; A2-6, according to the formula MCS = 1 / n∑P i The model confidence MCS is calculated; where n represents the number of industrial tags, P i represents the predicted probability of the i-th industrial label; A2-7, using the weighted summation method, the initial credibility IC of some structured knowledge is calculated according to the text quality score TQS, the data source score and the model confidence MCS; among them, α1, α2, α3 and β1, β2 represent the weight coefficients corresponding to each formula respectively, and the sum of the corresponding weight coefficients is 1.

4. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 3 is characterized in that: The text features of structured knowledge calculated according to the standard knowledge template include: The total number of attributes actually filled in the structured knowledge is N total The ratio of the number of necessary attributes to the standard knowledge template is marked as completeness C; The number of attributes in structured knowledge that conform to the standard knowledge template is divided by the total number of attributes N total The ratio of is marked as the standard degree S; The number of attributes in structured knowledge that do not conform to the standard knowledge template is marked as the number of contradictory attributes N. C , and according to the formula K=1-N C / N total Calculate the consistency K; Divide structured knowledge according to the source of industrial text data and set the source level to obtain the source level L; And according to the time interval between the generation time of industrial text data and the current time Calculate the timeliness score F, the formula is: .

5. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 1 is characterized in that: The change characteristics obtained according to the historical modification records include: B1-1, the modification frequency score FS is calculated based on the number of modifications of structured knowledge in the historical modification records. The formula is: ;in, Indicates the modification frequency within a preset short period of time. Indicates the frequency of modification within the preset medium term. Indicates the modification frequency within a preset long-term period. , , represents the weight coefficient; B1-2, divide the structured knowledge into numerical knowledge and categorical knowledge according to data type, and collect the structured knowledge V before modification in the historical modification records old and the modified structured knowledge V new ; B1-3, for numerical knowledge, according to the formula Calculate the modification score AS; for categorical knowledge, according to the formula The modification score AS is calculated; where: Indicates the preset reference value, represents the structured knowledge vector before modification, represents the modified structured knowledge vector; B1-4, establishing a knowledge graph association relationship based on a number of structured knowledge, and determining the direct associated node ND and the indirect associated node NI of each node; wherein, in the knowledge graph association relationship, the structured knowledge items are nodes; B1-5, according to the formula The influence range score IS is calculated; where: Indicates the node importance of the directly associated node, Indicates the node importance of indirectly related nodes, represents the level of indirectly related nodes, i represents the structured knowledge index, I max Indicates the maximum value of node importance among several associated nodes; B1-6, perform weighted summation on the knowledge change indicators to obtain the knowledge vulnerability score VS; wherein the knowledge change indicators include the modification frequency score FS, the modification amplitude score AS and the impact range score IS.

6. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 5 is characterized in that: The expert credibility is calculated based on the historical review data of several experts, including: C1-1, collect historical audit data of several experts, including the number of tasks assigned N A 、Number of tasks accepted R , audit content, audit results, audit records, audit quantity N T ; Among them, historical audit data is stored in the audit database; C1-2, based on the audit results, audit quantity, and the time difference between the audit time in the audit record and the current time , calculate the accuracy score EA of the current expert, the formula is: ;in, Indicates the audit results, and Represents structured knowledge of correct review, Represents structured knowledge of error review, represents the total number of reviews, i represents the structured knowledge index; C1-3, using the cosine similarity calculation formula, count the similar audit records whose similarity with the current expert's audit content is greater than the preset similarity threshold, and get the total number of similar cases , Number of cases with consistent audit results and the number of cases that agree with the high confidence expert review results ; Wherein, the highly credible experts include system-defined experts and experts whose expert credibility obtained by updating the expert credibility is greater than a preset credibility threshold; C1-4, according to the formula Calculate the consistency score EC of the current expert; C1-5, according to the formula Calculate the current expert's activity score EAct; where, Indicates the number of tasks with detailed descriptions included in the audit results; C1-6, using the weighted summation method, according to the accuracy score EA, consistency score EC and activity score EAct of the current expert, the expert credibility E is calculated; where, , and , , They represent the weight coefficients of the corresponding formulas respectively.

7. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 6 is characterized in that: The review priorities are determined based on the credibility of the experts and the characteristics of the changes, including: C2-1, weighted sum of the modification frequency score FS, impact scope score IS and knowledge vulnerability score VS in the change index to obtain the change risk CR; C2-2, according to the formula The audit priority Pr of structured knowledge is calculated; among them, , Represents the weight coefficient, and the sum is 1.

8. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 6 is characterized in that: The display and modification of structured knowledge according to the audit priority includes: Sort and display structured knowledge according to the review priority, and obtain a number of assigned experts for structured knowledge according to the preset expert assignment rules; Multiply the expert credibility E of the assigned expert by the domain weight to get the expert weight ; Wherein, j represents the expert index; Wherein, the field weight is the weight value preset by the system, which is used to evaluate the professional level of experts in several fields; For numerical knowledge, collect the proposed modification values ​​Vp of several assigned experts and the expert approval values ​​Ve of highly credible experts, and calculate them according to the formula Eo j =1- Vp-Ve / Vp calculates the expert opinion value Eo of the jth expert j ; For categorical knowledge, when the assigned expert agrees to modify, the expert opinion value Eo j =1; when the assigned expert disagrees with the modification, the expert opinion value Eo j =0; According to the formula Calculate and obtain the fusion result value FinalScore; According to the formula The degree of disagreement DS is calculated; Determine whether the disagreement degree DS is greater than the preset disagreement threshold; if yes, use the expert opinion value with the highest expert weight to modify the structured knowledge; if not, use the fusion result value to modify the structured knowledge.

9. The industrial knowledge extraction and editing system based on human-machine collaboration according to claim 8 is characterized in that: The updating of structured knowledge credibility and expert credibility according to the expert modification record includes: According to the formula IC new_i =IC i ×(1+0.1×(1-DS)) to calculate the updated structured knowledge credibility IC new_i Among them, IC i represents the initial credibility of the i-th structured knowledge; According to formula E new_j =E j ×(1+0.05×(1- Eo j -FinalScore )) Calculate the updated expert credibility E of the jth expert new_j ; where E j represents the expert credibility of the jth expert.

10. The method for extracting and editing industrial knowledge based on human-machine collaboration is applied to the system for extracting and editing industrial knowledge based on human-machine collaboration as claimed in any one of claims 1 to 9, characterized in that: include: The original industrial text is preprocessed and input into the knowledge extraction model to obtain the model output, and some structured knowledge and initial credibility are obtained based on the model output; wherein the knowledge extraction model is constructed based on the deep learning algorithm; Collect historical modification records of some structured knowledge, and obtain change features based on the historical modification records; wherein the change features include knowledge change indicators and knowledge vulnerability scores; Calculate the expert credibility based on the historical review data of several experts, and obtain the review priority based on the expert credibility and change characteristics; The structured knowledge is displayed and modified according to the review priority, and the structured knowledge credibility and expert credibility are updated according to the expert modification records.

Citation Information

Patent Citations

  • Knowledge credibility evaluation method and system combined with expert knowledge

    CN117251701A

  • Intelligent bid evaluation expert extraction and management system

    CN119597917A