A precise matching method for quality defect reports and status reports of nuclear power plants

By developing an accurate matching method between quality defect reports and status reports in nuclear power plants, the problem of dispersed quality defect reports and status reports in nuclear power plants is solved, and accurate push is achieved when filling out quality defect reports, improving the filling efficiency and accuracy.

CN114462399BActive Publication Date: 2025-05-13CNNC NUCLEAR POWER OPERATION MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011240359.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-09
Publication Date
2025-05-13
Estimated Expiration
2040-11-09

AI Technical Summary

Technical Problem

The data of quality defect reports and status reports in nuclear power plants are scattered in different business systems, forming information islands, resulting in business personnel being unable to obtain historical experience feedback data in a timely manner, affecting the efficiency and accuracy of quality defect reports.

Method used

Through an accurate matching method between the quality defect report and status report of nuclear power plants, the equipment encoding calculation rules, the semantic similarity calculation rules for nuclear power and keyword processing are used to accurately push historical status report information when filling in the quality defect report.

Benefits of technology

It improves the efficiency and accuracy of filling out quality defect reports, reduces the time and energy of the filler personnel, and ensures the time and energy of the experience feedback information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462399B_ABST
    Figure CN114462399B_ABST
Patent Text Reader

Abstract

This invention discloses a method for accurately matching quality defect reports and status reports in nuclear power plants, comprising the following steps: Step 1: Equipment coding calculation rules; Step 2: Nuclear power-specific semantic similarity calculation rules; Step 3: Calculation of equipment codes and semantic similarity scores for specific reactor types in each power plant; Step 4: Keyword processing to enhance the effectiveness of experience feedback data; Step 5: Intelligent recommendation. The beneficial effects of this invention are: Currently, nuclear power plants can only manually search for historical status report information when filling out quality defect reports, facing problems of low efficiency and low accuracy. The method provided by this invention can automatically and quickly locate and push status report experience feedback data during quality defect report filling, providing a reference for quality defect report filling and reducing the time and effort of personnel filling out quality defect reports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of nuclear power, and in particular to a method for accurately matching status report data according to quality defect reports in nuclear power plants, which provides accurate push of experience feedback information to quality defect report fillers when filling out quality defect reports. Background Art

[0002] After years of operation, nuclear power bases have accumulated a large amount of quality defect report data and status report data in their existing experience feedback systems and business systems. Because this data is scattered across different business systems, it forms information silos and lacks effective integration. Business departments primarily rely on regular push notifications from the experience feedback department for learning. However, these regular push notifications cannot meet the real-time needs of business personnel for historical experience feedback on their current work, often preventing them from receiving the most desired experience feedback data in a timely manner.

[0003] Typically, after a quality defect report is generated, a corresponding status report is developed to analyze the cause and formulate appropriate corrective actions. Therefore, it is necessary to provide an intelligent push method for experience feedback during the quality defect report preparation process. This method can accurately push historical status report information when the nuclear power plant quality defect report filler submits the quality defect report. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for accurately matching quality defect reports and status reports of nuclear power plants. It can perform data analysis based on historical quality defect reports and has a highly accurate recommendation method, which can serve as a guide and reference for the intelligent experience feedback rules of nuclear power plants.

[0005] The technical solution of the present invention is as follows: A method for accurately matching quality defect reports and status reports of nuclear power plants, comprising the following steps:

[0006] Step 1: Equipment code calculation rules;

[0007] Step 2: Nuclear power-specific semantic similarity calculation rules;

[0008] Step 3: Calculate the equipment code and semantic similarity score of each reactor type in each power plant;

[0009] Step 4: Keyword processing to enhance the effectiveness of experience feedback data;

[0010] Step 5: Smart recommendation.

[0011] The step 1 includes:

[0012] Count the rules for various equipment codes, as well as the rules between power plants and reactor types, and classify and calculate reactor types and equipment codes;

[0013] Use relevant regular expressions to determine whether the equipment code of the data complies with the equipment coding rules of its power plant;

[0014] The equipment coding does not comply with the equipment coding rules of the power plant

[0015] If not, then the "QDR Subject" field of the quality defect report and the "CR Subject" field of the status report are removed from the relevant equipment codes and related interference symbols based on natural language processing, and then the natural language semantic similarity matching is performed according to the semantic similarity method. The similarity score is normalized to obtain the matching score w 主题得分 If w 主题得分 Greater than or equal to the given correlation score w 限定分值 , then it is included in the set S 得分集合 ,

[0016] The equipment coding complies with the equipment coding rules of the power plant

[0017] If the input equipment code meets the rules of the power plant, the equipment code field data of the quality defect report is obtained and matched with the pre-processed database equipment code related data:

[0018] Specific device code matching rules:

[0019] Get the equipment field involved in the status report and fully match it with the input equipment code. If they are equal, obtain the relevant equipment code score. If they are not equal, remove the equipment codes on both sides from the unit and then fully match them. If they are equal, obtain the relevant equipment code score. If they are not equal, extract the system code + equipment number from the equipment codes on both sides and then fully match them. If they are equal, obtain the relevant equipment code score. If they are not equal, extract the system code from the equipment codes on both sides and then fully match them. If they are equal, obtain the relevant equipment code score.

[0020] If none of the above rules hold, use regular expressions to extract the relevant equipment codes in the subject and other relationship fields, and fully match them with the input equipment codes. If they are equal, the relevant equipment code score is obtained. If they are not equal, remove the unit from the equipment codes on both sides and then fully match them. If they are equal, the relevant equipment code score is obtained. If they are not equal, extract the system code + equipment number from the equipment codes on both sides and then fully match them. If they are equal, the relevant equipment code score is obtained. If they are not equal, extract the system code from the equipment codes on both sides and then fully match them. If they are equal, the relevant equipment code score is obtained.

[0021] Said step 2 comprises,

[0022] On the basis of matching the reactor type and equipment code type, a nuclear power-specific word segmentation semantic similarity matching method is introduced to achieve higher accuracy. Efficient word graph scanning is achieved based on the prefix dictionary to generate a directed acyclic graph (DAG) consisting of all possible word formation situations of Chinese characters in the sentence; dynamic programming is used to find the maximum probability path and the maximum segmentation combination based on word frequency; for unregistered words, an HMM model based on the word formation ability of Chinese characters is adopted, and the Viterbi algorithm is called. The cosine similarity algorithm is called according to the word segmentation results to obtain the similarity value; here, the semantic similarity calculation is performed using the input (quality defect report) QDR topic and topic-related description and the CR topic and related fields of the status report, and the semantic similarity score is obtained by multiplying the set weight.

[0023]

[0024] Said step 3 comprises,

[0025] (1) The quality defect report belongs to Qinyi Factory QS0

[0026] a) If the "Equipment Code" field in the quality defect report is consistent with the "Involved Equipment" field in the status report (the same equipment), the equipment code score is w a , consider the natural language semantic similarity matching between the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report, normalize the similarity scores, and the highest score w b , the total score of the match w 设备+主题得分 =w a +w b , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0027] b) If the "Equipment Code" field of the quality defect report is inconsistent with the "Involved Equipment" field of the status report, consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding criteria. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field, the score is w c , consider the natural language semantic similarity matching between the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report, normalize the similarity scores, and the highest score w d , the total score of the match w 设备提取+主题得分 =w c +w d , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0028] c) If the "equipment code" field of the quality defect report is inconsistent with the "involved equipment" field of the status report, and does not match the equipment code extracted from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding guidelines, the "QDR subject" field in the defect report data QDR is used to match the "CR subject" field in the status report for natural language semantic similarity matching, the similarity scores are normalized, and the total matching score w e , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0029] (2) The quality defect report belongs to QS3 of Qinsan Factory:

[0030] a) If the "Equipment Code" field in the quality defect report is consistent with the "Involved Equipment" field in the status report (the same equipment), the score is w f If the "Equipment Code" field in the quality defect report is identical to the "Involved Equipment" field in the status report except for the first unit number, the score is w g If the "Equipment Code" field in the quality defect report is the same as the "Involved Equipment" field in the status report, excluding the first unit number and up to the first digit after the second "-" sign, the score is w h , consider the natural language semantic similarity matching between the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report, normalize the similarity scores, and the highest score w i , the total score of the match w 系统得分 =w f 、w g 、w h The highest score in +w i , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0031] b) If the "Equipment Code" field of the quality defect report data QDR is not a), consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding criteria. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field according to a), the score is w j Then, the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report are considered for natural language semantic similarity matching, and the similarity scores are normalized. The highest score w k , the total score of the match w 设备提取+主题得分 =w j +wk , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0032] c) If the "equipment code" field of the quality defect report is not the case of a), and does not match the equipment code extracted from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding guidelines according to a), the "QDR subject" field in the defect report data QDR is used to match the "CR subject" field in the status report for natural language semantic similarity. The similarity score is normalized to 1 point, and the total score of the match is w l , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0033] (3) Quality defect reports belong to other power plants:

[0034] a) If the "Equipment Code" field in the quality defect report is consistent with the "Involved Equipment" field in the status report (the same equipment), the score is w m If the "Equipment Code" field in the quality defect report is identical to the "Involved Equipment" field in the status report except for the first unit number, the score is w n If the "Equipment Code" field in the quality defect report and the "Involved Equipment" field in the status report are the same except for the first unit number and the middle digit code, the score is w o , consider the natural language semantic similarity matching between the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report, normalize the similarity scores, and the highest score w p , the total score of the match w 设备+主题得分 =w m 、w n 、w o The highest score w in p , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0035] b) If the "Equipment Code" field of the quality defect report is not in a), consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding criteria. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field according to the rules in a), the score is w qThen, the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report are matched for natural language semantic similarity (excluding equipment coding and interference symbols), and the similarity scores are normalized. The highest score w r , the total score of the match w 设备提取+主题得分 =w q +w r , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ,

[0036] c) If the "equipment code" field of the quality defect report is not a), consider extracting the equipment code from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding criteria. If the "equipment code" field of the quality defect report matches the extracted equipment code field, and the "equipment code" field of the defect report data QDR does not match the extracted equipment code according to the rules in a), use the "QDR subject" field in the defect report data QDR and the "CR subject" field in the status report to perform natural language semantic similarity matching, normalize the similarity score, and the total matching score w s , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0037] Said step 4 comprises,

[0038] The data of keywords are scored, for example, the quality defect report belongs to the power plant S 得分集合 The same score as the other status report fields,

[0039] (1) If the scores are the same, the "CR Level" field in the status report data is judged. If it is A, an additional point a is added; if it is B, an additional point b is added; if it is C, an additional point c is added; if it is D or is blank, no point is added;

[0040] (2) If there are cases where the scores are the same, the “further action recommendations” field in the status report data is determined to be non-empty (excluding the word “none”), and d points are added;

[0041] (3) Weight value correction score. If there is a situation where the scores are the same, the "status description" of the status report and the "QDR subject" in the quality defect report can semantically match the following keywords. If any keyword is matched and is the same, add "e". For example, the keywords include the following: "inspection", "height work", "drowning", "welding", "ray", "RT", "flaw detection", "corrosion inspection", "scaffolding", "ultrasonic inspection", "radiograph", "electric welding", "gas cutting", "grinding wheel grinding and cutting", "grinding", "baking", "argon arc welding", "gas welding", "in-service inspection". If they match, the recommendation scores will be pushed first if they are the same.

[0042] Said step 5 comprises,

[0043] According to the above matching rules, the quality defect report corresponds to each S of the status report. 得分集合 , make recommendations in descending order of scores, adjust the weights between scores based on business rules and similarity calculation methods, and calculate the similarity, accuracy, and matching rates between the data to obtain the best matching results and achieve accurate push functionality.

[0044] The present invention has the beneficial effect of: existing nuclear power plants can only manually search for historical status report information when filling out quality defect reports, which is inefficient and inaccurate. The method provided by the present invention can automatically and quickly locate and push the status report experience feedback data when filling out quality defect reports, providing a reference for filling out quality defect reports and reducing the time and effort of quality defect report fillers. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a data relationship diagram - a diagram showing the matching of quality defect reports and status reports. DETAILED DESCRIPTION

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] The present invention provides a method for accurately matching quality defect reports and status reports of nuclear power plants. After inputting a quality defect report, the method performs correlation matching based on the power plant, QDR subject, equipment code, and equipment name fields in the quality defect report, and the CR subject, involved equipment, CR level, and status description fields in the status report, based on data accuracy and according to a debugged weighted ratio rule. The result with the highest similarity score between the rule matching and the natural language processing is normalized and used as the estimated value of the data. Based on the precise natural language word segmentation model, the method uses a semantic similarity matching method to adjust the weight of each score according to specific rules to achieve accurate matching between quality defect reports and status reports.

[0048] The present invention provides a method for accurately matching quality defect reports and status reports of nuclear power plants. The method can provide a quantitative technical means for nuclear power plants to achieve accurate push of status reports as experience feedback cases, and promote the effective use of experience feedback information. The semantic similarity method involved in the natural language processing similarity matching is based on the nuclear power professional vocabulary to segment the matching data, vectorize the segmentation results (feature engineering), and then calculate the similarity of the two vectors. The larger the calculated value, the higher the similarity, and vice versa. Finally, the final weight score is formed by the business data score and the semantic similarity matching score, and the weight scores are automatically pushed in descending order.

[0049] A method for accurately matching a quality defect report and a status report of a nuclear power plant comprises the following steps:

[0050] Step 1: Device code calculation rules

[0051] By counting the rules for various equipment codes and the rules between power plants and reactor types, the reactor types and equipment codes are classified and calculated, and the equipment codes are matched according to the matching rules, calculation errors and time can be reduced, while improving the accuracy of data matching.

[0052] The equipment codes are divided into the following categories according to the power plant and reactor type:

[0053] Qinyi Factory QS0 (e.g. "PYLQ-LQS-01-TPC": unit + system code (3 letters) + equipment number (2 / 4 digits) + equipment type (2 / 3 letters)),

[0054] Qinsan Plant QS3 (e.g. "1-21203-EP10008": unit - system code (5 digits) - equipment type (1 / 2 / 3 letters) + equipment number (2 / 3 digits)),

[0055] Other power plants (Qinhuangdao No. 2 Nuclear Power Plant QS2, Fangjiashan Nuclear Power Plant QS1, Changjiang Nuclear Power Plant CJ1, Fuqing Units 5-6 FQH, Fuqing Units 1-4 FQM) (e.g.: "1GSS207LP": unit + system code (3 letters) + equipment number (3 / 4 digits) + equipment type (2 / 3 letters)).

[0056] Use relevant regular expressions to determine whether the equipment coding of the data complies with the equipment coding rules of its power plant.

[0057] 1. The equipment coding does not comply with the equipment coding rules of the power plant

[0058] If not, then based on natural language processing, remove the relevant equipment codes and related interference symbols from the "QDR Subject" field of the quality defect report and the "CR Subject" field in the status report, and then perform natural language semantic similarity matching according to the semantic similarity method in step 3. The similarity score is normalized to obtain the matching score w 主题得分 If w 主题得分 Greater than or equal to the given correlation score w 限定分值 , then it is included in the set S 得分集合 .

[0059] 2. The equipment coding complies with the equipment coding rules of the power plant

[0060] If the input equipment code meets the rules of the power plant, the equipment code field data of the quality defect report is obtained and matched with the pre-processed database equipment code related data:

[0061] Specific device code matching rules:

[0062] Obtain the device field in the status report and fully match it with the entered device code. If they are equal, the relevant device code score is assigned. If they are not equal, remove the unit from both device codes and then fully match them. If they are equal, the relevant device code score is assigned. If they are not equal, extract the system code and device number from both device codes and then fully match them. If they are equal, the relevant device code score is assigned. If they are not equal, extract the system code from both device codes and then fully match them. If they are equal, the relevant device code score is assigned.

[0063] If none of the above rules hold, use regular expressions to extract the relevant device codes from the subject and other relationship fields. These codes are then fully matched against the input device codes. If they are equal, the relevant device code score is assigned. If they are not equal, remove the unit from both device codes and then fully match them. If they are equal, the relevant device code score is assigned. If they are not equal, extract the system code and device number from both device codes and then fully match them. If they are equal, the relevant device code score is assigned. If they are not equal, remove the system code from both device codes and then fully match them. If they are equal, the relevant device code score is assigned.

[0064] Step 2: Nuclear power-specific semantic similarity calculation rules

[0065] Based on the matching of reactor type and equipment code type, a nuclear power-specific word segmentation semantic similarity matching method is introduced to achieve higher accuracy. Efficient word graph scanning is achieved based on a prefix dictionary, generating a directed acyclic graph (DAG) consisting of all possible word formations of Chinese characters in a sentence. Dynamic programming is used to find the maximum probability path and the maximum segmentation combination based on word frequency. For unregistered words, an HMM model based on the word formation ability of Chinese characters is used, and the Viterbi algorithm is invoked. The cosine similarity algorithm is called based on the word segmentation results to obtain a similarity value. Here, the semantic similarity between the input (Quality Defect Report) QDR topic and related description and the CR topic and related fields of the status report is calculated and multiplied by the set weight to obtain the semantic similarity score.

[0066]

[0067] Step 3: Calculate the equipment code and semantic similarity score for each specific power plant reactor type

[0068] (1) The quality defect report belongs to Qinyi Factory QS0

[0069] a) If the "Equipment Code" field in the quality defect report is consistent with the "Involved Equipment" field in the status report (the same equipment), the equipment code score is w a Then, the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report are matched for natural language semantic similarity (excluding device codes and interference symbols), and the similarity scores are normalized. The highest score w b The total score of the match is w 设备+主题得分 =w a +w b , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0070] b) If the "Equipment Code" field of the quality defect report is inconsistent with the "Involved Equipment" field of the status report, consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding guidelines. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field, the score is w c Then, the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report are matched for natural language semantic similarity (excluding device codes and interference symbols), and the similarity scores are normalized. The highest score w d The total score of the match is w 设备提取+主题得分 =w c +w d , only push the total score within the given relevant score w 限定分值Data with scores above 100% are included in the set S 得分集合 .

[0071] c) If the "equipment code" field of the quality defect report is inconsistent with the "involved equipment" field of the status report, and does not match the equipment code extracted from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding guidelines, the "QDR subject" field in the defect report data QDR is used to match the "CR subject" field in the status report with the natural language semantic similarity (excluding equipment codes and interference symbols), and the similarity scores are normalized. The total score of the match is w e , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0072] (2) The quality defect report belongs to QS3 of Qinsan Factory:

[0073] a) If the "Equipment Code" field in the quality defect report is consistent with the "Involved Equipment" field in the status report (the same equipment), the score is w f If the "Equipment Code" field in the quality defect report is identical to the "Involved Equipment" field in the status report except for the first unit number, the score is w g If the "Equipment Code" field in the quality defect report is the same as the "Involved Equipment" field in the status report, excluding the first unit number and up to the first digit after the second "-" sign, the score is w h Consider the natural language semantic similarity matching between the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report (excluding equipment coding and interference symbols), and normalize the similarity score. The highest score w i The total score of the match is w 系统得分 =w f 、w g 、w h The highest score in +w i , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0074] b) If the "Equipment Code" field of the quality defect report data QDR is not a), consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding standards. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field according to a), the score is w jThen, the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report are matched for natural language semantic similarity (excluding device codes and interference symbols), and the similarity scores are normalized. The highest score w k The total score of the match is w 设备提取+主题得分 =w j +w k , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0075] c) If the "equipment code" field of the quality defect report is not the case of a), and does not match the equipment code extracted from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding criteria according to a), the "QDR subject" field in the defect report data QDR is used to match the "CR subject" field in the status report for natural language semantic similarity (excluding equipment code and interference symbols), and the similarity score is normalized to 1 point. The total score of the match is w l , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0076] (3) Quality defect reports belong to other power plants:

[0077] a) If the "Equipment Code" field in the quality defect report is consistent with the "Involved Equipment" field in the status report (the same equipment), the score is w m If the "Equipment Code" field in the quality defect report is identical to the "Involved Equipment" field in the status report except for the first unit number, the score is w n If the "Equipment Code" field in the quality defect report and the "Involved Equipment" field in the status report are the same except for the first unit number and the middle digit code, the score is w o Consider the natural language semantic similarity matching between the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report (excluding equipment coding and interference symbols), and normalize the similarity score. The highest score w p The total score of the match is w 设备+主题得分 =w m 、w n 、w o The highest score w in p , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0078] b) If the "Equipment Code" field of the quality defect report is not in a), consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding criteria. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field according to the rules in a), the score is w q Then, the “QDR Subject” field in the quality defect report and the “CR Subject” field in the status report are matched for natural language semantic similarity (excluding device codes and interference symbols), and the similarity scores are normalized. The highest score w r The total score of the match is w 设备提取+主题得分 =w q +w r , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0079] c) If the "equipment code" field of the quality defect report is not in a), consider extracting the equipment code from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding criteria. If the "equipment code" field of the quality defect report matches the extracted equipment code field, and the "equipment code" field of the defect report data QDR does not match the extracted equipment code according to the rules in a), use the "QDR subject" field in the defect report data QDR and the "CR subject" field in the status report to perform natural language semantic similarity matching (eliminating equipment codes and interference symbols), and normalize the similarity score. The total score of the match is w s , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 .

[0080] Note: above w 限定分值 A value between 0.6 and 1.

[0081] Step 4: Keyword processing to enhance the effectiveness of experience feedback data

[0082] The data of keywords are scored, for example, the quality defect report belongs to the power plant S 得分集合 The score in the same status report as the other fields.

[0083] (1) If the scores are the same, the "CR Level" field in the status report data is judged. If it is A, an additional point a is added; if it is B, an additional point b is added; if it is C, an additional point c is added; if it is D or is blank, no point is added;

[0084] (2) If there are cases where the scores are the same, the “further action recommendations” field in the status report data is determined to be non-empty (excluding the word “none”), and d points are added;

[0085] (3) Weight value correction score. If there is a match between the "status description" of the status report and the "QDR subject" in the quality defect report, and if any of the keywords are matched and the match is the same, add "e". For example, the keywords include the following: "inspection", "height work", "drowning", "welding", "radio", "RT", "flaw detection", "corrosion inspection", "scaffolding", "ultrasonic inspection", "radio inspection", "electric welding", "gas cutting", "grinding", "baking", "argon arc welding", "gas welding", "in-service inspection". If there is a match, the recommendation score will be pushed first if it is the same.

[0086] Step 5: Smart Recommendation

[0087] According to the above matching rules, the quality defect report corresponds to each S of the status report. 得分集合 , recommending in descending order of scores. According to business rules and similarity calculation methods, the weights between scores are adjusted, and the similarity, accuracy, and matching rates between data are analyzed to obtain the best matching results and achieve accurate push notifications.

Claims

1. A method for accurately matching quality defect reports and status reports of a nuclear power plant, characterized by: The following steps are included: Step 1: Equipment code calculation rules; The step 1 comprises: Count the rules for various equipment codes, as well as the rules between power plants and reactor types, and classify and calculate the reactor types and equipment codes; Use relevant regular expressions to determine whether the equipment code of the device complies with the equipment code rules of its power plant; If not, then the "QDR Subject" field of the quality defect report and the "CR Subject" field in the status report are removed from the relevant equipment codes and related interference symbols based on natural language processing, and then the natural language semantic similarity is matched according to the semantic similarity method. The similarity scores are normalized to obtain the matching score w 主题得分 , If w 主题得分 Greater than or equal to the given correlation score w 限定分值 , then it is included in the set S 得分集合 ; If the input equipment code meets the rules of the power plant, the equipment code field data of the quality defect report is obtained and matched with the pre-processed database equipment code related data; Specific device code matching rules: Get the equipment fields involved in the status report and fully match them with the input equipment codes. If they are equal, obtain the relevant equipment code score. If they are not equal, remove the equipment codes on both sides from the unit and then fully match them. If they are equal, obtain the relevant equipment code score. If they are not equal, extract the system code + equipment number from the equipment codes on both sides and then fully match them. If they are equal, obtain the relevant equipment code score. If they are not equal, extract the system code from the equipment codes on both sides and then fully match them. If they are equal, obtain the relevant equipment code score. If none of the above rules are true, use regular expressions to extract the relevant equipment codes in the relationship field of the subject, and fully match them with the input equipment codes. If they are equal, obtain the relevant equipment code score. If they are not equal, remove the equipment codes on both sides from the unit and then fully match them. If they are equal, obtain the relevant equipment code score. If they are not equal, extract the system code + equipment number from the equipment codes on both sides and then fully match them. If they are equal, obtain the relevant equipment code score. If they are not equal, extract the system code from the equipment codes on both sides and then fully match them. If they are equal, obtain the relevant equipment code score. Step 2: Nuclear power-specific semantic similarity calculation rules; The step 2 comprises: On the basis of matching the reactor type and equipment coding type, a nuclear power-specific word segmentation semantic similarity matching method is introduced to achieve higher accuracy. Efficient word graph scanning is realized based on the prefix dictionary to generate a directed acyclic graph DAG consisting of all possible word formation situations of Chinese characters in the sentence; dynamic programming is used to find the maximum probability path and find the maximum segmentation combination based on word frequency; for unregistered words, an HMM model based on the word formation ability of Chinese characters is adopted, and the Viterbi algorithm is called. The cosine similarity algorithm is called according to the word segmentation results to obtain the similarity value; here, the semantic similarity calculation is performed using the input quality defect report QDR topic and topic-related description and the status report CR topic and related fields, and the semantic similarity score is obtained by multiplying the set weight; Step 3: Calculate the equipment code and semantic similarity score of each reactor type in the power plant; The step 3 comprises: (1) Quality defect reports belong to QS0 of nuclear power plants; a) If the "Equipment Code" field of the quality defect report is consistent with the "Involved Equipment" field of the status report, the equipment code score is w a , consider the natural language semantic similarity matching between the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report, normalize the similarity scores, and the highest score w b , the total score of the match is w 设备+主题得分 =w a +w b , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; b) If the "Equipment Code" field of the quality defect report is inconsistent with the "Involved Equipment" field of the status report, consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding criteria. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field, the score is w c , consider the natural language semantic similarity matching between the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report, normalize the similarity scores, and the highest score w d , the total score of the match is w 设备提取+主题得分 =w c +w d , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; c) If the "equipment code" field of the quality defect report is inconsistent with the "involved equipment" field of the status report, and does not match the equipment code extracted from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding criteria, the "QDR subject" field in the defect report data QDR is used to match the "CR subject" field in the status report with the natural language semantic similarity, and the similarity score is normalized. The total score of the match is w e , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; (2) Quality defect reports belong to QS3 of nuclear power plants: a) If the "Equipment Code" field of the quality defect report is consistent with the "Involved Equipment" field of the status report, the score is w f ; If the "Equipment Code" field of the quality defect report is identical to the "Involved Equipment" field of the status report except for the first unit number, the score is w g ; If the "Equipment Code" field of the quality defect report is the same as the "Involved Equipment" field of the status report, starting from the first unit number and ending after the second "-" sign and before the first digit, the score is w h , consider the natural language semantic similarity matching between the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report, normalize the similarity scores, and the highest score w i , the total score of the match is w 系统得分 =w f 、w g 、w h The highest score in +w i , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; b) If the "Equipment Code" field of the quality defect report is not a), consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding criteria. If the "Equipment Code" of the quality defect report matches the extracted equipment code field according to a), the score is w j Then, the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report are considered for natural language semantic similarity matching, and the similarity scores are normalized. The highest score w k , the total score of the match is w 设备提取+主题得分 =w j +w k , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; c) If the "equipment code" field of the quality defect report is not the case of a), and it does not match the equipment code extracted from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding criteria according to a), the "QDR subject" field in the defect report data QDR is used to match the "CR subject" field in the status report for natural language semantic similarity, and the similarity score is normalized to 1 point. The total score of the match is w l , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; (3) Quality defect reports belong to other power plants: a) If the "Equipment Code" field of the quality defect report is consistent with the "Involved Equipment" field of the status report, the score is w m ; If the "Equipment Code" field of the quality defect report and the "Involved Equipment" field of the status report are all the same except for the first unit number; the score is w n ; If the "Equipment Code" field of the quality defect report and the "Involved Equipment" field of the status report are the same except for the first unit number and the middle digit code, the score is w o , consider the natural language semantic similarity matching between the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report, normalize the similarity scores, and the highest score w p , the total score of the match is w 设备+主题得分 =w m 、w n 、w o The highest score w in p , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; b) If the "Equipment Code" field of the quality defect report is not in a), consider extracting the equipment code from the "CR Subject" and "Status Description" fields of the status report according to the power plant equipment coding criteria. If the "Equipment Code" field of the quality defect report matches the extracted equipment code field according to the rules in a), the score is w q Then, the "QDR Subject" field in the quality defect report and the "CR Subject" field in the status report are considered for natural language semantic similarity matching, the equipment code and interference symbols are removed, and the similarity scores are normalized. The highest score w r , the total score of the match is w 设备提取+主题得分 =w q +w r , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; c) If the "equipment code" field of the quality defect report is not in a), consider extracting the equipment code from the "CR subject" and "status description" fields of the status report according to the power plant equipment coding criteria. If the "equipment code" field of the quality defect report can match the extracted equipment code field, and the "equipment code" field of the defect report data QDR and the extracted equipment code do not match according to the rules in a), use the "QDR subject" field in the defect report data QDR and the "CR subject" field in the status report to perform natural language semantic similarity matching, normalize the similarity score, and the total matching score w s , only push the total score within the given relevant score w 限定分值 Data with scores above 100% are included in the set S 得分集合 ; Step 4: Keyword processing to enhance the effectiveness of experience feedback data; The step 4 comprises: The data of keywords are scored, and the quality defect report belongs to the power plant S 得分集合 The other fields of the status report have the same score as in (1) If the scores are the same, the "CR Level" field in the status report data is determined. If it is A, an additional point is added; if it is B, an additional point is added; if it is C, an additional point is added; if it is D or is blank, no additional point is added; (2) If the scores are the same, the "further action suggestions" field in the status report data is checked. If it is not empty, d points are added; (3) Weight value correction score. If the scores are the same, the "status description" in the status report and the "QDR subject" in the quality defect report can semantically match the following keywords. If any keyword is matched and they are the same, add e. The keywords include the following: "inspection", "height work", "drowning", "welding", "ray", "RT", "flaw detection", "corrosion inspection", "scaffolding", "ultrasonic inspection", "ray inspection", "electric welding", "gas cutting", "grinding wheel grinding and cutting", "grinding", "baking", "argon arc welding", "gas welding", "in-service inspection". If they match, the recommendation scores will be pushed first if they are the same; Step 5: Intelligent recommendation; The step 5 comprises: According to the above matching rules, the quality defect report corresponds to each S of the status report. 得分集合 , recommendations are made in descending order of scores, and the weights of the scores are adjusted according to business rules and similarity calculation methods.

2. The method for accurately matching a nuclear power plant quality defect report with a status report according to claim 1, characterized in that: The step 5 includes similarity, accuracy and matching rate between the statistical data, obtaining the best matching result and realizing the accurate push function.

Citation Information

Patent Citations

  • A similarity defect report recommendation method by combining weighted word vectors and latent semantic analysis

    CN109165382A

  • Code defect report retrieval method and device

    CN111339272A