Nuclear power file knowledge element association method, computer equipment and storage medium

By using the nuclear industry semantic library to screen keywords and split molecular chapters to calculate similarity coefficients in nuclear power files and standards, the problem of inaccurate correlation in the existing technology is solved, and a more accurate correlation and utilization of nuclear power files and standards is achieved.

CN120337897APending Publication Date: 2025-07-18CNNC NUCLEAR POWER OPERATION MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379200.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, in the correlation method between nuclear power files and standards, when directly calculating text similarity coefficients, the correlation content is inaccurate due to too much matching text, which affects the calculation results of the similarity coefficients between nuclear power files and standards.

Method used

By counting the number of occurrences of each word in the nuclear power file and the full text of the standard, the keyword coefficient is calculated using the nuclear industry semantic library, the keyword coefficient is selected, and the file is split and the standard is used to compare the full text with the sub-chapter, and the sub-chapter similarity coefficient is calculated to reduce the impact of full text comparison.

Benefits of technology

It improves the correlation accuracy and utilization efficiency of nuclear power files and standards, realizes intelligent mining of keywords and intelligent comparison of content, and ensures the accuracy of the relationship between sub-chapters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337897A_ABST
    Figure CN120337897A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of data mining, and particularly relates to a nuclear power file knowledge element association method, computer equipment and a storage medium, the method comprises the following steps: S1, counting the occurrence frequency of each word in a nuclear power file and a standard full text, and sorting; s2, calculating keyword coefficients in the nuclear power file and the standard full text according to the nuclear industry semantic library, and screening out keywords according to the keyword coefficients; s3, matching nuclear power files and standards with the same or similar keywords, and calculating a file similarity coefficient; s4, sorting the file similarity coefficients, and determining a similarity standard according to the file similarity coefficients; s5, splitting the nuclear power file and the similar standard into sub-sections for full-text comparison one by one, and calculating to obtain a sub-section similar coefficient; and S6, when the sub-chapter similarity coefficient reaches a preset numerical value, considering that the nuclear power file is associated with a certain chapter of the similarity standard. According to the method, related standards can be fully and comprehensively utilized, and the accuracy of nuclear power standard association is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data mining, and particularly relates to a method for associating knowledge elements of nuclear power documents, a computer device, and a storage medium. Background Art

[0002] The management procedures in nuclear power documents refer to the written documents formed by analyzing the management objectives and management elements of nuclear power plants, gradually refining the management requirements of each management element according to the hierarchical logical relationship, and compiling them according to the format requirements, which clarify the work specifications and process requirements in various aspects such as the operation and production, and business services of nuclear power plants. As a guiding document for the management procedures of nuclear power plants, standards have binding force or reference guiding effects on the company's management activities and technical activities.

[0003] A knowledge element refers to the smallest independent unit that cannot be further divided in the knowledge system and is the basic node constituting the knowledge network.

[0004] The knowledge element of nuclear power documents refers to the single knowledge points or knowledge fragments (such as chapters, management terms, charts, formulas, etc.) that are directly extracted from the documents of nuclear power enterprises (such as standards, management procedures, technical outlines, technical guides, technical regulations, etc.), which are independent and complete in the nuclear power industry and can express specific facts, concepts, methods, management, or conclusions.

[0005] The management procedures of nuclear power plants strictly comply with the requirements in relevant standards. Therefore, employees need to search for the standards involved in their business to assist in the compilation of procedures. However, it is difficult to query all the relevant standards in some fields. Therefore, in order to improve the utilization efficiency of nuclear power standards, it is necessary to associate the standards.

[0006] After investigation, currently in the domestic nuclear power industry, the association is between documents. By using standard keywords, the correlation coefficient of the standards is calculated, and a full-text comparison is made between the standards with high coefficients and nuclear power documents to calculate the similarity coefficient between the standard documents and nuclear power documents, so as to determine the associated content. When a user uses a certain nuclear power document, the standard content matching the keyword and the relevant content of the nuclear power standards corresponding to the keywords clustered with the matching keyword can be displayed to the user.

[0007] However, the current method of associating between documents directly calculates the similarity coefficient between nuclear power documents and standard content. When used, the similarity coefficient between the standard documents and nuclear power documents calculated will be affected by too much matching text, resulting in inaccurate associated content.

[0008] Therefore, there is an urgent need to develop a method for associating nuclear power documents that can effectively improve the accuracy of the associated content. Summary of the Invention

[0009] The object of the present invention is to provide a method for associating knowledge elements of nuclear power documents, a computer device and a storage medium. This method can more accurately mine the association relationship between nuclear power documents and sub-chapters of standards, make more full and comprehensive use of relevant standards, and improve the accuracy and utilization efficiency of nuclear power standard association.

[0010] Technical solution for realizing the object of the present invention:

[0011] A method for associating knowledge elements of nuclear power documents, the method comprising:

[0012] S1. Count the number of occurrences of each word in the nuclear power document and the full text of the standard and sort them;

[0013] S2. Calculate the keyword coefficients in the nuclear power document and the full text of the standard according to the nuclear industry semantic library, and screen out keywords according to the keyword coefficients;

[0014] S3. Match nuclear power documents and standards with the same or similar keywords, and calculate the document similarity coefficient;

[0015] S4. Sort the document similarity coefficients, and determine the similar standards according to the document similarity coefficients;

[0016] S5. Split the nuclear power document and the similar standard into sub-chapters and conduct a full-text comparison one by one to calculate the sub-chapter similarity coefficient;

[0017] S6. When the sub-chapter similarity coefficient reaches a preset value, it is considered that a certain chapter of the nuclear power document and the similar standard is associated.

[0018] Further, the S2 includes:

[0019] S2.1. Calculate the full-text occurrence frequency of each word in the nuclear power document and the full text of the standard;

[0020] The formula for calculating the full-text occurrence frequency of each word in the nuclear power document and the full text of the standard is:

[0021] C i,j = N i,j / M j

[0022] Wherein, C i,j represents the full-text coefficient of the i-th word in the j-th nuclear power document or standard, N i,j represents the number of occurrences of the i-th word in the j-th nuclear power document or standard, and M j represents the total number of words in the j-th nuclear power document or standard;

[0023] S2.2. Based on the nuclear industry semantic library, calculate the keyword coefficients, and screen out keywords according to the keyword coefficients;

[0024] The formula for calculating the keyword coefficient is as follows:

[0025] K i,j = C i,j / Ln(N - H i / X)

[0026] Where K i,j represents the keyword coefficient of the i-th word in the j-th nuclear power document or standard, C i,j represents the full text coefficient of the i-th word in the j-th nuclear power document or standard, Ln represents the natural logarithm function, N is the correction coefficient, and H i represents that there are H i words in the semantic library that contain the i-th keyword, and X represents the total number of words in the semantic library;

[0027] N = (H i / X) T,max + 1

[0028] Where (H i / X) T,max is the maximum value of the ratio of each general vocabulary in the nuclear industry semantic library.

[0029] Furthermore, the specific method of screening keywords according to the keyword coefficient in S2 is: select the 3 keywords with the largest keyword coefficients as the screened keywords.

[0030] Furthermore, the formula for calculating the document similarity coefficient in S3 is:

[0031]

[0032] Where represents the document similarity coefficient between the i-th nuclear power document and the j-th standard, A i,t represents the vector of the t-th keyword in the i-th nuclear power document, B j,k represents the vector of the k-th keyword in the j-th standard, ||A i,t - B j,k || represents the distance between vector A i,t and B j,k .

[0033] Furthermore, the specific method of determining the similar standard according to the document similarity coefficient in S4 is: select the 8 standards with the smallest document similarity coefficients as the similar standards.

[0034] Furthermore, S5 includes:

[0035] S5.1. Split the nuclear power document and the similar standard by chapter to obtain sub-chapters;

[0036] S5.2. Segment the respective sub - chapters to be compared to obtain multiple sub - texts corresponding to each sub - chapter respectively;

[0037] S5.3. Obtain the multi - dimensional vectors corresponding to the sub - texts in the same chapter and combine them into a sub - text vector matrix;

[0038] S5.4. Calculate the cosine similarity between each sub - text vector matrix and obtain a chapter similarity matrix;

[0039] S5.5. Calculate the similarity coefficient of two sub - chapters to be compared according to the chapter similarity matrix.

[0040] Further, the S5.5 includes:

[0041] S5.5.1. Calculate the similarity S from sub - chapter A of the nuclear power document to standard sub - chapter B AB

[0042]

[0043] where m is the number of rows of the similarity matrix X, represents the value of the largest element in the j - th column of the i - th row in the similarity matrix X;

[0044] S5.5.2. Calculate the similarity S from standard sub - chapter B to sub - chapter A of the nuclear power document BA

[0045]

[0046] where n is the number of columns of the similarity matrix X, represents the value of the largest element in the i - th row of the j - th column in the similarity matrix X;

[0047] S5.5.3. Calculate the sub - chapter similarity coefficient S between sub - chapter A of the nuclear power document and standard sub - chapter B A-B

[0048] S A-B =(2 * S AB * S BA ) / (S AB + S BA ).

[0049] A computer device, including a memory, a processor, and a computer program stored in the memory and operable on the processor. When the computer program is executed by the processor, it implements the steps of the nuclear power document knowledge element association method described above.

[0050] A computer-readable storage medium has a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the nuclear power document knowledge element association method described above are implemented.

[0051] The beneficial technical effects of the present invention are as follows:

[0052] 1. For a nuclear power document knowledge element association method provided by the present invention, by splitting the sub-chapters to be compared between the nuclear power document and the standard document to obtain sub-texts, then constructing a sub-text similarity matrix by comparing the sub-texts, and calculating the similarity coefficients between the sub-chapters according to the sub-text similarity matrix, it reduces the influence of excessive text in the direct full-text comparison process on the calculation results and improves the accuracy of chapter association.

[0053] 2. For a nuclear power document knowledge element association method provided by the present invention, the mining of the association relationship between the sub-chapters of the nuclear power document and the standard is more accurate, and the standard association processing is more sufficient and comprehensive, which can effectively improve the association and utilization efficiency of nuclear power standards.

[0054] 3. For a nuclear power document knowledge element association method provided by the present invention, it realizes the intelligent mining of keywords and the intelligent comparison of content. When associating the sub-chapters of the nuclear power document and the standard, the mining of keywords is more accurate, and the association relationship between the sub-chapters of the nuclear power document and the standard can be mined more fully and comprehensively. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a flowchart of a nuclear power document knowledge element association method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0057] An embodiment of the present invention provides a nuclear power document knowledge element association method, including:

[0058] S1. Count the number of occurrences of each word in the full text of the nuclear power document and the standard and sort them.

[0059] S2. Calculate multiple keyword coefficients in the full text of the nuclear power document and the standard according to the nuclear industry semantic library, and screen out keywords according to the keyword coefficients.

[0060] Calculate the keyword coefficients in the full text of the nuclear power document and the standard respectively, and then take the 3 keywords with the largest keyword coefficients from the nuclear power document and the 3 keywords with the largest keyword coefficients from the full text of the standard as the screened keywords.

[0061] Specifically, according to the text features or semantic features of the keywords, the nuclear industry semantic library is used to calculate the coefficient of each keyword. Among them, the integration of professional and general vocabulary in the nuclear industry semantic library has been realized based on various existing machine learning models and other technologies, which will not be elaborated in this invention.

[0062] The high-frequency words statistically obtained in S1 cannot be directly used as standard keywords because the general vocabulary existing in the standard full text has a high usage frequency, which may affect the judgment of keywords in the subsequent steps.

[0063] Therefore, to eliminate the interference of general vocabulary, the nuclear industry semantic library is used to screen keywords. The specific steps of S2 are as follows:

[0064] S2.1. Calculate the full-text occurrence frequency of each word in the nuclear power documents and the standard full text;

[0065] Calculate the full-text occurrence frequency of each word in the nuclear power documents and the standard full text, and the full-text coefficient of the word in the document according to Formula 1.

[0066] C i,j =N i,j / M j Formula 1

[0067] Where, C i,j represents the full-text coefficient of the i-th word in the j-th nuclear power document or standard, N i,j represents the number of times the i-th word appears in the j-th nuclear power document or standard, and M j represents the total number of words in the j-th nuclear power document or standard.

[0068] S2.2. Based on the nuclear industry semantic library, calculate the keyword coefficient, and screen out keywords according to the keyword coefficient:

[0069] Calculate the keyword coefficient according to Formula 2, calculate multiple keyword coefficients, sort the keyword coefficients, and select the 3 keywords with the largest keyword coefficients.

[0070] K i,j =C i,j / Ln(N-H i / X) Formula 2

[0071] Where, K i,j represents the keyword coefficient of the i-th word in the j-th nuclear power document or standard, C i,j represents the full-text coefficient of the i-th word in the j-th nuclear power document or standard, Ln represents the natural logarithm function, N is a correction coefficient (with a value between 1 and 2), H i represents that there are H i words in the semantic library that contain the i-th keyword, and X represents the total number of words in the semantic library.

[0072] The calculation method of the correction coefficient N is as follows: By counting the ratio of H i / X of each general term in the nuclear industry semantic library, and taking the maximum H i / X value + 1 counted as the value of the correction coefficient N. That is, the calculation formula of the correction coefficient N is: N = (H i / X) T,max + 1.

[0073] When the ratio H i / X of the general term in the semantic library is higher than the maximum H i / X value counted, that is, H i / X > (H i / X) T,max ), then the value of Ln(N - H i / X) is negative. Furthermore, the value of the keyword coefficient C i,j / Ln(N - H i / X) is negative. By screening the cases where the keyword coefficient is negative, the general terms are screened out, thus avoiding the interference of general terms on the keyword coefficient.

[0074] The higher the keyword coefficient, the higher the importance of the keyword, and the more representative the keyword is of the main information contained in the document.

[0075] Through actual testing and fitting, it is found that the keyword coefficients calculated by Formula 1 and Formula 2 have a better screening effect on keywords. The reason analysis is as follows:

[0076] According to Formula 1, N i,j represents the number of times the i-th word appears in the j-th document, M j represents the total number of words in the j-th document, and N i,j / M j is a value greater than 0 and less than 1;

[0077] More importantly, for general terms, they appear frequently in the full text but are not suitable as keywords. Therefore, in Formula 2, the nuclear industry semantic library is used to screen keywords. H i represents that there are H i words in the semantic library that contain the i-th keyword, X represents the total number of words in the semantic library, and H i / X represents the proportion of words containing high-frequency words in the semantic library. General terms have more application scenarios, so the ratio of H i / X is higher and greater than 0 and less than 1. At the same time, the logarithmic function and the correction coefficient N are used to avoid the influence of general terms on the keyword coefficient.

[0078] In a specific embodiment, the full-text coefficients of "nuclear power" (the target keyword, i.e., a professional term) and "of" (the interfering keyword, i.e., a common term) in a certain document are calculated through Formula 1 as 0.00596896 and 0.0238758 respectively. There are 5,261 words containing "nuclear power" and 61,256 words containing "of" in the nuclear industry semantic library, and the total number of words in the nuclear industry semantic library is 21,892,410. The correction coefficient N is taken as 1.002, and the keyword coefficients of the two words "nuclear power" and "of" are calculated as 3.395039 and -29.905834 respectively. Through the correction coefficient, the keyword coefficients of common words can be better controlled to be negative, thus avoiding selecting interfering keywords as keywords.

[0079] S3. Match nuclear power documents and standards with the same or similar keywords, and calculate the document similarity coefficient.

[0080] Select the 3 keywords with the largest calculated keyword coefficients in S2, which can better reflect the theme information of this document, and calculate their vectors according to the nuclear industry semantic library.

[0081] The nuclear industry semantic library has established an association network of similar words, matches nuclear power documents and standards with the same or similar keywords, and calculates their keyword vectors according to the nuclear industry semantic library.

[0082] Use the GloVe model to count the co-occurrence frequency of word pairs in the semantic library, define a loss function to measure the difference between the predicted co-occurrence probability and the actual one, use gradient descent to iteratively adjust the keyword vectors, and finally generate low-dimensional dense vectors containing semantic information, making words with similar semantics closer in the vector space.

[0083] Calculate the keyword vectors of the calculated nuclear power documents and the keyword vectors of other standards, and calculate the document similarity coefficient according to Formula 3.

[0084]

[0085] Among them, represents the document similarity coefficient between the i-th nuclear power document and the j-th standard, A i,t represents the vector of the t-th keyword in the i-th nuclear power document, B j,k represents the vector of the k-th keyword in the j-th standard, ||A i,t -B j,k || represents the distance between vector A i,t and B j,k .

[0086] S4. Sort the document similarity coefficients and determine the similar standards according to the document similarity coefficients.

[0087] Arrange the file similarity coefficients calculated in S3, select the 8 standards with the smallest file similarity coefficients as the similarity criteria, and determine the association relationship between nuclear power files and the standards.

[0088] When calculating the knowledge element association relationship between the management procedures in nuclear power files and the standards, all the standards in the relevant field standard library will be selected as the files to be calculated (within about 300 for each field).

[0089] S5. Split the nuclear power files and the similar standards into sub-chapters and conduct a full-text comparison one by one to calculate the sub-chapter similarity coefficients.

[0090] Step S5 specifically includes:

[0091] S5.1. Split the nuclear power files and the similar standards determined in step S4 into sub-chapters according to the chapters;

[0092] S5.2. Perform text segmentation on each sub-chapter to be compared to obtain multiple sub-texts corresponding to each sub-chapter respectively;

[0093] S5.3. Obtain the multi-dimensional vectors corresponding to the sub-texts in the same chapter and combine them into a sub-text vector matrix;

[0094] S5.4. Calculate the cosine similarity between each sub-text vector matrix and obtain the chapter similarity matrix;

[0095] Among them, the cosine similarity calculation formula is:

[0096] cos(θ[i][j])=(A[i]·B[j]) / (||A[i]||*||B[j]||) Formula 4

[0097] Among them, cos(θ[i][j]) represents the cosine similarity between the sub-text vector matrix of the nuclear power file and the sub-text vector matrix of the standard, A[i] represents the i-th sub-text vector matrix in the nuclear power file, B[j] represents the j-th sub-text vector matrix in the standard, ||A[i]|| and ||B[j]|| respectively represent the norms of A[i] and B[j], and A[i]·B[j] represents the dot product between the matrices.

[0098] Construct a chapter similarity matrix X according to the calculated cosine similarity. The element in the i-th row and j-th column of the matrix is the cosine similarity between the i-th sub-text vector matrix in the nuclear power file and the j-th sub-text vector matrix in the standard, that is, cos(θ[i][j]).

[0099] S5.5. Calculate the similarity coefficient between the two sub-chapters to be compared according to the chapter similarity matrix.

[0100] Step S5.5 specifically includes:

[0101] S5.5.1. Calculate the similarity S between sub - chapter A of the nuclear power document and sub - chapter B of the standard AB 。

[0102]

[0103] Where m is the number of rows of the similarity matrix X, represents the value of the j - th column element with the largest value in the i - th row of the similarity matrix X (cosine similarity).

[0104] S5.5.2. Calculate the similarity S between sub - chapter B of the standard and sub - chapter A of the nuclear power document BA 。

[0105]

[0106] Where n is the number of columns of the similarity matrix X, represents the value of the i - th row element with the largest value in the j - th column of the similarity matrix X (cosine similarity).

[0107] S5.5.3. Calculate the sub - chapter similarity coefficient S between sub - chapter A of the nuclear power document and sub - chapter B of the standard A-B 。

[0108] S A-B =(2 * S AB * S BA ) / (S AB + S BA )

[0109] In a specific embodiment, A and B represent the nuclear power document and the standard, A1 and B1 represent the first chapter of the nuclear power document and the first chapter of the standard, and A1 - 1 and B1 - 1 represent the first sub - text of the first chapter of the nuclear power document and the first sub - text of the first chapter of the standard.

[0110] What S5.3 calculates is the sub - text vector matrix of A1 - 1 and B1 - 1

[0111] What S5.4 calculates is the sub - text similarity matrix composed of the cosine similarities between the sub - text vector matrices of A1 - 1 and B1 - 1. The element in the i - th row and j - th column of the matrix represents the cosine similarity between the two sub - texts (vector matrices) A1 - i and B1 - j.

[0112] S5.5 calculates the similarity of sub - chapters based on the sub - text similarity matrix.

[0113] It should be noted that the sub-text vector matrix only represents a sub-text and does not represent a chapter. In step S5, the vectors of multiple segments in the sub-text are first calculated and combined into a sub-text matrix, and then the similarity between the nuclear power document and a certain two sub-texts of the standard is calculated to form a chapter similarity matrix, and each element in it represents the similarity between the sub-texts.

[0114] S6. When the sub-chapter similarity coefficient reaches a preset value, it is considered that a certain chapter of the nuclear power document is associated with a certain chapter of the similarity standard.

[0115] In a specific embodiment, the preset value of the sub-chapter similarity coefficient in this step is 0.6. By calculating the historical nuclear power document and the standard through the above steps, it is found that setting the preset value to 0.6 can better distinguish whether there is an associated relationship between a certain chapter of the nuclear power document and a certain chapter of another standard.

[0116] A method for associating knowledge elements of nuclear power documents provided by the present invention is a standard association method based on a nuclear industry semantic library. When a user uses the method of the present invention to associate knowledge elements of nuclear power documents, the mining of keywords is more accurate, and relevant standards can be utilized more fully and comprehensively, which can effectively improve the association and utilization efficiency of nuclear power standards.

[0117] The embodiment of the present invention also provides a device for associating knowledge elements of nuclear power documents, including:

[0118] A construction module, configured to obtain nuclear power keywords and construct a nuclear power keyword map;

[0119] A first detection module, configured to detect the relationships of each keyword in the nuclear power document in the keyword map of the nuclear industry semantic library, and obtain a list of first entity relationship sets corresponding to the respective keywords;

[0120] A second detection module, configured to detect the relationships of each keyword in the standard in the keyword map of the nuclear industry semantic library, and obtain a list of second entity relationship sets corresponding to the respective keywords;

[0121] A first calculation module, configured to calculate the intersection of the entity relationship sets in the first entity relationship set list and the entity relationship sets in the second entity relationship set list to obtain multiple intersections;

[0122] A second calculation module, configured to calculate the similarity coefficient between each standard keyword and each nuclear power document keyword according to the multiple intersections;

[0123] An arrangement module, configured to arrange each nuclear power information according to the similarity coefficient to obtain the result of associating knowledge elements of nuclear power documents.

[0124] An embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, each step in the embodiment of the above-mentioned nuclear power document knowledge element association method can be implemented.

[0125] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor, each step in the embodiment of the above-mentioned nuclear power document knowledge element association method is implemented.

[0126] The present invention has been described in detail above in conjunction with the accompanying drawings and embodiments. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention. The content not described in detail in the present invention can all adopt the prior art.

Claims

1. A method for associating knowledge elements of nuclear power documents, characterized in that, The method includes: S1. Count the number of occurrences of each word in the full text of nuclear power documents and standards and sort them; S2. Calculate the keyword coefficients in the full text of nuclear power documents and standards according to the nuclear industry semantic library, and screen out keywords according to the keyword coefficients; S3. Match nuclear power documents and standards with the same or similar keywords, and calculate the document similarity coefficient; S4. Sort the document similarity coefficients, and determine the similarity criteria according to the document similarity coefficients; S5. Split the nuclear power documents and the similarity criteria into sub-chapters and compare them one by one throughout the text to calculate the sub-chapter similarity coefficient; S6. When the sub-chapter similarity coefficient reaches a preset value, it is considered that a certain chapter of the nuclear power document is associated with the similarity criteria.

2. The method for associating knowledge elements of nuclear power documents according to claim 1, wherein The S2 includes: S2.

1. Calculate the full-text occurrence frequency of each word in the full text of nuclear power documents and standards; The formula for calculating the full-text occurrence frequency of each word in the full text of nuclear power documents and standards is: C i,j = N i,j / M j Among them, C i,j represents the full text coefficient of the i-th word in the j-th nuclear power document or standard, N i,j represents the number of times the i-th word appears in the j-th nuclear power document or standard, M j represents the total number of words in the j-th nuclear power document or standard; S2.

2. Based on the nuclear industry semantic library, calculate the keyword coefficients, and screen out keywords according to the keyword coefficients; The formula for calculating the keyword coefficients is: K i,j = C i,j / Ln(N - H i / X) Among them, K i,j represents the keyword coefficient of the i-th word in the j-th nuclear power document or standard, and C i,j represents the full text coefficient of the i-th word in the j-th nuclear power document or standard. Ln represents the natural logarithm function, N is the correction coefficient, and H i indicates that there are H i words in the semantic library that contain the i-th keyword, and X represents the total number of words in the semantic library; N = (H i / X) T,max + 1 Among them, (H i / X) T,max is the maximum value of the occupancy ratio of each general term in the nuclear industry semantic library 3. A method for associating knowledge elements of nuclear power documents according to claim 1, characterized in that In S2, screening out keywords according to the keyword coefficients specifically means: taking the 3 keywords with the largest keyword coefficients as the screened keywords.

4. A nuclear power document knowledge element association method according to claim 1, characterized in that The formula for calculating the document similarity coefficient in S3 is: Among them, represents the similarity coefficient between the i-th nuclear power document and the j-th standard document, A i,t represents the vector of the t-th keyword in the i-th nuclear power document, B j,k represents the vector of the k-th keyword in the j-th standard, ||A i,t -B j,k || represents the distance between the vectors A i,t and B j,k therebetween.

5. A method for associating knowledge elements of nuclear power documents according to claim 1, characterized in that, In S4, determining the similarity criteria according to the document similarity coefficients specifically means: taking the 8 standards with the smallest document similarity coefficients as the similarity criteria.

6. The knowledge element association method for nuclear power documents according to claim 1, wherein The S5 includes: S5.

1. Split the nuclear power documents and the similarity criteria into sub-chapters according to the chapters; S5.

2. Perform text segmentation on each sub-chapter to be compared to obtain multiple sub-texts corresponding to each sub-chapter respectively; S5.

3. Obtain the multi-dimensional vectors corresponding to the sub-texts in the same chapter and combine them into a sub-text vector matrix; S5.

4. Calculate the cosine similarity between the sub-text vector matrices and obtain the chapter similarity matrix; S5.

5. Calculate the similarity coefficient of two sub-chapters to be compared according to the chapter similarity matrix.

7. A method for associating knowledge elements of nuclear power documents according to claim 6, characterized in that, The S5.5 includes: S5.5.

1. Calculate the similarity S between the sub-chapter A of nuclear power documents and the standard sub-chapter B AB where m is the number of rows of the similarity matrix X, represents the value of the j-th column element that is the largest in the i-th row of the similarity matrix X; S5.5.

2. Calculate the similarity S between the standard sub-chapter B and the nuclear power document sub-chapter A BA where n is the number of columns of the similarity matrix X, represents the value of the i-th row element that is the largest in the j-th column of the similarity matrix X; S5.5.

3. Calculate the sub-chapter similarity coefficient S between the nuclear power document sub-chapter A and the standard sub-chapter B A-B S A-B =(2 * S AB * S BA ) / (S AB + S BA )。 8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that When the computer program is executed by a processor, it implements the steps of the nuclear power document knowledge element association method according to any one of claims 1-7.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the steps of the nuclear power document knowledge element association method according to any one of claims 1-7.