Scientific and technological achievement duplicate checking method, device and equipment and readable storage medium
By constructing heterogeneous graphs and utilizing graph convolution models, the problem of similarity judgment between research results and reports in scientific research projects was solved, achieving efficient and accurate text similarity assessment and improving the automation level of scientific research management.
Patent Information
- Application Number
- CN202511043892.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies are insufficient to effectively and automatically determine the relevance of research results from scientific research projects to their corresponding projects, making it difficult to avoid problems such as duplicate project approvals and plagiarism in research reports.
Heterogeneous graphs are constructed using graph convolution models. Through deep preprocessing and feature learning, semantic relationships between words, paragraphs, and documents are fused. Graph convolution operations are used to capture complex semantic relationships between texts, and the similarity of vector codes of research results and reports is calculated.
This method improves the accuracy and efficiency of text similarity assessment, reduces reliance on manually labeled data, and provides an automated and efficient text similarity assessment method.
Smart Images

Figure CN120951981A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of text plagiarism detection technology, and more specifically, to a method, apparatus, device, and readable storage medium for detecting plagiarism in scientific and technological achievements. Background Technology
[0002] With the increasing number of science and technology project applications and the growing quantity of scientific and technological achievements, the manual review and duplicate verification model based on manual evaluation is no longer sufficient to meet the needs. Information technology tools such as plagiarism detection systems are required to automatically assess the similarity of research projects from different departments and at different stages of research activities, effectively preventing issues such as duplicate project approvals and plagiarism in research reports. For example, traditional string matching methods determine text similarity by comparing characters, words, or phrases. Current research focuses on duplicate project approvals during the project approval process and plagiarism in research reports during project completion. However, there is limited research on whether research results are outputs of research projects and their relevance to the corresponding research projects. This is a pressing issue that urgently needs to be addressed in the evaluation of research projects. Summary of the Invention
[0003] The purpose of this invention is to provide a method, apparatus, device, and readable storage medium for detecting plagiarism in scientific and technological achievements, thereby improving the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:
[0004] Firstly, this application provides a method for detecting plagiarism in scientific and technological achievements, including:
[0005] Obtain existing research findings and research reports to be checked for plagiarism;
[0006] Establish a first mapping relationship based on the paragraphs and words contained in the research findings, and establish a second mapping relationship based on the paragraphs and words contained in the research report;
[0007] Based on the first and second mapping relationships, an original heterogeneous graph is constructed with research reports, research results, paragraphs, and words as nodes;
[0008] Construct the initial adjacency matrix and initial feature matrix of all nodes based on the original heterogeneous graph;
[0009] A graph convolution model is constructed, and the initial adjacency matrix and the initial feature matrix are input into the graph convolution model for convolution training to obtain the learned feature matrix.
[0010] The first vector code of the research results and the second vector code of the research report are extracted from the feature matrix respectively. The similarity between the first vector code and the second vector code is calculated to form a plagiarism report.
[0011] Secondly, this application also provides a method and apparatus for detecting plagiarism in scientific and technological achievements, comprising:
[0012] Acquisition module: Acquires existing research results and research reports to be checked for plagiarism;
[0013] Relationship Establishment Module: Establish the first mapping relationship based on the paragraphs and words contained in the research findings, and establish the second mapping relationship based on the paragraphs and words contained in the research report;
[0014] Heterogeneous graph construction module: Based on the first and second mapping relationships, construct the original heterogeneous graph with research reports, research results, paragraphs, and words as nodes;
[0015] Matrix construction module: Constructs the initial adjacency matrix and initial feature matrix of all nodes based on the original heterogeneous graph;
[0016] Training module: Build a graph convolution model, input the initial adjacency matrix and initial feature matrix into the graph convolution model for convolution training, and obtain the learned feature matrix;
[0017] Calculation module: Extracts the first vector code of the research results and the second vector code of the research report from the feature matrix respectively, calculates the similarity between the first vector code and the second vector code to generate a plagiarism report.
[0018] Thirdly, this application also provides a method and equipment for detecting plagiarism in scientific and technological achievements, including:
[0019] Memory, used to store computer programs;
[0020] A processor is used to implement the steps of the scientific and technological achievement plagiarism detection method when executing the computer program.
[0021] Fourthly, this application also provides a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for detecting plagiarism based on scientific and technological achievements.
[0022] The beneficial effects of this invention are as follows:
[0023] This invention constructs a heterogeneous graph structure integrating words, paragraphs, and documents through deep preprocessing and feature learning of text, and effectively captures complex semantic relationships between texts using graph convolution operations. During model training, self-supervised learning using heterogeneous graph reconstruction and cross-entropy loss function fully integrates the features of words and paragraphs related to the text into the text's semantic features, improving the comprehensiveness of text feature representation and thus enhancing the model's ability to accurately assess text similarity. This invention not only reduces reliance on manually labeled data but also significantly improves the accuracy and efficiency of plagiarism detection, providing an automated and efficient text similarity assessment method for scientific research project management.
[0024] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of the scientific and technological achievement plagiarism detection method described in the embodiments of the present invention;
[0027] Figure 2 This is a schematic diagram of the structure of the scientific and technological achievement plagiarism detection method and device described in the embodiments of the present invention;
[0028] Figure 3 This is a schematic diagram of the equipment structure for the scientific and technological achievement plagiarism detection method described in this embodiment of the invention.
[0029] Marked in the image:
[0030] 800. Scientific and technological achievement plagiarism detection methods and equipment; 801. Processor; 802. Memory; 803. Multimedia components; 804. I / O interface; 805. Communication components. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0032] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0033] Example 1:
[0034] This embodiment provides a method for checking the plagiarism of scientific and technological achievements.
[0035] See Figure 1 The figure shows that this method includes:
[0036] S1. Obtaining existing research results T a And the research report to be checked for plagiarism T b ;
[0037] Based on the above embodiments, this method further includes:
[0038] S2. Establish a first mapping relationship based on the paragraphs and words contained in the research findings, and establish a second mapping relationship based on the paragraphs and words contained in the research report;
[0039] Specifically, step S2 includes:
[0040] S21. Divide the research findings into several first paragraphs P a ={P a1 ,P a2 ,…,P an}, where P a P represents the first paragraph set. an This represents the nth paragraph in the first paragraph set; it also involves splitting the research report into several second paragraphs P. b ={P b1 ,P b2 ,…,P bn′}, where Pb P represents the second paragraph set. bn′ This represents the n′th paragraph in the second paragraph set;
[0041] S22. Segment several first paragraphs and several second paragraphs respectively to obtain several words, and form a word library from these words;
[0042] Specifically, due to the presence of stop words and other noise in the paragraphs, a stop word dictionary D is set up to enhance the importance of core vocabulary in the text. The segmented words are then filtered based on the stop word dictionary D to obtain a vocabulary database W1 = {w1, w2, ..., w...} containing only core vocabulary. m}, where w m This represents the m-th word, and the word library W1 contains all the words from the first and second paragraphs.
[0043] S23. Select the words contained in the first paragraph from the vocabulary database and construct the first mapping relationship between the research results and the first paragraph and the words contained therein;
[0044] Specifically, based on the research findings and several first paragraphs, a mapping relationship between the research findings and the first paragraphs is constructed {(T a ,P a1 ),(T a ,P a2 ),…,(T a ,P an )};
[0045] Construct a mapping relationship between the first paragraph and the words it contains: {(P a1 ,w a11 ),(P a1 ,w a12 ),…,(P an ,w an1 ),(P an ,w an2 ...};
[0046] Therefore, the first mapping relationship is:
[0047] E tpw1 =
[0048] {(T a ,P a1 ),(T a ,P a2 ),…,(T a ,P an ),(P a1 ,w a11 ),(P a1 ,w a12 ...};
[0049] S24. Select the words contained in the second paragraph from the vocabulary database and construct a second mapping relationship between the research report and the second paragraph and the words contained therein;
[0050] Similarly, the second mapping relationship is:
[0051] E tpw2 =
[0052] {(T b ,P b1 ),(T b ,P b2 ),…,(T b ,P bn′ ),(P b1 ,w b11 ),(P b1 ,w b12 )...}.
[0053] Based on the above embodiments, this method further includes:
[0054] S3. Based on the first and second mapping relationships, construct the original heterogeneous graph G = (V, E) with research reports, research results, paragraphs, and words as nodes;
[0055] Among them, text node V t =(v ta ,v tb ), where v ta Indicates a research report node, v tb Indicates research findings nodes;
[0056] Paragraph node V p =(v pa1 ,v pa2 ,…,v pan ,v pb1 ,v pb2 ,…,v pbn′ ), where v pa1 ~v pan Both represent the nodes corresponding to the first paragraph, v pb1 ~v pbn′ Both represent the nodes corresponding to the second paragraph;
[0057] Word node V w =(v w1 ,v w2 ,…,v wm );
[0058] Edge E1 =
[0059] {(v ta ,vpa1 ),(v tb ,v pb1 ),(v pa1 ,v wa11 ),(v pa1 ,v wa12 ),…,(v pbn′ ,v wbn′m )};where, v wa11 v wa12 and v wbn′m For word node V w The selected nodes.
[0060] This embodiment uses the original heterogeneous graph to integrate the symbolic features of different types of nodes into the features of the research report.
[0061] Based on the above embodiments, this method further includes:
[0062] S4. Construct the initial adjacency matrix and initial feature matrix of all nodes based on the original heterogeneous graph;
[0063] Specifically, step S4 includes:
[0064] S41. Obtain the node set consisting of text nodes, paragraph nodes, and word nodes in the original heterogeneous graph, wherein the text nodes are research report nodes and research result nodes;
[0065] S42. Create a multidimensional matrix A = N × N based on the size of the node set, where N
[0066] This indicates the number of nodes in the node set;
[0067] S43. Obtain the edge set in the original heterogeneous graph, and update the multidimensional matrix according to each edge in the edge set to obtain the initial adjacency matrix A;
[0068] Specifically, if there is an edge between nodes i and j in the node set, the position of node i in the i-th row and j-th column in the initial adjacency matrix A is set to 1; otherwise, it is set to 0.
[0069] Specifically, step S4 further includes:
[0070] S44. Obtain the multi-dimensional vector corresponding to each word node from the Chinese Word Vectors pre-trained word vector set to construct the word vector corresponding to the word node:
[0071]
[0072] in, Word node v w1 Word vectors, express multidimensional vectors, Word node v wm Word vectors, express A multidimensional vector, D w Represents a set of word vectors;
[0073] S45. Obtain the word nodes contained in each paragraph node, and construct the paragraph vector of the paragraph node from the word vectors corresponding to the word nodes:
[0074]
[0075] In the formula, D p Represents a set of paragraph vectors. Represents paragraph node v pa1 paragraph vector, Represents paragraph node v pbn′ paragraph vector, and This indicates the word vector set D w The word vectors selected from the text.
[0076] S46. Obtain the paragraph nodes contained in the text node, and construct the text vector D of the text node from the paragraph vectors corresponding to the paragraph nodes. t :
[0077]
[0078] In the formula, Represents text node v ta The text vector, Represents text node v tb The text vector.
[0079] S47. Arrange the word vectors, paragraph vectors, and text vectors according to the order of the nodes to generate an initial feature matrix X = N × D, where D represents the dimension of each node vector.
[0080] Based on the above embodiments, this method further includes:
[0081] S5. Construct a graph convolution model, input the initial adjacency matrix and the initial feature matrix into the graph convolution model for convolution training, and obtain the learned feature matrix;
[0082] Specifically, step S5 includes:
[0083] S51. Build a graph convolution model and initialize the parameter matrices W0 and W1 of the graph convolution model. The number of rows of the parameter matrix W0 is equal to the number of columns of the feature matrix X, and the number of rows of W1 is equal to the number of columns of the feature matrix Z1. The number of columns of W0 and W1 is a custom number.
[0084] S52. Input the initial adjacency matrix and initial feature matrix into the graph convolution model for convolution training to generate the updated feature matrix Z2:
[0085]
[0086] In the formula, Z1 represents the feature matrix after one convolution, Z2 represents the feature matrix after two convolutions, i.e. the updated feature matrix, and ReLU is the activation function.
[0087] S53. Reconstruct the original heterogeneous graph using the updated feature matrix to generate a reconstructed adjacency matrix, and calculate the difference between the original adjacency matrix and the reconstructed adjacency matrix to obtain the loss value.
[0088] Specifically, step S53 includes:
[0089] S531. Calculate the product of the updated feature matrix and its transpose, and input the product into the activation function to generate the reconstructed adjacency matrix:
[0090]
[0091] In the formula, σ represents the activation function. Z2 represents the reconstructed adjacency matrix. T This represents the transpose of Z2.
[0092] S532. Input each element from the reconstructed adjacency matrix and the original adjacency matrix into the cross-entropy loss function, and calculate the difference between each element in turn:
[0093]
[0094] l i y represents the difference between the reconstructed adjacency matrix and the original adjacency matrix for the i-th element. i This represents the value of the i-th element in the initial adjacency matrix A, which can be either 0 or 1. This indicates the reconstruction of the adjacency matrix. Representing the adjacency matrix The value of the i-th element is [0, 1].
[0095] S533. The loss values for the original adjacency matrix and the reconstructed adjacency matrix are formed by the sum of the differences of all elements.
[0096]
[0097] In the formula, Loss represents the loss value, and N represents the number of elements.
[0098] S54. Determine whether the loss value is less than a preset threshold. In this embodiment, the preset threshold is set according to the actual situation, and this embodiment does not limit it.
[0099] S55. If not, update the graph parameter matrices W′0 and W′1 according to the loss value to obtain the updated graph convolution model:
[0100] S56. Input the updated feature matrix into the updated graph convolutional model for iterative training until the loss value between the original adjacency matrix and the reconstructed adjacency matrix is less than a preset threshold, and obtain the learned feature matrix, which includes all node features.
[0101] Specifically, replace X with Z2 and perform convolution on the feature matrix:
[0102]
[0103] In the formula, Z′1 and Z′2 represent the feature matrices after the second update.
[0104] This embodiment uses a graph self-supervised learning method to learn text features. The learning process only utilizes the existing graph structure and does not require the construction of similar and dissimilar texts, effectively avoiding the time-consuming and laborious problem of constructing positive and negative samples.
[0105] Based on the above embodiments, this method further includes:
[0106] S6. Extract the first vector code of the research results and the second vector code of the research report from the learned feature matrix, respectively, and calculate the similarity between the first vector code and the second vector code to form a plagiarism report;
[0107] In this embodiment, node features of the research results are extracted from the learned feature matrix to form the first vector code Z of the research report results. a =[h za1 ,h za2 ,…,h zak ], where h zak The k-th vector encoding is derived from the node features of the research report extracted from the learned feature matrix, forming the second vector encoding Z of the research report text. b =[h zb1 ,h zb2 ,…,h zbk′ ], h zbk′ This represents the encoding of the k′-th vector;
[0108] Calculate the cosine similarity (Cosine(Z)) between the first and second vector codes. a Z b ):
[0109]
[0110] In the formula, ‖·‖ represents the modulus;
[0111] In this embodiment, the cosine similarity (Z) a Z b The higher the value of ), the higher the relevance between the research report and the research results.
[0112] Specifically, the generated plagiarism report includes information about the research report to be checked and the plagiarism check results. The information about the research report to be checked includes the research report title, author, date, and main research content. The plagiarism check results include the research result name, main research content, and similarity.
[0113] Example 2:
[0114] like Figure 2 As shown, this embodiment provides a method and apparatus for detecting plagiarism in scientific and technological achievements. The apparatus includes:
[0115] Acquisition module: Acquires existing research results and research reports to be checked for plagiarism;
[0116] Relationship Establishment Module: Establish the first mapping relationship based on the paragraphs and words contained in the research findings, and establish the second mapping relationship based on the paragraphs and words contained in the research report;
[0117] Heterogeneous graph construction module: Based on the first and second mapping relationships, construct the original heterogeneous graph with research reports, research results, paragraphs, and words as nodes;
[0118] Matrix construction module: Constructs the initial adjacency matrix and initial feature matrix of all nodes based on the original heterogeneous graph;
[0119] Training module: Build a graph convolution model, input the initial adjacency matrix and initial feature matrix into the graph convolution model for convolution training, and obtain the learned feature matrix;
[0120] Calculation module: Extracts the first vector code of the research results and the second vector code of the research report from the feature matrix respectively, calculates the similarity between the first vector code and the second vector code to generate a plagiarism report.
[0121] Based on the above embodiments, the relationship establishment module includes:
[0122] Paragraph division: The research findings are divided into several first paragraphs, and the research report is divided into several second paragraphs;
[0123] Word segmentation unit: It performs word segmentation on several first paragraphs and several second paragraphs respectively to obtain several words, which constitute a word library;
[0124] First screening unit: Select the words contained in the first paragraph from the vocabulary database and construct the first mapping relationship between the research results and the first paragraph and the words contained therein;
[0125] The second filtering unit: filters out the words contained in the second paragraph from the vocabulary database, and constructs a second mapping relationship between the research report and the second paragraph and the words contained therein.
[0126] Based on the above embodiments, the matrix construction module includes:
[0127] First acquisition unit: Acquire the node set consisting of text nodes, paragraph nodes and word nodes in the original heterogeneous graph, wherein the text nodes are research report nodes and research result nodes;
[0128] Matrix creation unit: Creates a multidimensional matrix based on the size of the node set;
[0129] Update unit: Obtain the edge set in the original heterogeneous graph, update the multidimensional matrix according to each edge in the edge set, and obtain the initial adjacency matrix.
[0130] Based on the above embodiments, the matrix construction module further includes:
[0131] Word vector construction unit: Obtain the multi-dimensional vector corresponding to each word node from the Chinese pre-trained word vector set to construct the word vector corresponding to the word node;
[0132] Paragraph vector construction unit: Obtain the word nodes contained in each paragraph node, and construct the paragraph vector of the paragraph node from the word vectors corresponding to the word nodes;
[0133] Text vector construction unit: Obtain the paragraph nodes contained in the text node, and construct the text vector of the text node from the paragraph vector corresponding to the paragraph node;
[0134] Sorting Unit: Arranges the word vectors, paragraph vectors, and text vectors according to the order of the nodes to generate an initial feature matrix.
[0135] Based on the above embodiments, the training module for building the graph convolutional model includes:
[0136] Model building unit: Builds the graph convolution model and initializes the parameter matrix of the graph convolution model;
[0137] Convolutional training unit: Input the initial adjacency matrix and initial feature matrix into the graph convolutional model for convolutional training to generate the updated feature matrix;
[0138] Reconstruction Unit: The original heterogeneous graph is reconstructed using the updated feature matrix to generate a reconstructed adjacency matrix, and the difference between the original adjacency matrix and the reconstructed adjacency matrix is calculated to obtain the loss value.
[0139] Judgment unit: Determines whether the loss value is less than a preset threshold.
[0140] If not, update the graph parameter matrix based on the loss value to obtain the updated graph convolution model;
[0141] Iterative Unit: The updated feature matrix is input into the updated graph convolutional model for iterative training until the loss value between the original adjacency matrix and the reconstructed adjacency matrix is less than a preset threshold, thus obtaining the learned feature matrix.
[0142] Based on the above embodiments, the reconstruction unit includes:
[0143] First computational unit: Calculates the dot product of the updated feature matrix and inputs the dot product into the activation function to generate the reconstructed adjacency matrix;
[0144] The second calculation unit inputs each element from the reconstructed adjacency matrix and the original adjacency matrix into the cross-entropy loss function and calculates the difference between each element in turn.
[0145] The third computational unit: the loss value of the original adjacency matrix and the reconstructed adjacency matrix is formed by the sum of the differences of all elements.
[0146] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.
[0147] Example 3:
[0148] Corresponding to the above method embodiments, this embodiment also provides a method and device for checking scientific and technological achievements. The method and device for checking scientific and technological achievements described below can be referred to in correspondence with the method for checking scientific and technological achievements described above.
[0149] Figure 3 This is a block diagram illustrating a method and apparatus 800 for detecting plagiarism in scientific and technological achievements, according to an exemplary embodiment. Figure 3 As shown, the plagiarism detection method device 800 may include: a processor 801 and a memory 802. The plagiarism detection method device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0150] The processor 801 controls the overall operation of the plagiarism detection device 800 to complete all or part of the steps in the aforementioned plagiarism detection method. The memory 802 stores various types of data to support the operation of the plagiarism detection device 800. This data may include, for example, instructions for any application or method operating on the device 800, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the plagiarism detection device 800 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, and an NFC module.
[0151] In an exemplary embodiment, the technology achievement plagiarism detection method device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the aforementioned technology achievement plagiarism detection method.
[0152] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the above-described method for detecting plagiarism in scientific and technological achievements. For example, the computer-readable storage medium may be the memory 802 including the program instructions, which may be executed by the processor 801 of the apparatus 800 for detecting plagiarism in scientific and technological achievements to complete the above-described method for detecting plagiarism in scientific and technological achievements.
[0153] Example 4:
[0154] Corresponding to the above method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the scientific and technological achievement plagiarism detection method described above.
[0155] A readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the scientific and technological achievement plagiarism detection method described in the above method embodiments.
[0156] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.
[0157] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0158] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can explore the technical scope disclosed in the present invention.
[0159] Any variations or substitutions that can be easily conceived within the scope of this invention should be included within the protection scope of this invention.
[0160] Therefore, the scope of protection of this invention should be determined by the scope of the claims.
Claims
1. A method for detecting plagiarism in scientific and technological achievements, characterized in that, include: Obtain existing research findings and research reports to be checked for plagiarism; Establish a first mapping relationship based on the paragraphs and words contained in the research findings, and establish a second mapping relationship based on the paragraphs and words contained in the research report; Based on the first and second mapping relationships, an original heterogeneous graph is constructed with research reports, research results, paragraphs, and words as nodes; Construct the initial adjacency matrix and initial feature matrix of all nodes based on the original heterogeneous graph; A graph convolution model is constructed, and the initial adjacency matrix and the initial feature matrix are input into the graph convolution model for convolution training to obtain the learned feature matrix. The first vector code of the research results and the second vector code of the research report are extracted from the feature matrix respectively. The similarity between the first vector code and the second vector code is calculated to form a plagiarism report.
2. The method for detecting plagiarism of scientific and technological achievements according to claim 1, characterized in that... A first mapping relationship is established based on the paragraphs and words contained in the research findings, and a second mapping relationship is established based on the paragraphs and words contained in the research report, including: The research findings are divided into several first paragraphs, and the research report is divided into several second paragraphs; Several first paragraphs and several second paragraphs were segmented into words to obtain a number of words, which were then used to form a word database. The words contained in the first paragraph are selected from the vocabulary database, and the first mapping relationship between the research results and the first paragraph and the words contained therein is constructed. The words contained in the second paragraph are selected from the vocabulary database, and a second mapping relationship is constructed between the research report and the second paragraph and its contained words.
3. The method for detecting plagiarism of scientific and technological achievements according to claim 1, characterized in that... Construct an initial adjacency matrix for all nodes based on the original heterogeneous graph, including: Obtain the node set consisting of text nodes, paragraph nodes, and word nodes in the original heterogeneous graph, wherein the text nodes are research report nodes and research result nodes; Create a multidimensional matrix based on the size of the node set; Obtain the edge set in the original heterogeneous graph, and update the multidimensional matrix according to each edge in the edge set to obtain the initial adjacency matrix.
4. The method for detecting plagiarism of scientific and technological achievements according to claim 1, characterized in that... Construct a graph convolutional model, input the initial adjacency matrix and initial feature matrix into the graph convolutional model for convolution training, and obtain the learned feature matrix, including: Build a graph convolution model and initialize its parameter matrix; The initial adjacency matrix and initial feature matrix are input into the graph convolution model for convolution training to generate the updated feature matrix; The original heterogeneous graph is reconstructed using the updated feature matrix to generate a reconstructed adjacency matrix, and the difference between the original adjacency matrix and the reconstructed adjacency matrix is calculated to obtain the loss value. Determine whether the loss value is less than a preset threshold: If not, update the graph parameter matrix based on the loss value to obtain the updated graph convolution model; The updated feature matrix is input into the updated graph convolutional model for iterative training until the loss between the original adjacency matrix and the reconstructed adjacency matrix is less than a preset threshold, thus obtaining the learned feature matrix.
5. A method and device for detecting plagiarism in scientific and technological achievements, characterized in that, include: Acquisition module: Acquires existing research results and research reports to be checked for plagiarism; Relationship Establishment Module: Establish the first mapping relationship based on the paragraphs and words contained in the research findings, and establish the second mapping relationship based on the paragraphs and words contained in the research report; Heterogeneous graph construction module: Based on the first and second mapping relationships, construct the original heterogeneous graph with research reports, research results, paragraphs, and words as nodes; Matrix construction module: Constructs the initial adjacency matrix and initial feature matrix of all nodes based on the original heterogeneous graph; Training module: Build a graph convolution model, input the initial adjacency matrix and initial feature matrix into the graph convolution model for convolution training, and obtain the learned feature matrix; Calculation module: Extracts the first vector code of the research results and the second vector code of the research report from the feature matrix respectively, calculates the similarity between the first vector code and the second vector code to generate a plagiarism report.
6. The method and apparatus for detecting plagiarism of scientific and technological achievements according to claim 5, characterized in that, The relationship establishment module includes: Paragraph division: The research findings are divided into several first paragraphs, and the research report is divided into several second paragraphs; Word segmentation unit: It performs word segmentation on several first paragraphs and several second paragraphs respectively to obtain several words, which constitute a word library; First screening unit: Select the words contained in the first paragraph from the vocabulary database and construct the first mapping relationship between the research results and the first paragraph and the words contained therein; The second filtering unit: filters out the words contained in the second paragraph from the vocabulary database, and constructs a second mapping relationship between the research report and the second paragraph and the words contained therein.
7. The method and apparatus for detecting plagiarism in scientific and technological achievements according to claim 5, characterized in that, The matrix construction module includes: First acquisition unit: Acquire the node set consisting of text nodes, paragraph nodes and word nodes in the original heterogeneous graph, wherein the text nodes are research report nodes and research result nodes; Matrix creation unit: Creates a multidimensional matrix based on the size of the node set; Update unit: Obtain the edge set in the original heterogeneous graph, update the multidimensional matrix according to each edge in the edge set, and obtain the initial adjacency matrix.
8. The method and apparatus for detecting plagiarism in scientific and technological achievements according to claim 5, characterized in that, The training module includes: Model building unit: Builds the graph convolution model and initializes the parameter matrix of the graph convolution model; Convolutional training unit: Input the initial adjacency matrix and initial feature matrix into the graph convolutional model for convolutional training to generate the updated feature matrix; Reconstruction Unit: The original heterogeneous graph is reconstructed using the updated feature matrix to generate a reconstructed adjacency matrix, and the difference between the original adjacency matrix and the reconstructed adjacency matrix is calculated to obtain the loss value. Judgment unit: Determines whether the loss value is less than a preset threshold. If not, update the graph parameter matrix based on the loss value to obtain the updated graph convolution model; Iterative Unit: The updated feature matrix is input into the updated graph convolutional model for iterative training until the loss value between the original adjacency matrix and the reconstructed adjacency matrix is less than a preset threshold, thus obtaining the learned feature matrix.
9. A method and device for detecting plagiarism in scientific and technological achievements, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the scientific and technological achievement plagiarism detection method as described in any one of claims 1 to 4 when executing the computer program.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the scientific and technological achievement plagiarism detection method as described in any one of claims 1 to 4.