A scientific report achievement duplicate checking method based on neural network feature recognition

By using neural network feature recognition methods, the problems of low efficiency and low accuracy in plagiarism detection in science and technology project applications have been solved, realizing an efficient and reliable plagiarism detection system, enhancing the ability to identify synonyms and improving text security.

CN117033550BActive Publication Date: 2025-12-12THE QUARTERMASTER RES INST OF THE GENERAL LOGISTICS DEPT OF THE CPLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310968046.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-12-12
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

In existing technologies, the methods for checking the plagiarism of scientific and technological projects and achievements are inefficient and have low accuracy, and are prone to misjudgment and omission, especially the keyword set comparison method, which is easily circumvented.

Method used

A method for detecting plagiarism in scientific and technological reports based on neural network feature recognition is adopted, which includes report input, information data optimization and extraction, encoding conversion, discrimination and comparison, discrimination compensation and result output. Through word segmentation, sentence segmentation, extraction of keywords and synonyms compensation, and setting thresholds for accurate identification.

Benefits of technology

This improves the accuracy and reliability of the plagiarism detection system, effectively avoids the use of synonyms for circumvention, enhances the efficiency and judgment of the plagiarism detection system, and ensures the confidentiality and security of the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033550B_ABST
    Figure CN117033550B_ABST
Patent Text Reader

Abstract

The application discloses a scientific and technological report achievement duplication checking method based on neural network feature recognition, and belongs to the technical field of report duplication checking. In the application, the searching and comparing objects of the report to be checked are expanded through an external network compensation unit, and the near-synonymous words and short sentences of the extracted words and sentences are brought into the objects to be checked by using a word / sentence near-synonymous compensation unit, so that the report to be checked can effectively avoid duplication checking by using near-synonymous words and sentences to evade the duplication checking, and the coincidence degree and the similarity are reduced, the accuracy and the reliability of the duplication checking system are greatly enhanced, the recognizable length of the character code is set by each threshold setting unit, and the setting values are different, the result of the report is further optimized under the conditions of the accurate and multiple times of recognition, and the coding content is periodically transformed and updated by means of a code changing and transforming unit, so that one-way data transmission is constructed, and the extracted content is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of report duplication checking, in particular to a scientific report achievement duplication checking method based on neural network feature recognition. BACKGROUND

[0002] At present, a large number of students and scientific researchers in China will report various types of scientific projects and achievements at the national, provincial and local levels every year. In order to avoid the phenomenon of waste of scientific research funds caused by repeated reporting of scientific projects and achievements, in the reporting and auditing process, two kinds of duplication checking methods, manual review and simple comparison of the key word set in the scientific achievement report with the project database, are used to filter out the repeated reporting projects.

[0003] However, although the two filtering methods can reduce the repeated reporting of scientific projects and achievements to a certain extent, they still have the disadvantages of low efficiency and easy to make wrong and missed judgments, especially the simple comparison of the key word set in the project report. Once the key words in the text are changed or slightly changed, or the relative words and sentences are used to evade, the duplication checking system can pass, so the accuracy is low and the reliability is poor. Therefore, a scientific report achievement duplication checking method based on neural network feature recognition is proposed. SUMMARY

[0004] The purpose of the present application is to provide a scientific report achievement duplication checking method based on neural network feature recognition to solve the problem of low accuracy and poor reliability of the comparison of the key word set in the project report in the background art.

[0005] To achieve the above purpose, the present application provides the following technical scheme: a scientific report achievement duplication checking method based on neural network feature recognition, comprising the following steps:

[0006] S1, report input and screening: the report to be checked is input and screened through the input and screening module for corresponding text input and information data screening;

[0007] S2, information data optimization extraction: the information data screened out in S1 is further optimized and extracted by the target data extraction module;

[0008] S3, encoding conversion: the information data optimized and extracted in S2 is converted into character encoding by the encoding conversion module, and the data format is formatted;

[0009] S4, discrimination and comparison: the converted encoding in S3 is decrypted and divided by the discrimination module to make it within the set threshold, and then input to the database for comparison and comparison by the comparison module;

[0010] S5, discrimination compensation: after S4 comparison, through the database to find out the coincidence and similar comparison text directly output report results, and in the database is not found out the coincidence and similar comparison text, the code will jump to the discrimination compensation module for secondary data compensation, and then the optimal report results are obtained;

[0011] S6, result output: the results after S4 and S5 operation are output through the report result module.

[0012] As a further preferred technical solution of the present application: in S1, the input screening module comprises a text input unit, an information extraction unit and a data screening unit;

[0013] The text input unit is used for inputting the report to be checked.

[0014] The information extraction unit is used for extracting the total number of words, application field and text box structure in the report.

[0015] The data screening unit is used for screening the text directory, text chapter and text paragraph.

[0016] As a further preferred technical solution of the present application: in S2, the target data extraction module is used for word segmentation and sentence segmentation of the extracted total number of words, application field and text box structure, and the screened text directory, text chapter and text paragraph, and extracts key words, professional words, special words and multiple repetitive words.

[0017] As a further preferred technical solution of the present application: in S3, the code conversion module comprises a code change conversion unit, a character code conversion unit and a data formatting unit;

[0018] The code change conversion unit is used for periodic transformation and update of the code content.

[0019] The character code conversion unit is used for converting the text into character code.

[0020] The data formatting unit is used for partitioning and initializing the converted character code.

[0021] As a further preferred technical solution of the present application: in S4, the discrimination module comprises a data decryption unit, a first threshold setting unit and a database unit;

[0022] The data decryption unit is used for analyzing the converted character code.

[0023] The first threshold setting unit is used for setting the recognizable length of the character code, and the setting value is (0, 100).

[0024] The database unit is used for providing the report to be checked for searching and comparison.

[0025] A comparison module is configured to compare the coincidence and similarity of the report to be checked with the texts in the database unit.

[0026] As a further preferred embodiment of the present application, the specific operation method of the compensation module is determined in S5, including the following steps:

[0027] A1, the parameter BM jumps to the target network through T;

[0028] A2, the target network Wz (W1, W2, W3... Wz+1);

[0029] A3, extract the parameters in the target network Wz and input them into the DT for training;

[0030] A4, analyze the parameter BM through Key and construct the target network YG;

[0031] A5, the target network YG includes double experience pool groups Yz (Y1, Y2, Y3... Yz+1) and Gz (G1, G2, G3... Gz+1);

[0032] A6, extract the parameters in Yz and Gz and extract the parameters in the target network Wz and input them into the DT for training;

[0033] A7, set the training threshold (0, 25) for the target network;

[0034] A8, yes or no, meet the duplication checking result, when the duplication checking result is met, directly enter the next stage of operation, and when the duplication checking result is not met, according to the input path, respectively reverse back to A3 and A6 to re-run;

[0035] A9, save the network and end the training;

[0036] Wherein, the parameter BM is the code, T is the jump platform, Wz is the external link experience pool, DT is the comparison module, Key is the external port, YG is the word and sentence synonym experience pool, Yz is the word experience pool, and Gz is the sentence experience pool;

[0037] The compensation module includes an external network compensation unit, a word and sentence synonym compensation unit, and a second threshold setting unit;

[0038] The external network compensation unit expands the objects of the report to be checked through external links;

[0039] The word and sentence synonym compensation unit includes the near-synonymous words and short sentences of the extracted words and sentences into the objects to be checked;

[0040] The second threshold setting unit is used to set the recognizable length of the character code, and the set value is (0, 50).

[0041] As a further preferred of the technical solution: in S6, the report result module includes a determination result unit, a comprehensive evaluation unit and a similarity evaluation report;

[0042] The determination result unit is configured to output the determination result of the report to be checked for duplication.

[0043] The comprehensive evaluation unit is configured to output the determination point and the determination reason of the determination result of the report to be checked for duplication.

[0044] The similarity evaluation report is configured to output the compared text that coincides with and is similar to the report to be checked for duplication.

[0045] Compared with the prior art, the present application has the following advantages:

[0046] 1. In the present application, the external network compensation unit expands the search and comparison objects of the report to be checked for duplication, and the word and sentence near-synonym compensation unit optimizes the near-synonym words and short sentences of the extracted words and sentences into the objects to be checked for duplication, thereby effectively avoiding the report to be checked for duplication from evading duplication check by using near-synonym words and phrases to reduce its coincidence degree and similarity, greatly enhancing the accuracy and reliability of the duplication check system, and the character code recognizable length is set by the threshold setting unit, and different set values are used to further optimize the report results under the conditions of precision and multiple recognitions.

[0047] 2. The information extraction unit and the data screening unit of the present application extract the total number of characters, the application field and the text box structure in the text, and screen out the text directory, the text chapter and the text paragraph, and then the target data extraction module performs word segmentation and sentence segmentation on the extracted and screened contents, and extracts key words, professional words, special words and multiple repetitive words, thereby effectively improving the efficiency and judgment effect of the duplication check system.

[0048] 3. In the present application, the text is converted into character codes by the code conversion module, thereby greatly reducing the volume occupied by character bytes, and the code changing and converting unit can periodically change and update the coded content, thereby constructing one-way data transmission, which guarantees the confidentiality and security of the extracted content, the text in the database and the external link reference text. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 The figure is a schematic diagram of the architecture of the scientific report achievement duplication checking method based on neural network feature recognition of the present application.

[0050] Figure 2 The figure is a schematic diagram of the running process of the scientific report achievement duplication checking method based on neural network feature recognition of the present application.

[0051] Figure 3 An algorithm flow diagram of a discrimination compensation module in a scientific report achievement duplication checking method based on neural network feature recognition of the present application is shown in the figure.

[0052] Figure 4 An architecture diagram of an entry screening module in a scientific report achievement duplication checking method based on neural network feature recognition of the present application is shown in the figure.

[0053] Figure 5 An architecture diagram of an encoding conversion module in a scientific report achievement duplication checking method based on neural network feature recognition of the present application is shown in the figure.

[0054] Figure 6 An architecture diagram of a determination module in a scientific report achievement duplication checking method based on neural network feature recognition of the present application is shown in the figure.

[0055] Figure 7 An architecture diagram of a discrimination compensation module in a scientific report achievement duplication checking method based on neural network feature recognition of the present application is shown in the figure.

[0056] Figure 8 A report result module architecture diagram in a scientific report achievement duplication checking method based on neural network feature recognition of the present application is shown in the figure. DETAILED DESCRIPTION

[0057] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0058] EMBODIMENT

[0059] Please refer to Figures 1-8 The present application provides a technical solution: a scientific report achievement duplication checking method based on neural network feature recognition, comprising the following steps:

[0060] S1, report entry and screening: the report to be checked for duplication is subjected to corresponding text entry and information data screening work through an entry screening module;

[0061] S2, information data optimization extraction: the information data screened out in S1 is further optimized and extracted through a target data extraction module;

[0062] S3, encoding conversion: the information data optimized and extracted in S2 is converted into character codes by an encoding conversion module, and data formatting is performed;

[0063] S4, discrimination and comparison: the encoding converted by S3 is decrypted and divided by the discrimination module to be within the set threshold, and then input to the database for comparison by the comparison module;

[0064] S5, discrimination compensation: after comparison by S4, the database finds out the coincident and similar comparison text and directly outputs the report result, and the database does not find out the coincident and similar comparison text, and the encoding jumps to the discrimination compensation module for secondary data compensation, and then the optimal report result is obtained;

[0065] S6, result output: the results after S4 and S5 run are output by the report result module.

[0066] In this embodiment, specifically: in S1, the input screening module includes a text input unit, an information extraction unit and a data screening unit;

[0067] The text input unit is configured to receive the input report to be checked for duplication.

[0068] The information extraction unit is configured to extract the total number of words, application field and text box structure in the report.

[0069] The data screening unit is configured to screen the text directory, text chapter and text paragraph.

[0070] In this embodiment, specifically: in S2, the target data extraction module is configured to segment and sentence the extracted total number of words, application field and text box structure, and the screened text directory, text chapter and text paragraph, and extract key words, professional words, special words and multiple repetitive words.

[0071] In this embodiment, specifically: in S3, the encoding conversion module includes a code change conversion unit, a character code conversion unit and a data formatting unit;

[0072] The code change conversion unit is configured to periodically transform and update the encoding content, thereby constructing one-way data transmission to ensure the confidentiality and security of the report to be checked for duplication, the text in the database and the external link reference text.

[0073] The character code conversion unit is configured to convert the text into character encoding.

[0074] The data formatting unit is configured to partition and initialize the converted character encoding.

[0075] In this embodiment, specifically: in S4, the discrimination module includes a data decryption unit, a first threshold setting unit and a database unit;

[0076] The data decryption unit is configured to analyze the converted character encoding.

[0077] The first threshold setting unit is configured to set the recognizable length of the character code, and the set value is (0, 100);

[0078] The database unit is configured to provide the searched report with searching and comparing objects;

[0079] The comparison module is configured to compare the coincidence and similarity of the searched report and the text in the database unit.

[0080] In the embodiment, specifically, in S5, the specific operation method of the compensation module includes the following steps:

[0081] A1, the parameter BM jumps to the target network through T;

[0082] A2, the target network Wz (W1, W2, W3,..., Wz+1);

[0083] A3, extract the parameters in the target network Wz and input them into the DT for training;

[0084] A4, analyze the parameter BM through the Key and construct the target network YG;

[0085] A5, the target network YG includes double experience pool groups Yz (Y1, Y2, Y3,..., Yz+1) and Gz (G1, G2, G3,..., Gz+1);

[0086] A6, extract the parameters in Yz and Gz respectively and extract the parameters in the target network Wz and input them into the DT for training;

[0087] A7, set the training threshold (0, 25) to reset the target network;

[0088] A8, yes or no to meet the duplication checking result, when the duplication checking result is met, directly enter the next stage of operation, and when the duplication checking result is not met, according to the input path, respectively reverse back to A3 and A6 to re-run;

[0089] A9, save the network and end the training;

[0090] Wherein, the parameter BM is the code, T is the jump platform, Wz is the external link experience pool, DT is the comparison module, Key is the external port, YG is the word and sentence synonym experience pool, Yz is the word experience pool, and Gz is the sentence experience pool;

[0091] The discrimination compensation module includes an external network compensation unit, a word and sentence synonym compensation unit, and a second threshold setting unit;

[0092] The external link compensation unit expands the object of searching and comparing the duplicate report through external link, and the external link reference text includes log, anthology, monograph and literature, etc.

[0093] The phrase and sentence near-synonymous compensation unit optimizes the near-synonymous phrase and short sentence of the extracted phrase and sentence, and includes the near-synonymous phrase and short sentence in the object of searching and comparing, so as to effectively avoid the duplicate report from evading the duplicate search by using near-synonymous words and phrases, and reduce the coincidence degree and similarity.

[0094] The second threshold setting unit is used for setting the recognizable length of the character code, and the setting value is (0, 50).

[0095] In the embodiment, specifically, in S6, the report result module includes a determination result unit, a comprehensive evaluation unit and a similarity evaluation report.

[0096] The determination result unit is used for outputting the determination result of the duplicate report.

[0097] The comprehensive evaluation unit is used for outputting the determination point and the determination reason of the determination result of the duplicate report.

[0098] The similarity evaluation report is used for outputting the compared text which coincides and is similar to the duplicate report.

[0099] Working principle or structural principle: when in use, the staff first inputs the duplicate report through the text input unit, and then extracts the total number of words, application field and text frame structure in the text through the information extraction unit and data screening unit, and screens out the text directory, text chapter and text paragraph. Then, the target data extraction module is used to divide the extracted total number of words, application field and text frame structure, and the screened text directory, text chapter and text paragraph into words and sentences, and extract the key words, professional words, special words and multiple repetitive words.

[0100] The optimized extracted content is converted into character code through the character code conversion unit, and the converted character code is partitioned and initialized through the data formatting unit. The code changing and converting unit is used to periodically change and update the content of the code, so as to build one-way data transmission, thereby ensuring the confidentiality and security of the duplicate report, the text in the database and the external link reference text.

[0101] The parsed work is carried out by the data decryption unit after formatting the code, and after the parsing is completed, the length of the character code is within the set value (0, 100), which can guarantee the recognition degree and recognition effect of the text, and the code is input into the database and compared with the coincidence and similarity of the comparison module, and after S comparison, the comparison text with coincidence and similarity found in the database is directly output as a report result, and the comparison text with no coincidence and similarity found in the database is input into the discrimination compensation module, which expands the search and comparison object of the report to be checked by the external network compensation unit in the discrimination compensation module, and the cited text includes logs, anthologies, monographs and documents, etc., and the word and sentence near-synonymous compensation unit in the discrimination compensation module will include the near-synonymous words and short sentences of the optimized extracted words and sentences into the object to be checked, so as to effectively avoid the report to be checked by using near-synonymous words and phrases to evade duplication, so as to reduce the coincidence and similarity, and the input characters are set by the second threshold setting unit, and the recognizable length is set to (0, 50), so as to further optimize the report result under the conditions of accurate and multiple recognition.

[0102] The determination result of the report to be checked is output by the determination result unit, and the determination point and the determination reason of the report to be checked are output by the comprehensive evaluation unit, and the comparison text with coincidence and similarity of the report to be checked is output by the similarity evaluation report.

[0103] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1.A method for checking plagiarism of scientific report achievements based on neural network feature recognition, characterized in that, It comprises the following steps: S1, report entry and screening: the report to be checked for duplication is subjected to corresponding text entry and information data screening through the entry screening module; S2, information data optimization extraction: the information data screened out in S1 is further optimized and extracted through the target data extraction module; S3, encoding conversion: the information data optimized and extracted in S2 is converted into character codes through the encoding conversion module, and data formatting is performed; S4, discrimination and comparison: the converted codes in S3 are subjected to corresponding decryption and division through the discrimination module so as to be within the set threshold value, and then input to the database for comparison through the comparison module; S5, discrimination compensation: after comparison in S4, the overlapping and similar comparison texts are directly output for report results through the database, while the non-overlapping and similar comparison texts in the database jump to the discrimination compensation module for secondary data compensation, and then the optimal report results are obtained; S6, result output: the results after S4 and S5 are output through the report result module; In S5, the specific operation method of the discrimination compensation module comprises the following steps: A1, the parameter BM jumps to the target network through T; A2, the target network Wz (W1, W2, W3... Wz+1); A3, the parameters in the target network Wz and the parameter BM are input to the DT for training; A4, the parameter BM is analyzed through the Key and the target network YG is constructed; A5, the target network YG comprises double experience pool groups Yz (Y1, Y2, Y3... Yz+1) and Gz (G1, G2, G3... Gz+1); A6, the parameters in Yz and Gz and the parameters in the target network Wz are input to the DT for training; A7, the target network is reset, and the training threshold value is set to (0, 25); A8, yes, no, meet the duplication checking results, when meeting the duplication checking results, directly enter the next stage of operation, and when not meeting the duplication checking results, according to the input path, respectively reverse to A3 and A6 to re-run; A9, save the network and end the training; Wherein, the parameter BM is the code, T is the jump platform, Wz is the external link experience pool, DT is the comparison module, Key is the external port, YG is the word and sentence synonym experience pool, Yz is the word experience pool, and Gz is the sentence experience pool; The discrimination compensation module comprises an external network compensation unit, a word and sentence synonym compensation unit, and a second threshold setting unit; Wherein, the external network compensation unit expands the object of the report to be checked for duplication through external link; The word and sentence synonym compensation unit includes the synonym words and short sentences of the optimized extracted words and sentences into the object to be checked for duplication; The second threshold setting unit is used for setting the recognizable length of the character code, and the set value is (0, 50). In S1, the entry screening module comprises a text input unit, an information extraction unit, and a data screening unit; 2.The sci-tech report achievement duplicate checking method based on neural network feature recognition according to claim 1, characterized in that: Wherein, the text input unit is used for inputting the report to be checked for duplication; The information extraction unit is used for extracting the total number of words, application fields, and text box structures in the report; The data screening unit is used for screening text directories, text chapters, and text paragraphs. ​ 3.The sci-tech report achievement duplicate checking method based on neural network feature recognition according to claim 2, characterized in that: In S2, the target data extraction module is used for segmenting the total number of extracted texts, application fields, and text box structures, and screening text directories, text chapters, and text paragraphs, and extracting key words, professional words, special words, and multiple repetitive words. 4.The sci-tech report achievement duplicate checking method based on neural network feature recognition according to claim 1, characterized in that: In S3, the encoding conversion module includes a code change conversion unit, a text code conversion unit, and a data formatting unit. The code change conversion unit is used for periodic conversion and update of the encoding content. The text code conversion unit is used for converting texts into character codes. The data formatting unit is used for partitioning and initializing the converted character codes. 5.The sci-tech report achievement duplicate checking method based on neural network feature recognition according to claim 1, characterized in that: In S4, the discrimination module includes a data decryption unit, a first threshold setting unit, and a database unit. The data decryption unit is used for analyzing the converted character codes. The first threshold setting unit is used for setting the recognizable length of the character codes, and the setting value is (0, 100). The database unit is used for providing search and comparison objects for the duplicate report. The comparison module is used for comparing the coincidence points and similarity of the duplicate report with the database unit. 6.The sci-tech report achievement duplicate checking method based on neural network feature recognition according to claim 1, characterized in that: In S6, the report result module includes a judgment result unit, a comprehensive evaluation unit, and a similarity evaluation report. The judgment result unit is used for outputting the judgment result of the duplicate report. The comprehensive evaluation unit is used for outputting the judgment points and reasons of the duplicate report judgment result. The similarity evaluation report is used for outputting the compared texts that coincide with and are similar to the duplicate report.

Citation Information

Patent Citations

  • Patent duplicate checking method and system based on complex network text language intention coding mode

    CN110609932A

  • Science and technology project duplicate checking method for automatically extracting synonyms based on deep learning algorithm

    CN110928985A