Document originality assessment methods, devices, equipment and storage media

By training large-scale language models and introducing external parameters, the problem that small-scale models cannot understand long and difficult sentences is solved, the accuracy and efficiency of document originality assessment are improved, and accurate originality scores and review opinions are generated.

CN118862862BActive Publication Date: 2025-09-12WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410881678.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-03
Publication Date
2025-09-12
Estimated Expiration
2044-07-03

AI Technical Summary

Technical Problem

In existing technologies, small-scale large language models cannot accurately understand long and difficult sentences in literature, resulting in inaccurate originality assessment results.

Method used

By training the first and second language models, and utilizing the external parameters and contextual information of the literature, accurate originality scores and review comments are generated. A large-scale language model is used to process complex sentences, and corrections are made through the second language model.

Benefits of technology

It improves the accuracy and efficiency of document originality assessment, can accurately understand long and difficult sentences, and generate quantitative and qualitative originality assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118862862B_ABST
    Figure CN118862862B_ABST
Patent Text Reader

Abstract

Disclosed are a method, apparatus, device, and storage medium for assessing document originality, belonging to the field of computer technology. The method comprises: training a first large language model based on a first data set, the first data set comprising multiple documents and external parameters of each document, the first large language model being used to generate a first originality score for the first document based on the first document and the external parameters of the first document; and training a second large language model based on a second data set and the originality score of each document in the multiple documents, the second large language model being used to generate a first originality review opinion and a revised first originality score for the first document based on the public review opinions of the first document, the context in which the first document is cited, and the first originality score. This method can accurately and efficiently assess document originality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method, apparatus, device and storage medium for evaluating the originality of a document. Background Art

[0002] For scientific and technological literature, originality is an important indicator, and it is necessary to accurately and efficiently determine the originality of the literature in order to evaluate its value.

[0003] In related technologies, a method for evaluating document originality includes: inputting the document into a small-scale large language model, and performing text analysis on the document through the small-scale large language model.

[0004] However, since there are often long and difficult sentences in the literature, small-scale large language models cannot accurately understand the long and difficult sentences in the literature, resulting in incorrect originality evaluation results. Summary of the Invention

[0005] The present disclosure provides a method, apparatus, device, and storage medium for evaluating document originality, which can accurately and efficiently evaluate document originality. The technical solution includes at least the following solutions:

[0006] In a first aspect, a method for assessing document originality is provided, comprising: training a first large language model based on a first data set, the first data set including multiple documents and external parameters of each document, the external parameters including the number of citations and the number of downloads, the first large language model being used to generate a first originality score for the first document based on the first document and the external parameters of the first document, where the first document is any one of the multiple documents; training a second large language model based on a second data set and the originality score of each document in the multiple documents, the second data set including public review opinions for each document in the multiple documents and the context in which each document in the multiple documents is cited, the second large language model being used to generate a first originality review opinion and a revised first originality score for the first document based on the public review opinions of the first document, the context in which the first document is cited, and the first originality score, the first originality review opinion and the revised first originality score being used to assess the originality of the first document; wherein the parameter scales of the first large language model and the second large language model are greater than a first parameter threshold.

[0007] Optionally, the external parameters include parameters in multiple fields. In the first data set, the external parameters of different fields of any document have different weights. The training of the first large language model based on the first data set includes: based on the third large language model, representing each document in the multiple documents by a semantic vector, and the parameter scale of the third large language model is less than the first parameter threshold; based on the third large language model, determining N second documents corresponding to N semantic vectors closest to the first semantic vector, the first semantic vector is the semantic vector corresponding to the first document, and the second document is any one of the multiple documents; inputting the first document, the N second documents, the external parameters of the first document, the external parameters of the N second documents, and a first instruction into the first large language model to obtain the first originality score, and the first instruction is used to guide the first large language model to output the first originality score; wherein N is an integer and N is greater than or equal to 0.

[0008] Optionally, the method further includes: based on the third largest language model, dividing the first document and the N second documents into multiple text blocks according to function, each text block corresponding to a label; the step of inputting the first document, the N second documents, the external parameters of the first document, the external parameters of the N second documents, and the first instruction into the first large language model includes: inputting the multiple text blocks and the labels corresponding to each of the text blocks, the external parameters of the first document, the external parameters of the N second documents, and the first instruction into the first large language model.

[0009] Optionally, the external parameters also include: disruptive index, citation structure, citation content, citation function, citation emotion, number of views, number of collections, number of reposts, comment polarity, number of patent citations, and patent forwarding volume.

[0010] Optionally, the second data set also includes virtual review opinions for each document, and the virtual review opinions are generated based on a fourth language model. The second language model is trained based on the second data set and the originality scores of each document in the multiple documents, including: inputting the public review opinions of the first document and the N second documents, the virtual review opinions of the first document and the N second documents, the context when the first document is cited, the first originality score, and a second instruction into the second language model to obtain the first originality review opinion and the revised first originality score, and the second instruction is used to guide the second language model to output the first originality review opinion and the revised first originality score.

[0011] Optionally, the first originality review opinion includes an originality comment on the first document and an originality type of the first document, wherein the originality types include: original problem, original method, original theory, original result, and original application, and the originality comment is used to explain the originality type and the revised first originality score.

[0012] Optionally, the method further includes: inputting a third instruction into the second language model to obtain the first originality review opinion in a first format and the revised first originality score.

[0013] In the second aspect, a document originality assessment device is also provided, including: a first training module, used to train a first large language model based on a first data set, the first data set including multiple documents and external parameters of each document, the external parameters including the number of citations and the number of downloads, the first large language model is used to generate a first originality score for the first document based on the first document and the external parameters of the first document, the first document being any one of the multiple documents; a second training module, used to train a second large language model based on a second data set and the originality score of each document in the multiple documents, the second data set including the public review opinions of each document in the multiple documents and the context in which each document in the multiple documents is cited, the second large language model is used to generate a first originality review opinion and a revised first originality score for the first document based on the public review opinions of the first document, the context in which the first document is cited, and the first originality score, the first originality review opinion and the revised first originality score being used to perform originality assessment on the first document.

[0014] Optionally, the external parameters include parameters in multiple fields. In the first data set, the external parameters of different fields of any document have different weights. The first training module is also used to: based on a third language model, perform semantic vector representation on each of the multiple documents, and the parameter scale of the third language model is less than the first parameter threshold; based on the third language model, determine N second documents corresponding to N semantic vectors closest to the first semantic vector, the first semantic vector is the semantic vector corresponding to the first document, and the second document is any one of the multiple documents; input the first document, the N second documents, the external parameters of the first document, the external parameters of the N second documents, and a first instruction into the first language model to obtain the first originality score, and the first instruction is used to guide the first language model to output the first originality score; wherein N is an integer and N is greater than or equal to 0.

[0015] Optionally, the device further includes: a division module, used to divide the first document and the N second documents into multiple text blocks according to function based on the third language model, each text block corresponding to a label; the first training module is also used to input the multiple text blocks and the labels corresponding to each of the text blocks, the external parameters of the first document, the external parameters of the N second documents and the first instruction into the first language model.

[0016] Optionally, the second data set also includes virtual review opinions for each document, and the virtual review opinions are generated based on the fourth language model. The second training module is also used to input the public review opinions of the first document and the N second documents, the virtual review opinions of the first document and the N second documents, the context when the first document is cited, the first originality score and the second instruction into the second language model to obtain the first originality review opinion and the revised first originality score. The second instruction is used to guide the second language model to output the first originality review opinion and the revised first originality score.

[0017] Optionally, the second training module is further used to input the third instruction into the second large language model to obtain the first originality review opinion in a first format and the revised first originality score.

[0018] In a third aspect, a computer device is also provided, comprising: a memory and a processor, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor, thereby executing the document originality evaluation method described in the above embodiment.

[0019] In a fourth aspect, a computer-readable storage medium is also provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor, thereby executing the document originality evaluation method described in the above embodiment.

[0020] In a fifth aspect, a computer program product is provided, comprising a computer program / instruction, which implements the method described in the first aspect when executed by a processor.

[0021] The beneficial effects of the technical solutions provided by the embodiments of the present disclosure include at least:

[0022] In the disclosed embodiment, by training the first and second language models, an assessment of the originality of a document is achieved. Since the large-scale language model can process more complex sentences, it can accurately understand long and difficult sentences in the document. By introducing external parameters of the document, the accuracy of the generated first originality score is improved, and through the second language model, the first originality score is corrected, further improving the accuracy of the generated first originality score. The second language model can also generate a first originality review opinion, thereby enabling an assessment of the originality of the document. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0024] Figure 1 A flowchart of a method for evaluating document originality provided by an exemplary embodiment of the present disclosure is shown;

[0025] Figure 2 A flowchart of a method for evaluating document originality provided by another exemplary embodiment of the present disclosure is shown;

[0026] Figure 3 A schematic structural diagram of a document originality evaluation device provided by an exemplary embodiment of the present disclosure is shown;

[0027] Figure 4 It is a structural diagram of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] Unless otherwise defined, the technical or scientific terms used herein shall have the usual meanings understood by persons of ordinary skill in the field to which the present disclosure belongs. The words “first”, “second”, “third” and similar terms used in the patent application specification and claims of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as “one” or “a” do not indicate a quantity limitation, but rather indicate the presence of at least one. Words such as “include” or “comprising” and similar terms mean that the elements or objects appearing before “include” or “comprising” cover the elements or objects listed after “include” or “comprising” and their equivalents, and do not exclude other elements or objects.

[0029] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0030] Figure 1 A flowchart of a method for evaluating document originality provided by an exemplary embodiment of the present disclosure is shown. The method can be executed by a computer device. Figure 1 , the method comprising:

[0031] In step 101, a first language model is trained based on a first data set.

[0032] The first data set includes multiple documents and external parameters of each document, the external parameters including the number of citations and the number of downloads. The first language model is used to generate a first originality score for the first document based on the first document and the external parameters of the first document, where the first document is any one of the multiple documents.

[0033] When training the first language model based on the first data set, for other documents in the first data set except the first document, the first language model is trained in the same manner as the first document, that is, the other documents and external parameters of the other documents are input into the first language model.

[0034] Optionally, originality is scored on a scale of 1-100.

[0035] Optionally, the multiple documents in the first data set can be obtained through platforms such as WoS (Web of Science), PubMed, China National Knowledge Infrastructure, and Wanfang Database.

[0036] In step 102 , a second language model is trained based on the second data set and the originality score of each document in the plurality of documents.

[0037] The second data set includes the public review opinions of each document in the multiple documents and the context in which each document in the multiple documents is cited. The second largest language model is used to generate the first originality review opinion and the revised first originality score of the first document based on the public review opinions of the first document, the context in which the first document is cited, and the first originality score. The first originality review opinion and the revised first originality score are used to evaluate the originality of the first document.

[0038] The parameter scales of the first and second largest language models are greater than a first parameter threshold. For example, the first parameter threshold ranges from 1B to 2B (1 billion to 2 billion).

[0039] A large language model with a value greater than the first parameter threshold is also known as a large-scale language model. A large-scale language model can process more complex sentences and accurately understand long and difficult sentences in the literature.

[0040] For example, the first language model and the second language model are Mixtral-type large language models. Public review comments can be obtained through the Open Review platform.

[0041] After the first language model and the second language model are trained, the originality of the document can be evaluated based on the trained first language model and the second language model.

[0042] In the disclosed embodiment, by training the first and second largest language models, an assessment of the originality of a document is achieved. Since the large-scale language model can handle more complex sentences, it can accurately understand long and difficult sentences in the document. By introducing external parameters of the document, the accuracy of the generated first originality score is improved, and through the second largest language model, the first originality score is corrected, further improving the accuracy of the generated first originality score. The second largest language model can also generate a first originality review opinion, thereby enabling the originality assessment of the document.

[0043] Figure 2 A flowchart of a method for evaluating document originality provided by another exemplary embodiment of the present disclosure is shown. The method can be executed by a computer device. Figure 2 , the method comprising:

[0044] In step 201, a first language model is trained based on a first data set.

[0045] The first data set includes multiple documents and external parameters of each document, the external parameters including the number of citations and the number of downloads. The first language model is used to generate a first originality score for the first document based on the first document and the external parameters of the first document, where the first document is any one of the multiple documents.

[0046] Optionally, the external parameters include parameters in multiple fields, and in the first data set, the external parameters in different fields of any document have different weights.

[0047] In the disclosed embodiments, the multiple fields include: scientific impact, social impact, and economic impact. The external parameters for scientific impact include: disruptive index, citation count, citation width, citation breadth, citation function, and citation location; the external parameters for social impact include: download count, pageview count, favorite count, repost count, and comment polarity; and the external parameters for economic impact include: patent citation count and patent forwarding volume.

[0048] Citation count is the number of times a document has been cited.

[0049] The disruptive indicator, also known as the disruptive index, is a quantitative indicator proposed in recent years that can directly measure the degree of disruptive innovation of a paper. In the field of scientific impact, the originality contribution score of a scientific impact field can be calculated based on the disruptive indicator, citation structure, citation content, citation function, and citation sentiment. This score can represent the contribution of external parameters such as disruptive indicators, citation structure, citation content, citation function, and citation sentiment to the originality of the document. When inputting external parameters into the first language model, for external parameters such as disruptive indicators, citation structure, citation content, citation function, and citation sentiment, it is only necessary to input the calculated originality contribution score of the scientific impact field.

[0050] Optionally, the original contribution score in the field of scientific impact is calculated using formula (1).

[0051]

[0052] In the disruptive index, the types of papers cited by document k can be divided into Type I and Type J. Type I refers to papers that only cite document k and no other references, while Type J refers to papers that cite at least two different references, one of which is document k.

[0053] In formula (1), S(k) represents the originality contribution score of document k, The score related to Category I literature in the originality contribution score is calculated using formula (2); The score related to category J literature in the originality contribution score is calculated using formula (3). Express Know Perform element-wise summation of a vector.

[0054]

[0055] In formula (2), i∈I represents any document i in category I. The originality contribution score is the score of document i in category I. The other parameters in formula (2) have the same meaning as in formula (1) and are not described in detail here.

[0056]

[0057] In formula (3), j∈J represents any document j in category J. Indicates the score of original contribution related to document j in category J. R represents the score of the original contribution score of document k related to the reference r of document j in category J. j,kThe set of all references cited by document j. Reference r is any reference cited by document j. The other parameters in formula (3) have the same meanings as in formula (1) and are not described in detail here.

[0058] Represents the fraction of document k's originality contribution score relative to all references of document j in category J. When multiple references r contribute to document j, to avoid duplicate counting, only the maximum originality contribution in a specific dimension is selected. This fraction is then relative to all references of document j in category J in the originality contribution score of document k.

[0059] If there are multiple references r that contribute to the method of document j, only the most important relevant scores among the multiple references r are selected, and the relevant scores of the remaining references r may be repeated with the most important one, or have only a negligible impact.

[0060] In formula (2) It can be calculated using formula (4).

[0061] In formula (3) as well as Compared with the formula (2) The calculation principle is the same. When calculating When , it is sufficient to replace i with j and k with r in formula (4). The detailed description is omitted here.

[0062]

[0063] In formula (4), It represents the score related to the citation structure and is calculated by formula (5). i,k represents the set of sentences in which document i cites document k in the citation structure, and (i, k, n) represents the nth sentence in which document i cites document k in the citation structure. The other parameters in formula (4) have the same meanings as in formula (2) and are not described in detail here.

[0064] Optionally, the citation structure includes: introduction, related research, research methods, research results and research conclusions. When document i cites document k, there may be multiple references to document k in document i. If a reference to document k belongs to any of the above citation structures, then N i,k This sentence is included in .

[0065]

[0066] In formula (5), They represent the scores of the question Q, method M, result R, theory T, and application A in the cited content, respectively, and are calculated using formula (6). The other parameters in formula (5) have the same meaning as in formula (4) and are not described in detail here.

[0067] In formula (6), F is an indicator function used to indicate whether there is originality in problem, method, result, theory or application. (i,k,n) is the score related to the citation function, calculated by formula (7), E (i,k,n) is the score related to the reference sentiment, calculated using formula (8).

[0068]

[0069] In formula (7), P F For a reference function F (i,k,n) The probability of occurrence, is the set of sentences with all reference functions. Among them, reference functions include: background, based on, use, and comparison. For example, F (i,k,n) This represents the score associated with the function of the nth sentence of document k cited by document i in the citation structure belonging to any one or more of the following: background, based on, use, and comparison. The other parameters in formula (7) have the same meaning as in formula (6) and are not described in detail here.

[0070]

[0071] E (i,k,n) =0, E∈{Negative}

[0072] In formula (8), P E For a certain reference emotion E (i,k,n) The probability of occurrence, {Positive, Neutral} indicates that the quoted sentiment belongs to positive or neutral sentences; {Negative} indicates that the quoted sentiment belongs to negative sentences. Among them, the quoted functions include: background, based on, use, and comparison. For example, F (i,k,n) This is the score associated with the function of the nth sentence of document k cited by document i in the citation structure belonging to any one or more of the following: background, based on, use, and comparison. The other parameters in formula (8) have the same meaning as in formula (6) and are not described in detail here.

[0073] Download count is the number of times a document has been downloaded; view count is the number of times a document has been viewed; save count is the number of times a document has been saved; and repost count is the number of times a document has been reposted (forwarded). Comment polarity, like citation sentiment, also includes positive, negative, and neutral.

[0074] In the field of social impact, a social media score for a social impact field can be calculated based on comment polarity, download counts, view counts, favorite counts, and repost counts. When inputting external parameters into the first language model, for external parameters such as comment polarity, download counts, view counts, favorite counts, and repost counts, it is only necessary to input the calculated social media score for the field of scientific impact.

[0075] When calculating the social media score, the above formulas (1) to (8) can also be referred to, and the number of downloads, views, favorites, and reposts can be used as the above reference structure. By using the text information in the process of downloads, views, favorites, and reposts, the reference content, reference function, and reference emotion can be identified, and then the relevant social media score can be calculated.

[0076] Some documents also have corresponding patents, so these documents also have external parameters related to economic impact. Patent citations are the number of times a patent corresponding to a document has been cited, while patent forwarding is the number of times a patent corresponding to a document has been forwarded. Economic impact parameters can be directly weighted or calculated using social media scoring methods, but this is not limited in the present disclosure.

[0077] Each document has different weights for external parameters in different fields. The weights are set based on experience. External parameters in the same field have the same weight, while external parameters in different fields can have the same or different weights.

[0078] For example, if a document has the greatest impact in the scientific field, such as a large number of citations, the weight of the document's external parameters in the field of scientific impact may be higher than the external parameters in the fields of social impact and economic impact.

[0079] If a document has the greatest impact in the social field, such as high popularity, the weight of the document's external parameters in the field of social impact can be higher than the external parameters of the document in the fields of scientific impact and economic impact.

[0080] If a document has the greatest impact in the economic field, for example, the corresponding patent document is widely used and the patent income is high, then the external parameters of the document in the economic impact field can be higher than the external parameters of the document in the social impact field and the field that can be influenced.

[0081] In the embodiment of the present disclosure, by weighting the external parameters, when the external parameters are subsequently input into the first language model, the accuracy of the originality score generated by the first language model can be improved.

[0082] In this case, step 201 includes the following steps ac:

[0083] Step a: Based on the third language model, each document in the multiple documents is represented by a semantic vector.

[0084] The parameter scale of the third largest language model is smaller than the first parameter threshold. The relevant content of the first parameter threshold is referred to the aforementioned step 102, and the detailed description is omitted here.

[0085] Here, the large language model whose parameter scale is smaller than the first parameter threshold is also called a small model, that is, a small-scale large language model.

[0086] Semantic vector representation is the process of converting a document into a semantic vector. When performing semantic vector representation on a document, only the main body of the document is represented. The main body of the document includes the title, main text, abstract, etc., and does not include references.

[0087] After representing each document with a semantic vector, a document vector database can be established based on the semantic vectors of multiple documents. The document vector database includes multiple sub-databases, such as the title sub-database, the body text sub-database, and the abstract sub-database. Semantic vectors of the same type are stored in the same sub-database.

[0088] Exemplarily, the third largest language model may be a BERT-type model, such as the SBERT large language model.

[0089] Step b: Based on the third language model, determine N second documents corresponding to N semantic vectors closest to the first semantic vector.

[0090] The first semantic vector is a semantic vector corresponding to the first document, and the second document is any one of the multiple documents.

[0091] Wherein, N is an integer and N is greater than or equal to 0.

[0092] Step b essentially involves inputting the first document into the third language model and then calculating semantic vector similarity to determine the N closest semantic vectors in the document vector database to the first. The documents corresponding to these N semantic vectors are the second documents. These N second documents are the N closest documents to the first document.

[0093] Step c: inputting the first document, N second documents, external parameters of the first document, N external parameters of the second documents, and the first instruction into a first large language model to obtain a first originality score.

[0094] The first instruction is used to guide the first language model to output a first originality score. For example, the first instruction is: As a reviewer in the biomedical field, you need to compare the research content differences between similar papers and core papers and generate an originality score between 0 and 100.

[0095] In the disclosed embodiment, the first instruction is used to guide the first language model to generate the sequence of actions (or specific actions) of originality scores, and to limit the specific technical field to improve the accuracy of the generated originality scores. In the above example, the first instruction limits the technical field to the biomedical field, and guides the specific actions of the first language model, that is, to compare the research content differences between similar papers and core papers. Guiding the first language model to generate originality scores through the first instruction can also improve the accuracy of the originality scores generated by the first language model.

[0096] Optionally, the method further includes: based on the third language model, dividing the first document and N second documents into multiple text blocks according to their functions, each text block corresponds to a label, and text blocks with the same function have the same label. Here, each paragraph in the document has a different function, including: background, results (conclusions), experiments, etc. If the functions of several consecutive text paragraphs are the same, for example, they are used to describe the background, then these consecutive text paragraphs belong to the same text block and correspond to the same label, that is, the label corresponding to the background. This process can also be called chapter structuring processing.

[0097] In this case, step c includes: inputting multiple text blocks and labels corresponding to each text block, external parameters of the first document, N external parameters of the second documents, and the first instruction into the first large language model.

[0098] Inputting multiple text blocks and the labels corresponding to each text block into the first large language model, rather than directly inputting the full text of the document into the large language model, can improve the accuracy of the originality score generated by the large language model.

[0099] In step 202 , a second language model is trained based on the second data set and the originality score of each document in the plurality of documents.

[0100] Optionally, the second data set also includes virtual review comments for each document, where the virtual review comments are generated based on a fourth language model. Here, the fourth language model is a large language model pre-trained with a small sample, and illustratively, the fourth language model can be GPT-4.

[0101] Optionally, the fourth language model generates virtual review opinions in the following manner: multiple documents are input into the fourth language model in sequence to obtain initial virtual review opinions for each document; the initial virtual review opinions for each document are manually screened to filter out unreliable virtual review opinions and retain credible virtual review opinions; the documents corresponding to the unreliable virtual review opinions are input into the fourth language model again, and the above steps are repeated until credible virtual review opinions for each document are obtained, and the credible virtual review opinions are stored in the second data set.

[0102] In this case, step 202 includes:

[0103] The public review opinions of the first document and N second documents, the virtual review opinions of the first document and N second documents, the context when the first document is cited, the first originality score and the second instruction are input into the second language model to obtain the first originality review opinion and the revised first originality score.

[0104] Optionally, the first originality review opinion includes the originality comments of the first document and the originality type of the first document. The originality types include: original problem, original method, original theory, original result and original application. The originality comments are used to explain the originality type and the revised first originality score.

[0105] In the disclosed embodiment, after comparing the first document with N second documents, if the first document is original, the second largest language model can obtain the original text portion in the first document, and then the second largest language model can classify the original text portion in the first document to obtain the originality type of the first document, and explain the originality type through the originality comment.

[0106] The second instruction guides the second language model to output the first originality review opinion and the revised first originality score. For example, the second instruction reads: "As a review expert in the biomedical field, you need to understand the following definitions of originality types and dimensions: ...." You need to analyze the input text and originality value scores and generate the originality type and explanatory comments for the core paper according to the required evaluation steps.

[0107] In an embodiment of the present disclosure, the second instruction is used to guide the second largest language model to generate the first originality review opinion and the revised first originality score in sequence (or specific actions), and to limit the specific technical field, thereby improving the accuracy of the output of the second largest language model. In the above example, the second instruction limits the technical field to the biomedical field, and guides the second largest language model to perform the following actions, that is, first understand the definition of originality type and originality dimension, and then analyze the input text (such as the first document and N second documents) and the input originality value score (such as the first originality score), thereby generating the first originality review opinion and the revised first originality score of the input document. Guiding the second largest language model to generate the originality review opinion and the revised originality score through the second instruction can improve the accuracy of the originality review opinion and the revised originality score generated by the second largest language model.

[0108] After obtaining the first originality review opinion and the revised first originality score, the second language model needs to be subjected to human feedback reinforcement learning (RLHF) to improve the accuracy of the output of the second language model.

[0109] There are many methods for implementing RLHF in related technologies, and detailed description is omitted here.

[0110] Optionally, steps 201-202 are implemented by using efficient fine-tuning technology for model parameters, for example, by using LoRA, QLoRA, Prefix-tuning and other technologies.

[0111] In step 203 , the third instruction is input into the second language model to obtain a first originality review opinion in a first format and a revised first originality score.

[0112] In the embodiment of the present disclosure, the third instruction is used to guide the second large language model to convert the output format into the first format.

[0113] Exemplarily, the first format is json format.

[0114] Through steps 201-203, the first and second language models can be trained, and the accuracy of their outputs can be improved. After both the first and second language models are trained, the originality of the document can be evaluated based on the trained first and second language models.

[0115] When evaluating the originality of a document based on the trained first and second language models, the document to be evaluated is first input into the third language model to obtain N similar documents closest to the document to be evaluated, and the document to be evaluated and the N similar documents are subjected to text structuring processing; then the document to be evaluated and the N similar documents after text structuring processing are input into the first language model to obtain the originality score of the document to be evaluated; then the document to be evaluated, the N similar documents and the originality score of the document to be evaluated are input into the second language model to obtain the originality type, originality comments and the revised originality score of the document to be evaluated.

[0116] In the disclosed embodiment, by utilizing a large-scale large language model (such as the first large language model and the second large language model) and a small-scale large language model (such as the third large language model) in a collaborative manner, the originality value of the document is quantitatively evaluated (originality score) and qualitatively evaluated (originality review opinions), thereby improving the efficiency of originality evaluation of the document. Among them, the small-scale large language model is responsible for data preprocessing of the document (such as paragraph structuring processing and semantic vector representation) and finding similar documents. The large-scale large language model is responsible for performing semantic comparative analysis on the document to be evaluated and similar documents, and generating quantitative originality review opinions, that is, originality types and originality comments.

[0117] In this way, the semantic understanding, text generation and knowledge reasoning capabilities of the large model are fully utilized to generate originality review opinions and originality scores, and the generated originality review opinions and originality scores are highly accurate and interpretable.

[0118] The following are device embodiments of the present application. For details not described in detail in the device embodiments, reference may be made to the above method embodiments.

[0119] Figure 3 A schematic diagram of a document originality evaluation device provided by an exemplary embodiment of the present disclosure is shown. Figure 3 The document originality evaluation device 300 includes: a first training module 301 and a second training module 302.

[0120] The first training module 301 is used to train a first large language model based on a first data set. The first data set includes multiple documents and external parameters of each document. The external parameters include the number of citations and the number of downloads. The first large language model is used to generate a first originality score for the first document based on the first document and the external parameters of the first document. The first document is any one of the multiple documents.

[0121] The second training module 302 is used to train a second large language model based on a second data set and the originality score of each document in the multiple documents. The second data set includes the public review opinions of each document in the multiple documents and the context in which each document in the multiple documents is cited. The second large language model is used to generate a first originality review opinion and a revised first originality score of the first document based on the public review opinions of the first document, the context in which the first document is cited, and the first originality score. The first originality review opinion and the revised first originality score are used to evaluate the originality of the first document.

[0122] Optionally, the external parameters include parameters in multiple fields. In the first data set, the external parameters of different fields of any document have different weights. The first training module 301 is also used to: based on the third largest language model, represent each document in the multiple documents with a semantic vector, and the parameter scale of the third largest language model is less than the first parameter threshold; based on the third largest language model, determine the N second documents corresponding to the N semantic vectors closest to the first semantic vector, the first semantic vector is the semantic vector corresponding to the first document, and the second document is any one of the multiple documents; input the first document, N second documents, the external parameters of the first document, the external parameters of the N second documents, and the first instruction into the first large language model to obtain a first originality score, and the first instruction is used to guide the first large language model to output the first originality score; wherein N is an integer and N is greater than or equal to 0.

[0123] Optionally, the device also includes: a division module 303, which is used to divide the first document and N second documents into multiple text blocks according to function based on the third language model, and each text block corresponds to a label; the first training module 301 is also used to input multiple text blocks and the labels corresponding to each text block, the external parameters of the first document, the external parameters of the N second documents, and the first instruction into the first language model.

[0124] Optionally, the second data set also includes virtual review opinions for each document, and the virtual review opinions are generated based on the fourth language model. The second training module 302 is also used to input the public review opinions of the first document and N second documents, the virtual review opinions of the first document and N second documents, the context when the first document is cited, the first originality score and the second instruction into the second language model to obtain the first originality review opinion and the revised first originality score. The second instruction is used to guide the second language model to output the first originality review opinion and the revised first originality score.

[0125] Optionally, the second training module 302 is further configured to input the third instruction into the second large language model to obtain a first originality review opinion in a first format and a revised first originality score.

[0126] It should be noted that the document originality assessment device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate the document originality assessment. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the document originality assessment device provided in the above embodiment and the document originality assessment method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0127] The division of modules in the embodiments of the present disclosure is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the present disclosure may be integrated into a single processor, exist physically as separate modules, or be integrated into a single module. The integrated modules may be implemented in either hardware or software functional modules.

[0128] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a terminal device (which can be a personal computer, mobile phone, or communication device, etc.) or a processor (processor) to execute all or part of the steps of the method of each embodiment of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.

[0129] Figure 4 Schematic diagram of the structure of the computer device provided by the embodiment of the present disclosure. Figure 4 As shown, the computer device 400 includes a processor 401 and a memory 402 .

[0130] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0131] Memory 402 may include one or more computer-readable storage media, which may be non-transitory. Memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 402 is used to store at least one instruction, which is executed by processor 401 to implement the document originality assessment method provided in the embodiments of the present disclosure.

[0132] Those skilled in the art will understand that Figure 4 The structure shown in the figure does not constitute a limitation on the computer device 400, and the computer device 400 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.

[0133] The embodiment of the present disclosure also provides a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by the processor of a computer device, the computer device is enabled to execute the document originality evaluation method provided in the embodiment of the present disclosure.

[0134] The embodiments of the present disclosure also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the document originality evaluation method provided in the embodiments of the present disclosure.

[0135] The above description is merely an optional embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure.

Claims

1. A method for evaluating the originality of a document, characterized in that: The method comprises: training a first large language model based on a first data set, wherein the first data set includes a plurality of documents and external parameters of each document, the external parameters including a number of citations and a number of downloads, and the first large language model is used to generate a first originality score for the first document based on the first document and the external parameters of the first document, where the first document is any one of the plurality of documents; training a second language model based on a second data set and the originality score of each document in the plurality of documents, wherein the second data set includes public review opinions of each document in the plurality of documents and the context in which each document in the plurality of documents is cited, and the second language model is used to generate a first originality review opinion and a revised first originality score for the first document based on the public review opinions of the first document, the context in which the first document is cited, and the first originality score, and the first originality review opinion and the revised first originality score are used to perform originality assessment on the first document; The parameter scales of the first largest language model and the second largest language model are greater than a first parameter threshold.

2. The method according to claim 1, characterized in that The external parameters include parameters in multiple fields. In the first data set, the external parameters in different fields of any document have different weights. Training the first language model based on the first data set includes: Performing a semantic vector representation on each of the plurality of documents based on a third language model, wherein a parameter scale of the third language model is less than the first parameter threshold; Determining, based on the third language model, N second documents corresponding to N semantic vectors closest to the first semantic vector, where the first semantic vector is a semantic vector corresponding to the first document, and the second document is any one of the multiple documents; Inputting the first document, the N second documents, external parameters of the first document, external parameters of the N second documents, and a first instruction into the first large language model to obtain the first originality score, wherein the first instruction is used to guide the first large language model to output the first originality score; Wherein, N is an integer and N is greater than or equal to 0.

3. The method according to claim 2, characterized in that The method further comprises: Based on the third language model, the first document and the N second documents are divided into a plurality of text blocks according to their functions, each text block corresponding to a label; The step of inputting the first document, the N second documents, the external parameters of the first document, the external parameters of the N second documents, and the first instruction into the first large language model includes: The multiple text blocks and the label corresponding to each of the text blocks, the external parameters of the first document, the external parameters of the N second documents, and the first instruction are input into the first large language model.

4. The method according to any one of claims 1 to 3, characterized in that The external parameters also include: disruptive indicators, citation structure, citation content, citation function, citation sentiment, number of views, number of collections, number of reposts, comment polarity, number of patent citations, and patent forwarding volume.

5. The method according to claim 2, characterized in that The second data set also includes virtual review comments for each document, which are generated based on the fourth language model. The step of training a second language model based on the second data set and the originality score of each document in the plurality of documents comprises: The public review opinions of the first document and the N second documents, the virtual review opinions of the first document and the N second documents, the context when the first document is cited, the first originality score and the second instruction are input into the second large language model to obtain the first originality review opinion and the revised first originality score. The second instruction is used to guide the second large language model to output the first originality review opinion and the revised first originality score.

6. The method according to claim 5, characterized in that The first originality review opinion includes an originality comment on the first document and an originality type of the first document. The originality types include: original problem, original method, original theory, original result, and original application. The originality comment is used to explain the originality type and the revised first originality score.

7. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Input the third instruction into the second language model to obtain the first originality review opinion in a first format and the revised first originality score.

8. A document originality assessment device, characterized in that: The device comprises: a first training module configured to train a first large language model based on a first data set, wherein the first data set includes a plurality of documents and external parameters of each document, wherein the external parameters include a number of citations and a number of downloads, and wherein the first large language model is configured to generate a first originality score for the first document based on the first document and the external parameters of the first document, wherein the first document is any one of the plurality of documents; The second training module is used to train a second large language model based on a second data set and the originality score of each document in the multiple documents, wherein the second data set includes the public review opinions of each document in the multiple documents and the context in which each document in the multiple documents is cited. The second large language model is used to generate a first originality review opinion and a revised first originality score of the first document based on the public review opinions of the first document, the context in which the first document is cited, and the first originality score. The first originality review opinion and the revised first originality score are used to evaluate the originality of the first document.

9. A computer device, characterized in that: The computer device includes: a memory and a processor, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the method according to any one of claims 1 to 7.