A text evaluation method and device, electronic equipment and storage medium
By identifying and analyzing citation markers and document identification numbers in academic texts, and combining them with machine learning models for evaluation, the problem of low efficiency and poor accuracy in academic text evaluation is solved, achieving efficient and accurate text quality evaluation, and supporting the optimization of large language models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the evaluation of citation references in academic texts is inefficient and inaccurate, and usually relies on manual review, resulting in inefficiency and inaccurate evaluation.
By identifying citation markers and document identification numbers in the text to be evaluated, text attributes and citation attributes are determined. A machine learning model is used to quickly locate citation fields and find document identification numbers, and evaluation is carried out by combining preset rules and weight calculations.
It improves the accuracy and efficiency of academic text evaluation, can characterize text quality from multiple dimensions, provides quality feedback for academic texts generated by large language models, and supports model iterative optimization.
Smart Images

Figure CN121435959B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a text evaluation method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the continuous development of society and the significant progress of technology, an increasing number of academic articles and news releases are emerging in various industries and fields. The quality of these academic texts varies greatly. To facilitate reading, learning, and discussion among professionals in related fields, these academic texts can be evaluated. Therefore, how to evaluate academic texts has become one of the key research focuses for professionals in related fields.
[0003] Currently, there are few means to examine and evaluate the use of references in academic texts. Usually, the review and judgment are done manually by auditors, which is inefficient and the accuracy of the evaluation is also questionable. Summary of the Invention
[0004] This application provides a text evaluation method, apparatus, electronic device, and storage medium to improve the accuracy of text content evaluation.
[0005] According to one aspect of this application, a text evaluation method is provided, comprising:
[0006] Obtain the text to be evaluated;
[0007] Based on the citation marks in the text to be evaluated, determine all citation fields in the text to be evaluated and the corresponding document identification number for each citation field;
[0008] Based on each citation field and the corresponding document identification number, determine at least one text attribute and at least one citation attribute of the text to be evaluated;
[0009] The text to be evaluated is evaluated based on each text attribute and each citation attribute.
[0010] According to another aspect of this application, a text evaluation apparatus is provided, comprising:
[0011] The text acquisition module is used to acquire the text to be evaluated.
[0012] The field determination module is used to determine all cited fields in the text to be evaluated and the corresponding document identification number for each cited field, based on the citation marks in the text to be evaluated.
[0013] The attribute determination module is used to determine at least one text attribute and at least one citation attribute of the text to be evaluated based on each citation field and the corresponding document identification number.
[0014] The attribute evaluation module is used to evaluate the text to be evaluated based on each text attribute and each reference attribute.
[0015] According to another aspect of this application, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the text evaluation method described in any embodiment of this application.
[0019] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the text evaluation method described in any embodiment of this application.
[0020] According to another aspect of this application, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the text evaluation method according to any embodiment of this application.
[0021] In the technical solution of this application embodiment, the citation markers in the text to be evaluated generated by the large language model are identified, and all citation fields in the text to be evaluated and the corresponding document identification numbers of each citation field are determined. This allows for the rapid location of the citation fields in the text to be evaluated and their specific content, and the rapid retrieval of the corresponding document identification numbers for determining the citation attributes. Based on each citation field and the corresponding document identification number, at least one text attribute and at least one citation attribute of the text to be evaluated are determined. The analysis of the text to be evaluated yields text attributes and citation attributes, providing a basis for text evaluation. Different text attributes and different citation attributes can characterize the quality of the text to be evaluated from different dimensions, comprehensively evaluating the text and improving the accuracy of text evaluation.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a text evaluation method provided according to Embodiment 1 of this application;
[0025] Figure 2 This is a schematic diagram of a text evaluation process applicable according to Embodiment 2 of this application;
[0026] Figure 3 This is a schematic diagram of the structure of a text evaluation device according to Embodiment 3 of this application;
[0027] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the text evaluation method of the embodiments of this application. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1This application provides a flowchart of a text evaluation method according to Embodiment 1. This embodiment is applicable to the detection and evaluation of academic text generated by a large language model. The method can be executed by a text evaluation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0032] S110. Obtain the text to be evaluated.
[0033] The text to be evaluated can be any academically relevant text that needs to be evaluated, such as texts that cite other literature in papers or news articles, or academic text generated by a large language model. The large language model can be any model capable of generating natural language text or understanding language text. This application embodiment and its various implementations will take the evaluation of academic text generated by a target large language model as an example. Any large language model from related technologies can be used in the examples of this application embodiment and its various implementations. After receiving a user's question instruction, the target large language model generates answer text related to the question. It is understood that general large language models will search for information on the internet during the analysis of questions and answers as the basis for answering questions and generating text, that is, they will cite information from the internet; however, in the example of this application embodiment, the target large language model is used to generate learning text, and of course, it refers to the content of other academic literature. Therefore, the generated text to be evaluated contains citations of literature.
[0034] S120. Based on the citation marks in the text to be evaluated, determine all citation fields in the text to be evaluated and the corresponding document identification number for each citation field.
[0035] The citation marker can be a preset marker near certain words or phrases in the text to be evaluated. This preset marker is used to mark which words or phrases in the text to be evaluated are cited from other documents. These words or phrases that cite other documents can be the citation field. The document identification number can be a unique identifier corresponding to the citation field. Using this document identification number, the original document can be found through retrieval platforms, document databases, etc. For example, it can be a DOI (Digital Object Unique Identifier), etc. This application embodiment does not limit this.
[0036] It is understandable that the large language model will have citation marks in the text to be evaluated where other literature is cited, such as "[1]" or "①". By identifying these citation marks in the text to be evaluated, it is determined which sentences in the text to be evaluated cite other literature, that is, the citation field is identified. The method for identifying citation marks can be any character recognition method among related technologies such as OCR (Optical Character Recognition). Then, the punctuation marks before the citation marks (such as periods or exclamation marks, especially the symbols that indicate the end of a sentence) are identified by the character recognition method to determine the sentences between the punctuation marks and the aforementioned citation marks. The interface is used as the citation field for subsequent processing. At the same time, at the end of the text generated by the large language model or during the thinking process of the output of the large language model, the information of the cited literature will be displayed. The information of the cited literature can be found in the displayed information through the citation marks, including the literature identification number of the literature.
[0037] S130. Based on each citation field and the corresponding document identification number, determine at least one text attribute and at least one citation attribute of the text to be evaluated.
[0038] Among them, text attributes can be inherent attributes of the text to be evaluated, such as the number of words in the text or the number of words in the cited content. These are attribute information that can be obtained directly without processing, calculation, or analysis. Citation attributes can be attribute information related to the cited content in the text to be evaluated, such as including but not limited to the number of citations and the number of cited documents. Citation attributes are not limited to the attributes of the citation field itself. The original text of the document can also be obtained from public sources based on the document identification number and further compared with the citation field to analyze whether the citation is correct or reasonable. These attributes are given quantitative standards through preset evaluation criteria. This application does not exhaustively list all citation attributes.
[0039] Text attributes can be directly obtained by scanning the text in the text to be evaluated. Reference attributes can be processed for reference fields, such as by identifying reference attributes of different reference fields through pre-trained machine learning models. This application does not limit this.
[0040] S140. Evaluate the text to be evaluated based on each text attribute and each reference attribute.
[0041] Different text attributes and different citation attributes can be quantified by assigning values according to preset rules. For example, the number of words can be categorized into different tiers. If the number of words in the generated text to be evaluated matches a preset tier, then the text to be evaluated is assigned a score corresponding to that tier for the text attribute of word count. By assigning values to all information in the text and citation attributes, and pre-setting different weights based on the different importance of each attribute to the evaluation, the scores are weighted and calculated to obtain the total score of the text to be evaluated. For example, a higher score indicates a higher quality of the text generated by the large language model. Of course, the evaluation of the text to be evaluated does not necessarily require a quantitative result. It is also possible to summarize the information from each dimension of the text and citation attributes to generate a report that reflects the quality of the text to be evaluated.
[0042] In the technical solution of this application embodiment, the citation markers in the text to be evaluated generated by the large language model are identified, and all citation fields and their corresponding document identification numbers are determined in the text to be evaluated. This allows for the rapid location of the citation fields and their specific content within the text to be evaluated, and the rapid retrieval of the corresponding document identification numbers for determining citation attributes. Based on each citation field and its corresponding document identification number, at least one text attribute and at least one citation attribute of the text to be evaluated are determined. The analysis of the text to be evaluated yields text attributes and citation attributes, providing a basis for text evaluation. Different text attributes and different citation attributes can characterize the quality of the text to be evaluated from different dimensions, comprehensively evaluating the text and improving the accuracy of text evaluation. For academic texts generated by the large language model, this also helps adjust the parameters of the large language model based on the quality feedback of the generated text, facilitating the iterative evolution of the large language model.
[0043] In one optional implementation, the citation attribute includes the number of cited references and the number of citations; the step of determining at least one citation attribute of the text to be evaluated based on each citation field and the corresponding document identification number in S130 may include:
[0044] A1. Determine the number of citations in the text to be evaluated based on each citation field.
[0045] The citation count can be the total number of cited fields appearing in the text to be evaluated, i.e., the number of times a reference is cited. Even if the same reference is cited multiple times, multiple citation counts need to be accumulated. Based on the citation fields identified in the aforementioned embodiments, the total number of times such citation fields appear can be determined, or the number of citation identifiers can be directly identified to determine the number of times references are cited in the text to be evaluated.
[0046] A2. Perform a plagiarism check based on the document identification number to determine the number of cited references in the text to be evaluated.
[0047] The number of cited references can be the total number of references cited in the text to be evaluated. If multiple references belong to the same reference, they are only counted as one reference. Therefore, by checking the reference identification numbers for duplicates, multiple references cited in the text to be evaluated are identified, and the total number of cited references is counted as the number of cited references. Of course, any plagiarism detection algorithm in the relevant technical field can be used, and this application embodiment does not limit this.
[0048] In the above implementation, by analyzing and statistically analyzing each citation field and the corresponding document identification number, the citation count and number of cited documents in the text to be evaluated are identified, providing references from different dimensions for subsequent comprehensive evaluation and helping to improve the accuracy of academic text evaluation.
[0049] In one optional implementation, the citation attribute further includes citation accuracy and citation correctness; the step of determining at least one citation attribute of the text to be evaluated based on each citation field and the corresponding document identification number in S130 may include:
[0050] B1. Based on each document identification number, search the preset document database to determine the correct document citation in each citation field.
[0051] The literature database can be a pre-deployed database or an existing online database that stores a large number of existing documents. Theoretically, the corresponding document can be retrieved in this database based on the correct document identification number. A correct citation field is a field that accurately cites the correct document among all citation fields.
[0052] If the corresponding document can be retrieved in the literature database using the document identification number, meaning the document cited in the citation field does indeed exist, then that citation field can be used as the correct citation entry for the document.
[0053] B2. Based on the correct citations in each document and the corresponding cited documents, determine the correct citations for the content.
[0054] In this context, a correctly cited item can be a citation field that accurately reflects the content of the cited literature. It's important to note that citing a literature reference is not a mere excerpt, but rather a reasonable expression within the generated text's context without altering the original content and meaning. However, due to the illusion inherent in large language models, which can easily misinterpret or alter the meaning of the original literature, the meaning of a correctly cited item is compared with the meaning of the original literature, based on the established existence of the literature, to determine whether the citation field is correctly cited in terms of content. Of course, any method from related technologies can be used to compare the meaning of a correctly cited item with the meaning of the original literature; this application does not limit this approach.
[0055] It is important to distinguish that a correct citation only indicates that the cited literature actually exists, but the accuracy of the content is uncertain. Only by verifying the content based on a correct citation can we obtain a correct citation that is true and accurate in both citation and content.
[0056] B3. Determine the citation accuracy rate based on the correct citations and the number of cited references for each content.
[0057] The citation accuracy rate can be the accuracy rate of citations in the text being evaluated. In this embodiment, the ratio of the number of unique references in correctly cited entries to the total number of cited references can be used as the citation accuracy rate. In other words, the proportion of all genuine and correctly cited references to the total number of cited references can be considered the citation accuracy rate. For example, a large language model outputs a text with 20 citations (i.e., citation counts). A search in a literature database reveals 18 genuine references, one of which is cited three times. Therefore, a total of 16 different references are cited (i.e., the number of cited references). Further content comparison is performed. It's important to note that repeatedly cited references can be considered correct or incorrect citations depending on the accuracy of the content cited multiple times. Continuing the previous example, if a reference is cited three times, and all three citations are correctly detected, the citation of that reference in the text is considered a correct citation; if any one of them is an incorrect citation, the citation of that reference in the text is considered an incorrect citation. Of course, the definition can be broadened appropriately. If a document is cited multiple times and the content is correctly cited a significant number of times, then it can be considered a correct citation of that document. Continuing the previous example, if a document is cited three times and the content detection result is correct twice, then this citation of that document in this text is considered a correct citation.
[0058] If we assume that 14 out of 16 references are correctly cited, then the citation accuracy of this text is 14 ÷ 16 × 100% = 87.5%.
[0059] B4. Determine the accuracy of citations based on the correct citations and the number of times each item is cited.
[0060] Citation accuracy can be the accuracy of citation behavior in the text to be evaluated. In this embodiment, the ratio of the number of correctly cited items to the number of citations can be used as citation accuracy. It should be noted that, unlike B3, the citation accuracy rate can examine repeated correctly cited items. Regardless of whether the same reference is cited repeatedly, the number of correctly cited items represents the number of times the citation behavior is correct.
[0061] Building on the previous example, if 18 out of 20 citations in the text output by the large language model are from real sources (there may be cases of repeated citations of the same source), and 17 of these are correct citations, then the citation accuracy rate in the text is 17 ÷ 20 × 100% = 85%.
[0062] In the above implementation method, by analyzing the correct citation items of literature and the correct citation items of content, the citation accuracy and reference accuracy in the text to be evaluated are further determined, providing more dimensions for subsequent evaluation of the text and helping to improve the accuracy of the evaluation of academic texts output by the large language model.
[0063] In a further optional implementation, the step B1 of searching a preset document database based on each document identification number to determine the correct document citation in each citation field may include:
[0064] C1. For any document identification number, search the document database according to the document identification number. If the document metadata corresponding to the document identification number exists in the document database, then the citation field corresponding to the document identification number is determined to be a correct citation item of the document.
[0065] Document metadata can be a series of standardized information used to uniquely identify, describe, and locate a document. It is the core content of citation and reference lists, and may include, but is not limited to, author, title, abstract, publication name, and publication date. This application's implementation does not exhaustively list these. If a retrieval using the document identification number in the document database yields the corresponding document metadata, then the citation field is considered a correct citation.
[0066] C2. If a document identification number does not exist in the document database, or if the document metadata corresponding to the document identification number does not exist, then the citation field corresponding to the document identification number will be determined as a fictitious citation item.
[0067] On the other hand, fictitious citations can be citation fields that reference fictitious sources. If the document identification number cannot be found in the document database, or if the corresponding document metadata does not exist, it is due to the illusion problem of the large language model, causing fictitious references in the output text. Therefore, when a citation field cannot find a corresponding reference, that citation field is determined to be a fictitious citation.
[0068] In the above embodiments, the correctness of the citation of the literature is determined. By identifying and verifying the literature metadata, it is determined whether the reference corresponding to the citation field exists. This provides a practical solution for distinguishing between correct and fictitious citations in the literature in this application, provides a basis for the evaluation of the text to be evaluated, and helps to improve the efficiency of text evaluation.
[0069] In another optional implementation, the citation attribute further includes a fictitious document rate; the step of determining at least one citation attribute of the text to be evaluated based on each citation field and the corresponding document identification number in S130 may include: determining the fictitious document rate based on the fictitious citation items and the number of cited documents.
[0070] In this implementation, based on the aforementioned method, the fictitious citation item characterizes the phenomenon of false references in the citation field, and also reflects the illusion of the large language model. Therefore, the ratio of the number of unique references corresponding to fictitious citation items to the total number of cited references is used as the fictitious citation rate.
[0071] In the above implementation, the fictitious citation rate is determined by identifying fictitious citations and the number of cited references, providing another dimension of basis for subsequent text evaluation, enriching the perspective of text evaluation, and helping to improve the accuracy of text evaluation.
[0072] Similarly, another evaluation dimension can be added: fictitious citation rate. This fictitious citation rate can be the ratio of the number of fictitious citations to the total number of citations in the aforementioned literature. This fictitious citation rate reflects, to some extent, the degree of fictitious citation behavior in the large language model.
[0073] In another optional implementation, the step B2, which determines the correct citation of content based on the correct citation of each document and the corresponding cited document, may include: comparing the semantics of the correct citation of the document and the corresponding cited document based on a preset semantic discrimination model; if the comparison result is that the semantics are the same, then the correct citation of the document is determined to be the correct citation of content.
[0074] Semantic discrimination models, which are used to understand, distinguish, or judge the semantic information of text, are an important type of model in natural language processing. They determine the semantics of text by classifying, matching, judging, or scoring the input text. Any semantic discrimination model from related technologies can be used in this application embodiment, and this application embodiment does not impose any limitations. The semantics of correctly cited items and their corresponding references are identified and compared. If the semantics are the same, it indicates that they are consistent in content meaning, thus proving that the content citation of the correctly cited item is also correct. Only when a genuine document is cited, its content is understood, and the meaning expressed by the cited text is the same as that of the original document, can it be considered a successful citation.
[0075] In the above implementation, the semantic discrimination model identifies whether the semantics between the cited field and the reference are consistent, providing a method and basis for judging whether the cited field belongs to the correct citation of the content, which effectively improves the efficiency of identifying the correct citation of the content, and thus helps to improve the efficiency of text evaluation.
[0076] In another optional implementation, determining the text attributes of the text to be evaluated based on the text to be evaluated and each reference field in S130 may include: determining the number of text characters and the number of reference characters of the text to be evaluated based on the text to be evaluated and each reference field; and using the number of text characters and the number of reference characters as text attributes.
[0077] Both text word count and citation word count can exist as text attributes. Text word count can be the total number of words in the text output by a large language model, while citation word count can be the total number of words in the cited fields of the text. Furthermore, the ratio of citation word count to text word count can even be used as an indicator of citation density or citation richness to participate in text evaluation, which is the same as the way different dimensions of indicators are used in the evaluation in the aforementioned embodiments and various implementations. This application embodiment does not limit this approach.
[0078] Example 2
[0079] Figure 2 This is a schematic diagram of a text evaluation process provided in Embodiment 2 of this application. This embodiment is a specific example provided based on the foregoing embodiments and implementation methods. Figure 2 As shown, the details are as follows:
[0080] First, the academic texts generated by the large language model are formatted.
[0081] To address the issue of inconsistent formats generated by large language models, the system first parses the academic texts to be evaluated.
[0082] First, identifier recognition is performed, using regular expressions to identify all reference citation marks in the text (such as "[1]", "(Author,Year)", etc.).
[0083] Then, the IDs (Identity Documents) are mapped and standardized: all citation markers are replaced with a uniform format "Cited Content [XXX]". Here, "Cited Content" is one or more descriptive sentences preceding the citation marker. "[XXX]" is the unique identifier ID of the document in standard academic databases (such as DOI, arXiv, Web of Science, or self-built knowledge bases), which is the document identification number in the aforementioned example. [XXX] can contain multiple documents. Finally, a structured intermediate text is generated, providing a standardized basis for subsequent work.
[0084] Second, extract entities from the referenced fields.
[0085] Traverse the formatted text and extract all (cited content, citation ID) pairs one by one: The system constructs a list L=[(C_1,ID_1),(C_2,ID_2),...,(C_n,ID_n)], where C_i represents the text content corresponding to the i-th citation, and ID_i represents the document ID pointed to by the i-th citation.
[0086] Third, verify the authenticity of the documents (that is, detect the illusion of the large language model).
[0087] Use the extracted [XXX] (i.e., ID) to perform a search in the preset database.
[0088] If the database contains metadata (title, author, abstract, etc.) corresponding to the ID, the ID is marked as a "genuine document"; if the ID is not found in the database, it is marked as a "fake document" (i.e., a severe illusion generated by the large language model). Fake documents are also filtered out. If a document is determined to be a "fake document," the citation field is directly judged as "incorrect" before subsequent content verification and is not included in the semantic discrimination model to save computational resources. Finally, the number of fictitious documents and the number of incorrectly cited IDs are counted.
[0089] Fourth, semantic consistency judgment (i.e., verification of referenced content).
[0090] For the "real documents" identified in the above verification, a pre-defined semantic discrimination model is further invoked for content comparison. The "cited content" and the abstract or full-text excerpt of the real document corresponding to the ID mentioned above are input as context. The semantic discrimination model performs inference to determine whether the "cited content" is supported by the actual content of the reference. If the cited field faithfully reflects the original text's viewpoint, it is determined to be a correctly cited item. If the cited field misinterprets the original text, or if the original text does not mention the content, it is determined to be an incorrectly cited item.
[0091] Fifth, conduct text evaluation based on the multi-dimensional indicators mentioned above.
[0092] Based on the results of the above steps, the following five core indicators are calculated to construct a multi-dimensional quality profile of the evaluation text:
[0093] 1) Text word count (W): Count the total number of words in the generated text to assess the richness of the generated content.
[0094] 2) Number of cited references (N_cite): The number of unique occurrences of [XXX] in the text after deduplication. It reflects the breadth of knowledge referenced when amplifying the language model to generate content.
[0095] 3) Citation Count (N_ref): The total number of times [XXX] appears in the text (without deduplication). It reflects the frequency of the model's academic argumentation.
[0096] 4) Citation accuracy (Acc_cite): Acc_cite = Number of IDs judged as "correct" after deduplication / Number of cited references N_cite.
[0097] It should be noted that if a document is cited 3 times, with 2 correct and 1 incorrect, the calculation of this metric should be based on either a strict mode (any one incorrect citation is considered incorrect) or a lenient mode (most correct citations are considered correct). This scheme prefers the strict mode: that is, all citations corresponding to the ID must be "correct" for the citation to be considered correct.
[0098] 5) Reference accuracy (Acc_ref): Acc_ref = Total number of reference pairs judged as "correct" / Number of references N_ref.
[0099] 6) Calculate the fictitious literature rate: number of fictitious literature / number of cited literature.
[0100] Sixth, generate an evaluation report based on the text.
[0101] Based on the above calculation results, a structured evaluation report is generated:
[0102] 1) Basic statistics: display word count, citation density, and number of fictitious references.
[0103] 2) Citation quality score: a weighted score based on citation accuracy and reference accuracy. The weights can be pre-set by relevant technical personnel based on a large number of experiments or human experience.
[0104] 3) Error Analysis: Listing specific "fake literature" and "citation mismatch segments" helps users adjust the parameters of the large language model or manually polish the text.
[0105] Example 3
[0106] Figure 3 This is a schematic diagram of the structure of a text evaluation device provided in Embodiment 3 of this application. Figure 3 As shown, the device 300 includes:
[0107] The text acquisition module 310 is used to acquire the text to be evaluated.
[0108] The field determination module 320 is used to determine all the cited fields in the text to be evaluated and the corresponding document identification number of each cited field based on the citation marks in the text to be evaluated.
[0109] The attribute determination module 330 is used to determine at least one text attribute and at least one citation attribute of the text to be evaluated based on each citation field and the corresponding document identification number.
[0110] The attribute evaluation module 340 is used to evaluate the text to be evaluated based on each text attribute and each reference attribute.
[0111] In the technical solution of this application embodiment, the citation markers in the text to be evaluated generated by the large language model are identified, and all citation fields and their corresponding document identification numbers are determined in the text to be evaluated. This allows for the rapid location of the citation fields and their specific content within the text to be evaluated, and the rapid retrieval of the corresponding document identification numbers for determining citation attributes. Based on each citation field and its corresponding document identification number, at least one text attribute and at least one citation attribute of the text to be evaluated are determined. The analysis of the text to be evaluated yields text attributes and citation attributes, providing a basis for text evaluation. Different text attributes and different citation attributes can characterize the quality of the text to be evaluated from different dimensions, comprehensively evaluating the text and improving the accuracy of text evaluation. For academic texts generated by the large language model, this also helps adjust the parameters of the large language model based on the quality feedback of the generated text, facilitating the iterative evolution of the large language model.
[0112] In one optional implementation, the citation attribute includes the number of cited documents and the number of citations; the attribute determination module 330 may include:
[0113] The citation count determination unit is used to determine the number of citations in the text to be evaluated based on each citation field.
[0114] The citation count determination unit is used to perform deduplication based on the document identification number and determine the number of cited documents in the text to be evaluated.
[0115] In one optional implementation, the citation attribute further includes citation accuracy and citation correctness; the attribute determination module 330 may include:
[0116] The unit for determining correct citations is used to search a pre-set literature database based on each literature identification number to determine the correct citations in each citation field.
[0117] The correct citation item determination unit is used to determine the correct citation item based on the correct citation items of each document and the corresponding cited documents;
[0118] The citation accuracy determination unit is used to determine the citation accuracy based on the number of correctly cited items and cited references for each content.
[0119] The citation accuracy determination unit is used to determine the citation accuracy based on the correct citation items and the number of citations for each content item.
[0120] In one alternative implementation, the correct citation determination unit may include:
[0121] The correct citation item determination sub-unit is used to search the document database according to any document identification number. If the document metadata corresponding to the document identification number exists in the document database, the citation field corresponding to the document identification number is determined to be the correct citation item of the document.
[0122] The fictitious citation item determination subunit is used to determine the citation field corresponding to the document identification number as a fictitious citation item if the document identification number does not exist in the document database, or if the document metadata corresponding to the document identification number does not exist.
[0123] In one alternative implementation, the citation attribute further includes a fictitious document rate; the attribute determination module 330 may include:
[0124] The fictitious reference rate determination unit is used to determine the fictitious reference rate based on fictitious citations and the number of cited references.
[0125] In one optional implementation, the content correct reference determination unit may be specifically used for:
[0126] Based on a pre-defined semantic discrimination model, the semantics of the correct citation in a document are compared with those of the corresponding cited document. If the comparison results show that the semantics are the same, the correct citation in the document is determined to be the correct citation in the content.
[0127] In one alternative implementation, the attribute determination module 330 may include:
[0128] The quotation word count determination unit is used to determine the text word count and quotation word count of the text to be evaluated based on the text to be evaluated and each quotation field;
[0129] The text attribute determination unit is used to treat the number of text words and the number of quoted words as text attributes.
[0130] The text evaluation apparatus provided in this application embodiment can execute the text evaluation method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing each text evaluation method.
[0131] Example 4
[0132] Figure 3 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0133] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0134] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0135] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as text evaluation methods.
[0136] In some embodiments, the text evaluation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the text evaluation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the text evaluation method by any other suitable means (e.g., by means of firmware).
[0137] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0139] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0141] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0142] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0143] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the text evaluation method provided in any embodiment of this application. This program product and the text evaluation methods disclosed in the embodiments of this application belong to the same inventive concept, and therefore will not be described in detail here.
[0144] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0145] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A text evaluation method, characterized in that, include: Obtain the text to be evaluated generated by the large language model; Based on the citation markers in the text to be evaluated, determine all citation fields in the text to be evaluated and the corresponding document identification number for each citation field; Based on each of the cited fields and the corresponding document identification number, at least one text attribute and at least one citation attribute of the text to be evaluated are determined; wherein, the citation attribute includes the number of cited documents and the number of citations; the number of citations is the number of cited fields; wherein, if the same document is cited repeatedly, the number of citations is accumulated; The text to be evaluated is evaluated based on each of the stated text attributes and each of the stated reference attributes; The citation attributes also include citation accuracy and citation accuracy. The step of determining at least one citation attribute of the text to be evaluated based on each of the citation fields and the corresponding document identification number includes: Based on the document identification number, a search is performed in the preset document database to determine the correct document citation in each of the citation fields; Based on the correct citations of each document and their corresponding references, determine the correct citations for the content; The citation accuracy rate is determined based on the correct citations of each item and the number of cited references. The citation accuracy rate is determined based on the correct citation items and the number of citations for each item. The step of searching a preset document database based on each document identification number to determine the correct document citation in each citation field includes: For any of the document identification numbers, a search is performed in the document database according to the document identification number. If the document metadata corresponding to the document identification number exists in the document database, the citation field corresponding to the document identification number is determined to be the correct citation item of the document. If the document identification number does not exist in the document database, or if the document metadata corresponding to the document identification number does not exist, then the citation field corresponding to the document identification number will be determined as a fictitious citation item. The citation attribute also includes the fictitious reference rate; The step of determining at least one citation attribute of the text to be evaluated based on each of the citation fields and the corresponding document identification number includes: The fictitious citation rate is determined based on the fictitious citations in the literature and the number of cited literatures.
2. The method according to claim 1, characterized in that, The step of determining at least one citation attribute of the text to be evaluated based on each of the citation fields and the corresponding document identification number includes: Based on each of the cited fields, determine the number of citations in the text to be evaluated; The number of cited references in the text to be evaluated is determined by performing a plagiarism check based on the document identification number.
3. The method according to claim 1, characterized in that, The process of determining the correct citations based on the correct citations of each document and the corresponding cited documents includes: Based on a preset semantic discrimination model, the semantics of the correct citation item in the document are compared with those of the corresponding cited document. If the comparison result shows that the semantics are the same, the correct citation item in the document is determined to be the correct citation item in the content.
4. The method according to claim 1, characterized in that, The step of determining the text attributes of the text to be evaluated based on the text to be evaluated and each of the reference fields includes: Based on the text to be evaluated and each of the cited fields, determine the number of text characters and the number of cited characters in the text to be evaluated; The number of words in the text and the number of words in the quoted text are used as the text attributes.
5. A text evaluation device, characterized in that, include: The text acquisition module is used to acquire the text to be evaluated generated by the large language model; The field determination module is used to determine all the reference fields in the text to be evaluated and the document identification number corresponding to each reference field based on the reference marks in the text to be evaluated. The attribute determination module is used to determine at least one text attribute and at least one citation attribute of the text to be evaluated based on each of the citation fields and the corresponding document identification number; wherein, the citation attribute includes the number of cited documents and the number of citations; the number of citations is the number of citation fields; wherein, if the same document is cited repeatedly, the number of citations is accumulated. The attribute evaluation module is used to evaluate the text to be evaluated based on each of the text attributes and each of the reference attributes. The citation attributes further include citation accuracy and citation correctness; the attribute determination module includes: The document correct citation determination unit is used to search in a preset document database based on each document identification number to determine the correct document citation in each citation field; The correct citation item determination unit is used to determine the correct citation item based on the correct citation items of each document and the corresponding cited documents; The citation accuracy determination unit is used to determine the citation accuracy based on the correct citation items of each of the contents and the number of cited references; The citation accuracy determination unit is used to determine the citation accuracy based on the correctly cited items of each of the contents and the number of citations. The document correct citation determination unit includes: The correct citation item determination subunit is used to search the document database according to any of the document identification numbers. If the document metadata corresponding to the document identification number exists in the document database, the citation field corresponding to the document identification number is determined as the correct citation item of the document. The fictitious citation item determination subunit is used to determine the citation field corresponding to the document identification number as a fictitious citation item if the document identification number does not exist in the document database, or if the document metadata corresponding to the document identification number does not exist. The citation attribute further includes the fictitious document rate; the attribute determination module includes: The fictitious document rate determination unit is used to determine the fictitious document rate based on the fictitious citations in the document and the number of cited documents.
6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the text evaluation method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the text evaluation method of any one of claims 1-4.