Text intelligent calibration method and system based on large language model

Multi-dimensional intelligent text proofreading using a large language model solves the problems of insufficient accuracy and adaptability in text proofreading in existing technologies, achieves efficient and accurate text proofreading, adapts to text proofreading in different fields and styles, and improves text quality.

CN120805898APending Publication Date: 2025-10-17苏州日报社
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511013409.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing proofreading tools based on rule matching or shallow semantic analysis have difficulty identifying errors in complex contexts, resulting in poor accuracy and adaptability of text proofreading, especially in the expression of policy-related normative texts.

Method used

It adopts an intelligent text proofreading method based on a large language model, through a multi-dimensional proofreading process, including word-level spelling correction, sensitive word correction, and sentence-level semantic expression, logical structure and professional terminology correction, combined with self-built professional vocabulary and association network for error correction, to adapt to text proofreading in different fields and styles.

Benefits of technology

It significantly improves the accuracy and adaptability of text proofreading, ensures that the text meets industry standards and professional requirements, improves proofreading efficiency, and achieves efficient and high-quality text proofreading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805898A_ABST
    Figure CN120805898A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent text calibration method and system based on a large language model, and is used for solving the technical problem that the accuracy and adaptability of text calibration are poor. According to the intelligent text calibration scheme based on the large language model, through preliminary calibration of parallel processing, the linear mode of a traditional calibration process is broken through, and the calibration efficiency is greatly improved. And the text is comprehensively and meticulously proofread from multiple dimensions such as semantic expression, logic structure and terminology by using the large language model, so that the quality of the text in the professional field can be remarkably improved, the text is ensured to meet the industrial standard and professional requirement, and the accuracy and adaptability of text calibration are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of text proofreading, and in particular to a text intelligent proofreading method and system based on a large language model. BACKGROUND

[0002] With the rapid development of Internet technology, text information has become an important carrier for people to communicate, acquire knowledge, and relax. However, in the process of text creation, editing, and publishing, especially in the process of news writing in the news media industry, various language errors such as spelling errors, grammatical errors, and semantic errors inevitably occur. These errors not only affect the quality of the text and reduce user experience, but also may cause unnecessary misunderstandings and confusion. Therefore, it is particularly important to control these text errors in the news dissemination industry. Therefore, text proofreading tools have a wide range of application needs in the news media industry.

[0003] In the process of implementing the prior art, the inventors have found that:

[0004] Existing proofreading tools are often based on rule matching or shallow semantic analysis, which is difficult to identify errors in complex contexts, resulting in certain errors in the proofreading results. The adaptability of individualized proofreading for texts in different fields and styles is poor, especially for the proofreading of some policy type standard writing expressions, and the proofreading effect is usually unsatisfactory.

[0005] Therefore, it is necessary to provide a text intelligent proofreading scheme based on a large language model to solve the technical problem of poor accuracy and adaptability of text proofreading. SUMMARY

[0006] The embodiments of the present application provide a text intelligent proofreading method and system based on a large language model to solve the technical problem of poor accuracy and adaptability of text proofreading.

[0007] Specifically, a text intelligent proofreading method based on a large language model includes the following steps:

[0008] Obtain text data to be proofread;

[0009] According to the text data to be proofread, a first dimension proofreading is performed according to a first proofreading rule to generate a first proofreading result;

[0010] According to the text data to be proofread, a second dimension proofreading is performed according to a second proofreading rule to generate a second proofreading result;

[0011] Based on the first proofreading result or the second proofreading result, a preliminary proofreading correction prompt is generated;

[0012] Obtain preliminary proofreading text data updated based on the preliminary proofreading correction prompt;

[0013] According to the preliminary inspection and correction text data, the third inspection and correction rule is used for comprehensive dimension inspection and correction to generate a third inspection and correction result;

[0014] Based on the third inspection and correction result, a comprehensive correction prompt is generated;

[0015] Among them, the first dimension inspection and correction of the first inspection and correction rule according to the to-be-inspected text data generates a first inspection and correction result, including the following steps:

[0016] According to the to-be-inspected text data, word object recognition is performed to generate a word recognition result;

[0017] For the word recognition result, a wrong word inspection and correction is performed to determine the error word object and the correction word object corresponding to the error word object as the first inspection and correction result;

[0018] The second dimension inspection and correction of the second inspection and correction rule according to the to-be-inspected text data generates a second inspection and correction result, including the following steps:

[0019] According to the to-be-inspected text data, word object recognition is performed to generate a word recognition result;

[0020] For the word recognition result, a sensitive word inspection and correction is performed to determine the sensitive word object and the correction word object corresponding to the sensitive word object as the second inspection and correction result;

[0021] The third inspection and correction rule is used for comprehensive dimension inspection and correction according to the preliminary inspection and correction text data to generate a third inspection and correction result, specifically including:

[0022] According to the preliminary inspection and correction text data, sentence object recognition is performed to generate a sentence recognition result;

[0023] For the sentence recognition result, at least one of semantic expression inspection and correction, logical structure inspection and correction, or professional term inspection and correction is performed to determine the problem sentence object as the problem area;

[0024] Based on the association network of a plurality of problem areas, at least one of the semantic expression correction, the logical structure correction, and the professional term correction is matched to the correction sentence object as the third inspection and correction result.

[0025] Further, the method further includes:

[0026] Identify the category attribute of the preliminary inspection and correction text data;

[0027] Based on the category attribute of the preliminary inspection and correction text data, the third inspection and correction rule corresponding to the category attribute is matched;

[0028] Based on the third inspection and correction rule, the language model is controlled to perform comprehensive dimension inspection and correction to generate a third inspection and correction result;

[0029] The category attribute of the preliminary inspection and correction text data at least includes at least one of the following categories: current affairs news, social news, cultural education, medical health, environment and nature, sports news, economic news, and entertainment news.

[0030] The third inspection and correction rule is an instruction template, and at least includes at least one of the following instruction templates: system instruction, scene prompt instruction, role prompt instruction, ability prompt instruction, target prompt instruction, restriction rule instruction, output rule instruction, workflow instruction, and example, which are used to control the language model output.

[0031] Further, the obtaining of the to-be-inspected text data specifically includes:

[0032] Obtaining text data of a text editing area or text data of a cursor selected area as original text data;

[0033] Performing data cleaning, word segmentation, and part-of-speech tagging on the original text data to generate the to-be-inspected text data.

[0034] Further, the method further includes:

[0035] Based on an association network of a plurality of problem areas and corresponding context sentences of the problem areas, performing at least one of the following corrections: semantic expression correction, logical structure correction, and professional term correction, as the third inspection and correction result.

[0036] Further, the method further includes:

[0037] Based on the error word object or the sensitive word object, identifying an error type corresponding to the word object;

[0038] Determining a context word object of the error word object or the sensitive word object;

[0039] According to the error type and the context word object, determining a plurality of correction word objects;

[0040] Evaluating the plurality of correction word objects, and selecting a best correction word object as the preliminary inspection and correction prompt through a sorting algorithm.

[0041] Further, the method further includes:

[0042] Based on the problem area, identifying an error type corresponding to the problem area;

[0043] Determining a context sentence object of the problem area;

[0044] According to the error type and the context sentence object, determining a plurality of correction sentence objects;

[0045] Evaluating the plurality of correction sentence objects, and selecting a best correction sentence object as the comprehensive correction prompt through a sorting algorithm.

[0046] Further, the method further comprises:

[0047] Obtaining first-round proofreading text data updated based on comprehensive correction prompts;

[0048] According to the first-round proofreading text data, performing sentence object recognition to generate a sentence recognition result;

[0049] For the sentence recognition result, at least one of semantic expression proofreading, logical structure proofreading, or professional term proofreading is performed to determine a question sentence object as a question area;

[0050] Based on the association network of a plurality of question areas and corresponding question area context sentences, at least one of semantic expression correction, logical structure correction, and professional term correction is performed as a fourth proofreading result;

[0051] Based on the fourth proofreading result, a plurality of comprehensive correction prompts are generated.

[0052] The embodiments of the present application also provide a text intelligent proofreading system based on a large language model.

[0053] Specifically, a text intelligent proofreading system based on a large language model comprises:

[0054] A text editing module is configured to obtain to-be-proofread text data;

[0055] A content proofreading module is configured to perform first-dimension proofreading according to the to-be-proofread text data based on a first proofreading rule to generate a first proofreading result, and is further configured to perform second-dimension proofreading according to the to-be-proofread text data based on a second proofreading rule to generate a second proofreading result;

[0056] An error correction prompt module is configured to generate a preliminary proofreading correction prompt based on the first proofreading result or the second proofreading result;

[0057] The text editing module is further configured to obtain preliminary proofreading text data updated based on the preliminary proofreading correction prompt according to user feedback;

[0058] The content proofreading module is further configured to perform comprehensive-dimension proofreading according to the preliminary proofreading text data based on a third proofreading rule to generate a third proofreading result;

[0059] The error correction prompt module is further configured to generate a comprehensive correction prompt based on the third proofreading result;

[0060] The content proofreading module is configured to perform first-dimension proofreading according to the to-be-proofread text data based on a first proofreading rule to generate a first proofreading result, and comprises the following steps:

[0061] According to the to-be-proofread text data, performing word object recognition to generate a word recognition result;

[0062] The error word object and the correction word object corresponding to the error word object are determined as the first correction result through the error word correction based on the word recognition result.

[0063] The content correction module is configured to perform second-dimension correction on the to-be-corrected text data according to a second correction rule to generate a second correction result, including the following steps:

[0064] The word recognition result is generated through word object recognition based on the to-be-corrected text data.

[0065] The sensitive word object and the correction word object corresponding to the sensitive word object are determined as the second correction result through the sensitive word correction based on the word recognition result.

[0066] The content correction module is configured to perform comprehensive-dimension correction on the preliminary correction text data according to a third correction rule to generate a third correction result, specifically including:

[0067] The sentence recognition result is generated through sentence object recognition based on the preliminary correction text data.

[0068] The problem sentence object is determined as the problem area through at least one of semantic expression correction, logic structure correction, or professional term correction based on the sentence recognition result.

[0069] At least one correction sentence object is matched based on the association network of a plurality of problem areas, and the correction sentence object is semantic expression correction, logic structure correction, or professional term correction, as the third correction result.

[0070] Further, the error correction prompt module is configured to:

[0071] The preliminary correction prompt or the comprehensive correction prompt is displayed.

[0072] The operation instruction of the user for the preliminary correction prompt or the comprehensive correction prompt is collected.

[0073] The operation instruction includes one-by-one adoption of the correction word object or the correction sentence object.

[0074] Or the operation instruction includes all adoption of the correction word object or the correction sentence object.

[0075] Further, the text editing module is configured to collect the adoption feedback of the user for the comprehensive correction prompt, and perform multiple rounds of correction on the correction text data.

[0076] The text editing module is further configured to obtain the first-round correction text data updated based on the comprehensive correction prompt.

[0077] The content correction module is further configured to perform sentence object recognition based on the first-round corrected text data, and generate a sentence recognition result; further configured to perform at least one of semantic expression correction, logical structure correction, or professional term correction on the sentence recognition result, determine a question sentence object as a question area; and further configured to perform at least one of semantic expression correction, logical structure correction, or professional term correction based on an association network of a plurality of question areas and corresponding context sentences on the question areas, and generate a fourth correction result.

[0078] The error correction prompt module is further configured to generate a multi-round comprehensive correction prompt based on the fourth correction result.

[0079] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:

[0080] Through the preliminary correction by parallel processing, the linear mode of the traditional correction process is broken, and the correction efficiency is greatly improved. By using a large language model to correct the text from multiple dimensions such as semantic expression, logical structure, and professional terms, the quality of the text in the professional field can be significantly improved, ensuring that it meets the industry standards and professional requirements, and improving the accuracy and adaptability of the text correction. BRIEF DESCRIPTION OF DRAWINGS

[0081] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application and illustrate the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0082] Figure 1 A flowchart of a text intelligent correction method based on a large language model provided by the embodiments of the present application;

[0083] Figure 2 A flowchart of a text intelligent correction method based on a large language model provided by the embodiments of the present application;

[0084] Figure 3 A structural schematic diagram of a text intelligent correction system based on a large language model provided by the embodiments of the present application. DETAILED DESCRIPTION

[0085] To make the purpose, technical scheme and advantages of the present application clearer, the technical scheme of the present application will be described in detail below with reference to the embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0086] Please refer to Figure 1With Figure 2 To solve the technical problems of poor accuracy and adaptability of text correction, the present application provides a text intelligent correction method based on a large language model. Through comprehensive intelligent correction of wrong characters, self-built professional vocabulary correction, and general large model correction from multiple dimensions such as semantic expression, logical structure, and professional terms, the method realizes automatic correction of "word-sentence-paragraph" four-level content, provides a systematic and automated solution for media content "three corrections and three corrections", and facilitates efficient and high-quality operation to improve user experience.

[0087] The specific implementation process of the intelligent correction method is introduced as follows:

[0088] S110: Obtain text data to be corrected.

[0089] It can be understood that the text data to be corrected is usually the first draft in daily text editing work. Usually, the first draft has errors in vocabulary selection due to the influence of non-standard word use or expression, unfamiliarity with local characteristics or industry-specific language. Therefore, the text to be corrected often needs to be corrected in multiple dimensions.

[0090] Further, in a specific embodiment provided by the present application, the obtaining of the text data to be corrected specifically includes:

[0091] Obtaining text data in a text editing area or text data in a cursor selected area as original text data;

[0092] Performing data cleaning, word segmentation, and part-of-speech tagging on the original text data to generate text data to be corrected.

[0093] The text editing area here refers to the interface area in a text editor or word processing software where text can be directly input, modified, and viewed. The present application obtains text data in the text editing area or cursor selected area of the text editor as original text data.

[0094] After obtaining the original text data, the original text data needs to be preprocessed, such as data cleaning, word segmentation, and part-of-speech tagging, and the text data after preprocessing is taken as the text data to be corrected.

[0095] S120: According to the text data to be corrected, a first dimension correction is performed according to a first correction rule to generate a first correction result.

[0096] It should be noted that, in order to improve the quality of the text, the inventor synchronously corrects the text data in multiple dimensions. The present application is distinguished by first dimension correction and second dimension correction. In a specific application scenario, the first dimension correction can be a conventional error correction, and the first correction rule can be a conventional wrong word library. The conventional wrong word library can be further divided into a similar shape word sub-library, a homonym word sub-library, a fuzzy sound wrong word sub-library, and a small syllable segmentation wrong word sub-library. By comparing the conventional wrong word library, the error text in the text data to be corrected, the error type of the error text, and the correction text of the error text can be matched.

[0097] Specifically, the first dimension correction of the text data to be corrected according to the first correction rule generates a first correction result, including the following steps:

[0098] According to the text data to be corrected, a word object recognition is performed to generate a word recognition result.

[0099] For the word recognition result, a wrong word correction is performed to determine the error word object and the correction word object corresponding to the error word object as the first correction result.

[0100] It should be noted that the first dimension correction is a word-level correction, so the text data to be corrected needs to be subjected to word object recognition. The word object recognition can be a word segmentation process in preprocessing, or a word-level screening of a word segmentation set.

[0101] Further, in a specific embodiment provided by the present application, the method further comprises:

[0102] Based on the error word object, the error type corresponding to the word object is identified.

[0103] The context word object of the error word object is determined.

[0104] According to the error type and the context word object, a plurality of correction word objects are determined.

[0105] The plurality of correction word objects are evaluated, and the best correction word object is selected as the initial correction prompt by a sorting algorithm.

[0106] By combining the error type and the context word object information, the accuracy of the correction word object can be improved. Moreover, the present application can recommend the correction word object with the highest accuracy, or recommend several correction word objects with higher accuracy, so as to facilitate the user to select.

[0107] In the specific application scenario provided in the present application, step S120 specifically represents performing word-level correction on common wrong characters, alternative characters, etc. in the text editor. Then, a correction report is output to prompt the common word errors such as wrong characters and alternative characters in the text. The next step requires user intervention to view all prompt information of the conventional error correction, and the user can accept all prompted words one by one (i.e., replace the error words in the original text) or click the "accept all" button to modify all at once, and enter the "prompt whether to accept" link.

[0108] S130: performing second-dimension correction according to the text data to be corrected, to generate a second correction result.

[0109] It should be noted that the second-dimension correction here can be professional correction, and the second correction rule can be a self-built word library, which can cover multiple-dimension sensitive word libraries (such as special sensitive words, legal and regulatory sensitive words, etc.) and standard language libraries (such as special language, local unit common event standard expression, etc.). These word libraries are deeply customized and continuously updated and optimized according to the actual needs of a specific field (such as news reporting).

[0110] In one specific embodiment provided in the present application, the professional word library built by the inventors includes: a sensitive word library (including special sensitive words: words related to leaders, policies, historical events, special movements, etc. Legal and regulatory sensitive words: words related to illegal activities, legal sanctions, etc. Social stability sensitive words: words that may cause social unrest, group events, etc. National and religious sensitive words: words related to nationality, religious beliefs, etc. that may cause national and religious conflicts. Moral and ethical sensitive words: words related to ethics, unhealthy atmosphere, etc. that may affect social customs. Information security sensitive words: words related to state secrets, trade secrets, etc. that may endanger information security. Personal privacy sensitive words: words related to personal privacy information, such as ID number, home address, etc. Vulgar sensitive words: words related to vulgar content, etc.) and a standard language library (such as local common event standard expression, news standard language, local place name, and local place name and personal name standard language library).

[0111] Such a deep customized professional vocabulary can accurately detect the expressions in the text that are inconsistent with the professional field specifications. For example, using the vocabulary can effectively avoid legal, social and other disputes caused by inappropriate word use or non-standard expression, ensuring the rigor and professionalism of news reporting. At the same time, for some local characteristics or industry-specific language, accurate correction suggestions can be given according to the standard vocabulary, so that the text is more in line with professional standards and the reading habits of the audience. Specifically, the professional vocabulary built by the inventors can also include a local characteristic vocabulary. Here, by combining the special nature of news media and the local characteristics of a specific region, in addition to general sensitive words such as sensitive words, terrorism-related words, vulgar language, etc., sensitive words related to local elements such as the history and culture, social customs of a specific region are also specially integrated.

[0112] Specifically, the second-dimension correction according to the to-be-corrected text data and the second correction rule to generate a second correction result includes the following steps:

[0113] According to the to-be-corrected text data, a word object recognition is performed to generate a word recognition result.

[0114] For the word recognition result, a sensitive word correction is performed to determine a sensitive word object and a correct word object corresponding to the sensitive word object as the second correction result.

[0115] It should be noted that the second-dimension correction here is a word-level correction, so the to-be-corrected text data needs to be subjected to word object recognition. The word object recognition here can be a word segmentation process in preprocessing, or a word-level screening on a word segmentation set. In specific application scenarios, the present application will perform text detection on the to-be-corrected text data based on a sensitive word library and a news standard vocabulary library, respectively, to ensure the rigor and professionalism of news reporting.

[0116] Further, in a specific embodiment provided by the present application, the method further includes:

[0117] Based on the sensitive word object, an error type corresponding to the sensitive object is identified;

[0118] The context word object of the sensitive word object is determined;

[0119] According to the error type and the context word object, a plurality of correction word objects are determined;

[0120] The plurality of correction word objects are evaluated, and the best correction word object is selected as the preliminary correction prompt through a sorting algorithm.

[0121] By combining the error type and the information of the context word object, the accuracy of the corrected word object can be improved. Moreover, the application can recommend the corrected word object with the highest accuracy or recommend several corrected word objects with higher accuracy, so as to facilitate the user to select.

[0122] In the specific application scenario provided by the application, step S130 specifically represents text detection in two links of sensitive word detection and news specification user detection on the text in the text editor. Then, a self-built word library correction suggestion report is output to prompt the text inconsistent with the self-built word expression. Next, the editing user needs to intervene, view all prompt information of professional correction, and can adopt (i.e., replace the error words in the original text) one by one or click the "all adopt" button to modify all at once, and enter the "prompt whether to adopt" link.

[0123] S140: generating a preliminary inspection correction prompt based on the first inspection result or the second inspection result.

[0124] It can be understood that the first dimension inspection and the second dimension inspection are performed synchronously, and the preliminary inspection correction prompt can be any one of the first inspection result or the second inspection result, that is, any one of the general correction or the professional correction is given, so that the user can select according to the use demand. It can also be combined with the first inspection result and the second inspection result, that is, the general correction and the professional correction are combined as a comprehensive suggestion to improve the text quality.

[0125] S150: obtaining the preliminary inspection text data updated based on the preliminary inspection correction prompt.

[0126] It can be understood that the preliminary inspection text data updated based on the preliminary inspection correction prompt refers to the feedback operation or instruction of the user after the preliminary inspection correction prompt is given in step S140, for example, the user can partially adopt (i.e., replace the error words in the original text) or click the "all adopt" button to modify all at once, and enter the "prompt whether to adopt" link.

[0127] The application collects the operation feedback or instruction feedback of the user on the preliminary inspection correction prompt, and updates the text data to be inspected to the preliminary inspection text data based on the operation of the user.

[0128] S160: performing comprehensive dimension inspection according to the third inspection rule based on the preliminary inspection text data, to generate a third inspection result.

[0129] It should be noted that, in order to further improve the quality of the text in the professional field, the large language model is used for deep analysis and proofreading of the text. In a specific application scenario, the third proofreading rule here can be at least one of semantic expression proofreading, logical structure proofreading, or professional term proofreading. This multi-dimensional proofreading method can provide users with comprehensive reference, so that they can clearly understand various problems in the text and make targeted modifications according to the report.

[0130] It should be noted that the third proofreading rule can be at least one of semantic expression proofreading, logical structure proofreading, or professional term proofreading.

[0131] The following describes a relatively basic implementation, including the following steps:

[0132] According to the initial proofreading text data, a sentence object recognition is performed to generate a sentence recognition result;

[0133] According to the sentence recognition result, at least one of semantic expression proofreading, logical structure proofreading, or professional term proofreading is performed to determine a problem sentence object as a problem area.

[0134] Based on the association network of the problem areas, at least one of semantic expression correction, logical structure correction, or professional term correction is matched to a corrected sentence object as the third proofreading result.

[0135] It should be noted that the third dimension proofreading here is a sentence-level proofreading, so the sentence object recognition needs to be performed on the text data to be proofread.

[0136] Then, the semantic expression proofreading, logical structure proofreading, and professional term proofreading are performed on the sentence objects sentence by sentence. As long as any proofreading is incorrect, the sentence object is marked as a problem sentence object as a problem area. Therefore, the initial proofreading text data can have multiple problem areas. Although the multiple problem areas are scattered in different positions, they have an association network, i.e., a coherent relationship, because they have the same article theme or concept. The association network here represents at least one quantifiable, comparable, or similar relationship between any two problem areas in terms of semantic coherence, logical consistency, functional consistency, or rhetorical consistency.

[0137] According to the association network of the problem areas, at least one of semantic expression correction, logical structure correction, or professional term correction is matched to a corrected sentence object, which can make the corrected sentence objects corresponding to the multiple problem areas also have an association network, so as to generate a corrected text with higher quality.

[0138] Further, in a preferred embodiment provided by the present application, the method further includes:

[0139] Based on the association network of a plurality of problem areas and corresponding problem area context sentences, at least one of semantic expression correction, logical structure correction, and professional term correction is performed as the third proofreading result.

[0140] In addition to correcting the problem areas based on the association network of a plurality of problem areas, the association network of a plurality of problem areas and corresponding problem area context sentences is introduced, that is, the full-text coherent relationship is used to correct the problem areas, to ensure the logical structure coherence, smooth connection, and natural fluency between sentences.

[0141] Further, in another embodiment provided by the present application, the method further comprises:

[0142] Based on the problem area, an error type corresponding to the problem area is identified;

[0143] A context sentence object of the problem area is determined;

[0144] Based on the error type and the context sentence object, a plurality of corrected sentence objects are determined;

[0145] The plurality of corrected sentence objects are evaluated, and a best corrected sentence object is selected as a comprehensive correction prompt through a sorting algorithm.

[0146] By combining the information of the error type and the context word object, the accuracy of the corrected sentence object can be improved. Furthermore, the most accurate corrected sentence object can be recommended, or several corrected sentence objects with relatively high accuracy can be recommended, so as to facilitate the user to select.

[0147] Further, in a more specific embodiment provided by the present application, the method further comprises:

[0148] A category attribute of the initial proofreading text data is identified;

[0149] Based on the category attribute of the initial proofreading text data, a third proofreading rule corresponding to the category attribute is matched;

[0150] Based on the third proofreading rule, a comprehensive dimension proofreading is controlled by using the language model, and a third proofreading result is generated.

[0151] It should be noted that the text has a category attribute. According to different category attributes of the text, a corresponding large language model proofreading prompt word template is selected, so that the text can be more targeted for in-depth analysis and proofreading. Here, the third proofreading rule is a more detailed instruction template, which is used to improve the output quality of the language model through a plurality of prompt word instructions.

[0152] Specifically, the category attribute of the preliminary inspection and correction text data at least includes at least one of Political News, Social News, Culture and Education News, Health and Medical News, Environment and Nature News, Sports News, Economic News, and Entertainment News.

[0153] The third inspection and correction rule is an instruction template, at least including system instructions, scene prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and at least one of example instructions, used to control the language model output.

[0154] In the specific application scenarios provided in the present application, based on the category attribute of the preliminary inspection and correction text data, the third inspection and correction rule corresponding to the category attribute is matched, which specifically shows that:

[0155] For news text of different categories, the corresponding large language model inspection prompt word template is used for intelligent inspection and correction of the large language model.

[0156] For example, for the text with the classification result of "Political News", the system applies the prompt word Prompt template as follows:

[0157] - Role (role prompt instruction): professional text correction expert and experienced political news editor

[0158] - Background (scene prompt instruction): the user needs to correct the political news text, which involves international relations, policies and regulations, government dynamics, elections and special movements, etc. Professional fields require high accuracy of semantic expression, logical structure and professional terms. Any error may lead to misinformation and damage to professionalism.

[0159] - Profile (role prompt instruction): you are an experienced text correction expert with a deep editing background in the field of political news. You have a deep understanding and research of domestic and international dynamics, policies and regulations, and can accurately grasp the accuracy and standardization of professional terms. You are good at meticulous correction of text from multiple dimensions such as semantics, logic and professionalism.

[0160] - Skills (Ability Prompt Instructions): You possess acute linguistic perception, enabling you to quickly identify ambiguities and errors in semantic expression; you have a solid grasp of rigorous logical analysis methods, ensuring coherent and natural text structure; you have precise memory and understanding of professional terminology related to current affairs news, enabling accurate verification of the use of such terminology.

[0161] - Goals (Objective Prompt Instructions):

[0162] 1. Accurately understand and verify each sentence in the text, ensuring clear and error-free semantic expression.

[0163] 2. Check the logical structure of the text for coherence, ensuring natural and smooth transitions between sections.

[0164] 3. Strictly proofread professional terminology related to current affairs news, ensuring accuracy.

[0165] 4. Suggest multiple rounds of proofreading, focusing on different aspects each time, to ensure the highest text quality.

[0166] - Constrains (Restriction Rule Instructions): Strictly adhere to professional standards for current affairs news, ensuring accurate use of terminology; maintain objectivity and accuracy in the text, avoiding any potentially misleading expressions; during proofreading, focus on details and do not overlook any potential errors.

[0167] - OutputFormat (Output Rule Instructions): The output format is a detailed proofreading report, including semantic issues, logical structure issues, professional terminology issues, and improvement suggestions, along with examples of modified text.

[0168] - Workflow (Work Flow Instructions):

[0169] 1. Preliminary reading of the text, overall understanding of content and structure, and marking of potential problem areas.

[0170] 2. Sentence-by-sentence analysis of the text, verification of clear and accurate semantic expression, checking of coherent logical structure, and special attention to the accuracy of professional terminology related to current affairs news.

[0171] 3. Multiple rounds of proofreading, with the first round focusing on semantic and logical issues, the second round focusing on the accuracy of professional terminology, and the third round focusing on overall polishing and connection optimization.

[0172] 4. Output a detailed proofreading report and provide examples of modified text, ensuring that the text quality meets professional standards.

[0173] - Examples (Examples):

[0174] - Example 1: Original text: "At the recent international conference, representatives from various countries had in-depth discussions on global climate change and reached a preliminary consensus."

[0175] After correction: "At the recent international conference, representatives from various countries had in-depth discussions on global climate change and reached a consensus." (Modification reason: To make the semantics more accurate and avoid the ambiguity that "preliminary consensus" may bring)

[0176] - Example 2: Original text: "Our government has introduced a series of new economic policies aimed at promoting rapid economic growth."

[0177] After correction: "Our government has introduced a series of new economic policies aimed at promoting high-quality economic development." (Modification reason: To make the policy goal statement more in line with current policy direction and avoid the one-sidedness that "rapid growth" may bring)

[0178] - Example 3: Original text: "In this election, candidate A and candidate B engaged in fierce competition, and finally candidate A won by a narrow margin."

[0179] After correction: "In this election, candidate A and candidate B engaged in fierce competition, and finally candidate A won by a narrow margin." (Modification reason: To make the statement more standardized and in line with the professional language of political news)

[0180] For the text classified as the "medical health" field, the system applies the following prompt template:

[0181] - Role (role prompt instruction): Medical health field professional terminology correction expert and senior editor

[0182] - Background (scene prompt instruction): The user needs to perform text content correction on news articles in the medical health field, which involves public health, medical policy, new drug research and development, healthy lifestyle, disease prevention and control, and other professional fields. The accuracy of professional terms is crucial to improving the professionalism of the text. Any misuse of terminology may affect the accurate communication and professionalism of the information, and may even lead to misunderstanding or misdirection.

[0183] - Profile (role prompt instruction): You are a professional terminology correction expert and senior editor in the medical health field with deep experience, who has in-depth research and rich practical experience in professional terminology in the medical health field, can accurately identify and correct the non-standard places in the use of terminology, and ensure the professionalism and accuracy of the text.

[0184] - Skills (Ability Instructions): You have keen language perception abilities, enabling you to quickly identify whether the use of terminology is in line with professional standards; you are familiar with the latest developments and trends in the medical and health field and have precise memories and understanding of the formal usage of professional terminology; you are skilled at analyzing the context of terminology and determining whether it is appropriate for use in specific medical and health news.

[0185] - Goals (Objective Instructions):

[0186] 1. Ensure the accuracy of professional terminology in the medical and health field.

[0187] 2. Provide official definitions and usage specifications for terminology.

[0188] 3. Determine whether the use of terminology is appropriate in the context of specific medical and health news.

[0189] 4. Provide improvement suggestions to ensure that the use of terminology meets the professional requirements of medical and health news.

[0190] - Constrains (Restriction Rules Instructions): Strictly follow the professional terminology usage specifications in the medical and health field to ensure the accuracy and authority of terminology; during proofreading, focus on the context and usage scenarios of terminology to avoid one-size-fits-all judgments; language expression should be concise and easy to understand.

[0191] - OutputFormat (Output Rules Instructions): The output format is a detailed list of terminology differentiation instructions, official definitions, usage specifications, and improvement suggestions, along with specific text examples.

[0192] - Workflow (Work Flow Instructions):

[0193] 1. Identify professional terminology in the text and mark possible problematic usage.

[0194] 2. Query the official definitions and usage specifications of the terminology to verify its formal usage.

[0195] 3. Determine whether the use of terminology is appropriate in the context of specific medical and health news and provide improvement suggestions.

[0196] - Examples (Examples):

[0197] - Example 1: Accuracy of Professional Terminology

[0198] - Incorrect Example: "This public health action aims to prevent the spread of 'influenza virus'." ("Influenza virus" is not a professional enough expression)

[0199] - Improvement suggestion: "This public health action aims to prevent the spread of 'influenza virus'." ("Influenza virus" is more appropriate in a professional context)

[0200] - Example 2: Term usage in context

[0201] - Error example: "New drug development has made'major breakthroughs'." ("Major breakthroughs" is not specific enough)

[0202] - Improvement suggestion: "New drug development has made'significant progress' in clinical trials." ("Significant progress" is more appropriate in a medical policy context)

[0203] - Example 3: Official definition of terms

[0204] - Error example: "A healthy lifestyle includes'reasonable diet' and'moderate exercise'." ("Reasonable diet" and "moderate exercise" are not professional enough)

[0205] - Improvement suggestion: "A healthy lifestyle includes 'balanced diet' and'regular exercise'." ("Balanced diet" and "regular exercise" are more appropriate in a healthy lifestyle context)

[0206] This dynamic adjustment of prompt word templates based on text categories allows the large language model to more accurately identify and correct potential problems in the text. For example, for political news texts, the system applies a prompt word template specifically for the political field, guiding the large language model to focus on the accuracy of special terms, the rigor of semantics, and the coherence of logic, etc. Key issues. This targeted correction method can significantly improve the quality of text in the professional field, ensuring that it meets industry standards and professional requirements, providing high-quality text content for users.

[0207] In the specific application scenarios provided in this application, based on the third correction rule, the language model is controlled to perform comprehensive dimensional correction to generate a third correction result, which is specifically manifested as:

[0208] The language model analyzes the initial correction text data sentence by sentence according to the instruction template. It performs sentence-by-sentence correction from three different dimensions: verifying semantic expression, checking logical structure, and correcting professional terms. Determine the problem sentence object and mark it as a problem area. Then match at least one correction sentence object for semantic expression correction, logical structure correction, and professional term correction in the problem area as the third correction result.

[0209] S170: Based on the third correction result, generate a comprehensive correction prompt.

[0210] It can be understood that the third calibration result can include error correction suggestions of multiple dimensions of a single problem area, so that the user can select according to the use demand. It can also include comprehensive error correction suggestions based on multiple dimensions of a single problem area to improve the quality of the text. The comprehensive correction prompt here can be a prompt for multiple problem areas in the text, a display of correction examples, and the user can partially adopt the modified sentence for all prompts (i.e. replace the incorrect sentence in the original text). It can also be a button that is clicked once to modify all, and then enter the "prompt whether to adopt" link.

[0211] Further, in a specific embodiment provided by the present application, the method further comprises:

[0212] Obtaining the first-round calibration text data updated based on the comprehensive correction prompt;

[0213] According to the first-round calibration text data, performing sentence object recognition to generate a sentence recognition result;

[0214] According to the sentence recognition result, performing at least one of semantic expression calibration, logical structure calibration, or professional term calibration to determine a problem sentence object as a problem area;

[0215] Based on the association network of a plurality of problem areas and corresponding context sentences of the problem areas, performing at least one of semantic expression correction, logical structure correction, and professional term correction as a fourth calibration result;

[0216] Based on the fourth calibration result, generating a plurality of rounds of comprehensive correction prompts.

[0217] It can be understood that the present application can perform multiple rounds of calibration on the text data to be calibrated, and focus on different calibration points, such as focusing on semantic and logical problems in the first round, focusing on the accuracy of professional terms in the second round, and performing overall polishing and connection optimization in the third round. Finally, a plurality of rounds of comprehensive correction prompts are output, which lists all problems and modification suggestions in detail, and provides a modified text example.

[0218] The multiple-round calibration mechanism can gradually and deeply excavate the problems in the text to ensure that each detail is fully reviewed and corrected. The multiple-round comprehensive correction prompt provides a comprehensive reference for the editor, enabling him to clearly understand the various problems in the text and make targeted modifications based on the report. In this way, the final quality of the text can be significantly improved to meet the standards of professional publishing or release, effectively reducing the risks that may be caused by text quality problems, such as misinformation and loss of professionalism.

[0219] In summary, the text intelligent proofreading method based on a large language model provided in the application breaks the linear mode of the traditional proofreading process through preliminary proofreading by parallel processing, greatly improving the proofreading efficiency. And using the large language model to proofread the text from multiple dimensions such as semantic expression, logical structure and professional terms can significantly improve the quality of the text in the professional field, ensure that it meets the industry standards and professional requirements, and improve the accuracy and adaptability of text proofreading.

[0220] Please refer to Figure 3 To support the text intelligent proofreading method based on a large language model, the application further provides a text intelligent proofreading system 100 based on a large language model.

[0221] It should be noted that the text intelligent proofreading system 100 based on a large language model integrates the proofreading mode of error detection, error classification and error correction of the text in one interface, realizes efficient and high-quality operation, and improves the user experience.

[0222] Specifically, a text intelligent proofreading system 100 based on a large language model comprises:

[0223] A text editing module 11 is configured to obtain text data to be proofread.

[0224] A content proofreading module 12 is configured to perform first-dimension proofreading on the text data to be proofread according to a first proofreading rule to generate a first proofreading result, and perform second-dimension proofreading on the text data to be proofread according to a second proofreading rule to generate a second proofreading result.

[0225] An error correction prompt module 13 is configured to generate a preliminary proofreading correction prompt based on the first proofreading result or the second proofreading result.

[0226] The text editing module 11 is further configured to obtain preliminary proofread text data updated based on the preliminary proofreading correction prompt according to user feedback.

[0227] The content proofreading module 12 is further configured to perform comprehensive-dimension proofreading on the preliminary proofread text data according to a third proofreading rule to generate a third proofreading result.

[0228] The error correction prompt module 13 is further configured to generate a comprehensive correction prompt based on the third proofreading result.

[0229] Specifically, the content proofreading module 12 is configured to perform first-dimension proofreading on the text data to be proofread according to a first proofreading rule to generate a first proofreading result, comprising the following steps:

[0230] According to the text data to be proofread, word object recognition is performed to generate a word recognition result.

[0231] The error word object and the corresponding correction word object of the error word object are determined as the first correction result according to the word recognition result.

[0232] The content correction module 12 is configured to perform second-dimension correction on the to-be-corrected text data according to a second correction rule to generate a second correction result, including the following steps:

[0233] The word object recognition is performed according to the to-be-corrected text data to generate a word recognition result.

[0234] The sensitive word correction is performed according to the word recognition result to determine the sensitive word object and the corresponding correction word object of the sensitive word object as the second correction result.

[0235] The content correction module 12 is configured to perform comprehensive-dimension correction on the preliminary correction text data according to a third correction rule to generate a third correction result, including the following steps:

[0236] The sentence object recognition is performed according to the preliminary correction text data to generate a sentence recognition result.

[0237] The semantic expression correction, the logic structure correction or the professional term correction are performed according to the sentence recognition result to determine the problem sentence object as the problem area.

[0238] The semantic expression correction, the logic structure correction or the professional term correction are performed according to the sentence recognition result to determine the problem sentence object as the problem area.

[0239] In a specific implementation provided in the application, the text editing module 11 acquires the text data of the text editing area or the text data of the cursor selected area as the original text data. Then, the original text data is subjected to data cleaning, word segmentation and part-of-speech tagging to generate the to-be-corrected text data.

[0240] The content correction module 12 performs word-level correction on the common wrong words and other wrong words in the text editor. Then, the error correction prompt module 13 outputs the correction report to prompt the common word errors such as wrong words and other wrong words in the text. In the next step, the user needs to intervene to check all the prompt information of the common error correction, and can accept all the prompted words one by one (i.e., replace the error words in the original text) or click the “all accept” button to modify all at once, and enter the “prompt whether to adopt” link.

[0241] In parallel, the content correction module 12 also performs text detection in two aspects: sensitive word detection and news standard user detection in the text editor. Then the error correction prompt module 13 outputs a self-built word library correction suggestion report to prompt the text that does not conform to the self-built word expression. The next step requires the editor user to intervene, view all prompt information for professional correction, and can accept all modified words one by one (i.e. replace the wrong words in the original text) or click the "all accept" button to modify all at once, and enter the "prompt whether to adopt" link.

[0242] The above is the preliminary correction of the text data to be corrected. Through parallel processing, the user can obtain two different correction reports at the same time, one for common word errors and the other for sensitive words and standard language problems in the professional field. This not only saves time, but also ensures the accuracy of the text in multiple dimensions. For example, in the news editing scenario, editors can quickly find and correct possible misuse of special sensitive words, improper expression of laws and regulations, and common misspelling, effectively improving the comprehensiveness and efficiency of text quality control.

[0243] After that, the text editing module 11 also collects the operation feedback or instruction feedback of the user for the preliminary correction correction prompt, and updates the text data to be corrected to the preliminary correction text data based on the user's operation.

[0244] The content correction module 12 identifies the category attribute of the preliminary correction text data. According to the category attribute of the preliminary correction text data, a corresponding large language model correction prompt template is selected. Then the language model analyzes the preliminary correction text data according to the instruction template. From three different dimensions of verifying semantic expression, checking logical structure, and correcting professional terms, the text is analyzed sentence by sentence. Determine the problem sentence object and mark it as a problem area. Then match at least one correction sentence object for semantic expression correction, logical structure correction, and professional term correction in the problem area.

[0245] The error correction prompt module 13 outputs a large language model correction suggestion report to prompt the text that has semantic expression, checking logical structure, and correcting professional terms errors in the text after the large language model analysis. The next step requires the editor user to intervene, view all prompt information for large language model correction, and can accept all modified words one by one (i.e. replace the wrong words in the original text) or click the "all accept" button to modify all at once, and enter the "prompt whether to adopt" link.

[0246] After the large language model analyzes the text task, the large language model will automatically perform multiple rounds of correction on the text again, including "first round: semantic and logical", "second round: professional terms", and "third round: polishing and cohesion" operations.

[0247] After the multiple rounds of proofreading are completed, the error correction prompt module 13 outputs a proofreading report, which is mainly a comprehensive suggestion report of the large language model after the above links, and gives the modification reason. The large language model intelligently lists all modification examples for the editor to refer to.

[0248] Next, in the text editing area, the editor can review each modification suggestion one by one, and can click “adopt” on each modification suggestion to directly adopt and replace the corresponding text in the editing area.

[0249] In addition, after the editor completes all the suggestions and replaces the text, the text intelligent proofreading system 100 based on the large language model can also perform a round of proofreading on the modified text from the beginning again for multiple times to ensure the quality of the text.

[0250] Further, in a specific embodiment provided by the present application, the error correction prompt module 13 is configured to:

[0251] display the initial proofreading correction prompt or the comprehensive correction prompt;

[0252] collect operation instructions of the user for the initial proofreading correction prompt or the comprehensive correction prompt;

[0253] The operation instructions include individual adoption of the corrected word object or the corrected sentence object.

[0254] Or the operation instructions include all adoption of the corrected word object or the corrected sentence object.

[0255] Further, in a specific embodiment provided by the present application, the text editing module 11 is configured to collect user feedback on the adoption of the comprehensive correction prompt, and perform multiple rounds of proofreading on the proofreading text data.

[0256] The text editing module 11 is also used to obtain the first-round proofreading text data updated based on the comprehensive correction prompt.

[0257] The content proofreading module 12 is also used to perform sentence object recognition based on the first-round proofreading text data to generate a sentence recognition result, and is also used to perform at least one of semantic expression proofreading, logic structure proofreading, or professional term proofreading on the sentence recognition result to determine a problem sentence object as a problem area, and is also used to perform at least one of semantic expression correction, logic structure correction, or professional term correction based on an association network of a plurality of problem areas and corresponding context sentences of the problem areas to generate a fourth proofreading result.

[0258] The error correction prompt module 13 is also used to generate multiple rounds of comprehensive correction prompts based on the fourth proofreading result.

[0259] On this basis, the text intelligent correction system 100 based on the large language model can also implement any of the above-mentioned embodiments of the text intelligent correction method based on the large language model, which will not be repeated here.

[0260] It should be noted that the terms "comprising", "including", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0261] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer usable program code.

[0262] The above only describes the embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various changes and modifications to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the scope of the claims of the present application.

Claims

1. A text intelligent proofreading method based on a large language model, characterized in that: The following steps are involved: Obtain the text data to be verified; Performing a first dimension calibration based on the text data to be calibrated using a first calibration rule to generate a first calibration result; Performing a second dimension calibration based on the text data to be calibrated using a second calibration rule to generate a second calibration result; Generate preliminary inspection correction prompts based on the first inspection result or the second inspection result; Obtaining initial inspection and correction text data updated based on initial inspection and correction prompts; Based on the initial calibration text data, perform comprehensive dimensional calibration using the third calibration rules to generate the third calibration results; Generate comprehensive correction prompts based on the third calibration results; The step of performing first dimension calibration based on the text data to be calibrated using a first calibration rule to generate a first calibration result includes the following steps: According to the text data to be checked, word object recognition is performed to generate word recognition results; Based on the word recognition results, perform typo correction, determine the wrong word object and the corrected word object corresponding to the wrong word object as the first correction result; The method of performing second dimension calibration according to the text data to be calibrated using a second calibration rule to generate a second calibration result includes the following steps: According to the text data to be checked, word object recognition is performed to generate word recognition results; Based on the word recognition results, perform sensitive word verification to determine the sensitive word object and the corrected word object corresponding to the sensitive word object as the second verification result; The method of performing comprehensive dimensional calibration based on the initial calibration text data using the third calibration rule to generate the third calibration result specifically includes: Based on the initial proofread text data, sentence object recognition is performed to generate sentence recognition results; Based on the sentence recognition results, at least one of semantic expression verification, logical structure verification, or professional terminology verification is performed to determine the problem sentence object as the problem area; Identify the category attributes of the initial proofreading text data; Based on the category attributes of the initial proofread text data, matching the third proofreading rule corresponding to the category attributes; Based on the third correction rule, the language model is controlled to target the association network of several problem areas, match at least one correction sentence object of semantic expression correction, logical structure correction, and professional terminology correction, perform comprehensive dimension correction, and generate the third correction result; The category attributes of the preliminary proofread text data include at least one category attribute of current affairs news, social news, culture and education, medical and health care, environment and nature, sports news, economic news, and entertainment news; The third verification rule is an instruction template, which includes at least system instructions, scene prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and at least one instruction template in the example, and is used to control the output of the language model.

2. The text intelligent proofreading method based on a large language model according to claim 1, characterized in that: The obtaining of the text data to be verified specifically includes: Get the text data in the text editing area or the text data in the area selected by the cursor as the original text data; Perform data cleaning, word segmentation, and part-of-speech tagging on the original text data to generate text data to be proofread.

3. The text intelligent proofreading method based on a large language model according to claim 1, characterized in that: The method further comprises: Based on the association network between several problem areas and the context sentences of the corresponding problem areas, at least one correction among semantic expression correction, logical structure correction and professional terminology correction is performed as the third verification result.

4. The text intelligent proofreading method based on a large language model according to claim 1, characterized in that: The method further comprises: Based on the error word object or the sensitive word object, identify the error type corresponding to the word object; Determine the context word object of the incorrect word object or the sensitive word object; Determine several corrected word objects according to the error type and context word objects; Several correction word objects are evaluated, and the best correction word object is selected as the initial correction prompt through a ranking algorithm.

5. The text intelligent proofreading method based on a large language model according to claim 1, characterized in that: The method further comprises: Based on the problem area, identify the error type corresponding to the problem area; Identify the context statement object for the problem area; Determine several corrected sentence objects according to the error type and context sentence objects; Several correction sentence objects are evaluated, and the best correction sentence object is selected as the comprehensive correction prompt through a ranking algorithm.

6. The text intelligent proofreading method based on a large language model according to claim 1, characterized in that: The method further comprises: Obtain the first round of proofreading text data based on comprehensive correction prompt updates; Based on the first round of proofreading text data, sentence object recognition is performed to generate sentence recognition results; Based on the sentence recognition results, at least one of semantic expression verification, logical structure verification, or professional terminology verification is performed to determine the problem sentence object as the problem area; Based on the association network of the plurality of problem areas and the context sentences of the corresponding problem areas, performing at least one correction of semantic expression, logical structure, and professional terminology as a fourth verification result; Based on the fourth calibration results, multiple rounds of comprehensive correction prompts are generated.

7. A text intelligent proofreading system based on a large language model, characterized by: include: A text editing module is used to obtain text data to be checked; A content verification module, configured to perform a first dimension verification based on the text data to be verified using a first verification rule to generate a first verification result; It is also used to perform a second dimension calibration based on the text data to be calibrated using a second calibration rule to generate a second calibration result; An error correction prompt module is used to generate an initial inspection correction prompt based on the first inspection result or the second inspection result; The text editing module is further used to obtain the initial proofreading text data updated based on the initial proofreading correction prompt according to the user's feedback; The content verification module is further configured to perform comprehensive dimension verification based on the initial verification text data using a third verification rule to generate a third verification result; The error correction prompt module is further used to generate a comprehensive correction prompt based on the third verification result; The content verification module is configured to perform a first dimension verification based on the text data to be verified using a first verification rule to generate a first verification result, and includes the following steps: According to the text data to be checked, word object recognition is performed to generate word recognition results; Based on the word recognition results, perform typo correction, determine the wrong word object and the corrected word object corresponding to the wrong word object as the first correction result; The content verification module is used to perform second dimension verification based on the text data to be verified using a second verification rule to generate a second verification result, including the following steps: According to the text data to be checked, word object recognition is performed to generate word recognition results; Based on the word recognition results, perform sensitive word verification to determine the sensitive word object and the corrected word object corresponding to the sensitive word object as the second verification result; The content verification module is used to perform comprehensive dimension verification based on the initial verification text data using the third verification rule to generate a third verification result, specifically including: Based on the initial proofread text data, sentence object recognition is performed to generate sentence recognition results; Based on the sentence recognition results, at least one of semantic expression verification, logical structure verification, or professional terminology verification is performed to determine the problem sentence object as the problem area; Identify the category attributes of the initial proofreading text data; Based on the category attributes of the initial proofread text data, matching the third proofreading rule corresponding to the category attributes; Based on the third correction rule, the language model is controlled to target the association network of several problem areas, match at least one correction sentence object of semantic expression correction, logical structure correction, and professional terminology correction, perform comprehensive dimension correction, and generate the third correction result; The category attributes of the preliminary proofread text data include at least one category attribute of current affairs news, social news, culture and education, medical and health care, environment and nature, sports news, economic news, and entertainment news; The third verification rule is an instruction template, which includes at least system instructions, scene prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and at least one instruction template in the example, and is used to control the output of the language model.

8. The text intelligent proofreading system based on a large language model according to claim 7, characterized in that: The error correction prompt module is configured to: Display initial inspection correction prompts or comprehensive correction prompts; Collecting the user's operation instructions for the initial inspection correction prompts or comprehensive correction prompts; The operation instructions include adopting the corrected word objects or corrected sentence objects one by one; Alternatively, the operation instruction includes all adoptions of the corrected word object or the corrected sentence object.

9. The text intelligent proofreading system based on a large language model according to claim 7, characterized in that: The text editing module is configured to collect user feedback on the comprehensive correction prompts and perform multiple rounds of proofreading on the proofread text data; The text editing module is further used to obtain the first round of proofreading text data updated based on the comprehensive correction prompt; The content review module is further configured to perform sentence object recognition based on the first round of review text data to generate a sentence recognition result; and is further configured to perform at least one of semantic expression review, logical structure review, or professional terminology review on the sentence recognition result to determine the problem sentence object as the problem area; further configured to perform at least one of semantic expression correction, logical structure correction, and professional terminology correction based on an association network between the plurality of problem areas and contextual sentences corresponding to the problem areas, as a fourth verification result; The error correction prompt module is further used to generate multiple rounds of comprehensive correction prompts based on the fourth calibration result.

Citation Information

Cited By

  • Intelligent document checking method based on large language model

    CN121859891A

  • A document intelligent proofreading method based on a large language model

    CN121859891B