Multi-language text quality evaluation method and intelligent text processing system
Through multilingual text quality assessment methods and intelligent text processing systems, and using pre-trained models for text quality assessment and optimization, the problem that traditional methods are difficult to adapt to multilingual environments is solved, and efficient and accurate text quality assessment and optimization are achieved.
Patent Information
- Application Number
- CN202510727543.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional multilingual text quality assessment methods are difficult to fully and accurately adapt to complex multilingual environments and lack effective evaluation and optimization means.
Adopting multilingual text quality assessment methods and intelligent text processing systems, through language category recognition, rapid error detection and detailed evaluation and analysis, using pre-trained machine learning models to evaluate and optimize text quality, generate diagnostic reports and perform text re-production to achieve self-optimization of the model.
It achieves efficient processing of multilingual texts, rapid error detection and in-depth evaluation, improves the accuracy and adaptability of text quality assessment, and adapts to changing network environments and user needs.
Smart Images

Figure CN120670939A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent text data technology, and more particularly, to a multilingual text quality assessment method and an intelligent text processing system. Background Art
[0002] In today's globalized era, the exchange and dissemination of information transcends language boundaries, and multilingual texts appear increasingly frequently in various fields. Whether it is the documents of multinational companies, the results of international cooperation in academic research, or the rich and colorful multilingual content on the Internet, there is an urgent need to accurately assess the quality of multilingual texts.
[0003] Effective multilingual text quality assessment is not only a means of measuring the quality of texts, but also the key to ensuring accurate information transmission, promoting cross-cultural communication, and improving the efficiency and quality of text processing. However, different languages have their own unique grammatical structures, vocabulary systems, and semantic expressions, which makes multilingual text quality assessment a challenging task. Traditional assessment methods are often limited to a single language or specific field, and it is difficult to fully and accurately adapt to the complexity of the multilingual environment.
[0004] To this end, we provide a multilingual text quality assessment method and an intelligent text processing system to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a multilingual text quality assessment and intelligent text management system. By performing language category identification on user text data, designing two stages of rapid error detection and detailed evaluation and analysis, and utilizing a pre-trained machine learning model to complete rapid evaluation and in-depth analysis and evaluation of multilingual text quality, the text quality score is quantitatively generated. At the same time, the original text is optimized and reproduced based on a detailed diagnostic report, which is provided to the user for selection and statistical feedback, thereby achieving self-optimization of the model, improving the system robustness and the accuracy of the final evaluation, and providing an accurate text quality assessment solution to solve the problems in the above-mentioned background technology.
[0006] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solutions:
[0007] S1. Obtain user text data; perform language identification and preprocessing on the user text data;
[0008] S2. Perform error detection on the preprocessed user text data to obtain an error category list; based on the error category list, use a pre-trained quality scoring model to evaluate and analyze the user text data to obtain a diagnosis report and text quality score;
[0009] S3. Based on the text quality score, the pre-trained optimization suggestion generation model is used to optimize and refine the diagnostic report to obtain optimization suggestions; based on the optimization suggestions, the user text data is rewritten to obtain rewritten text; the optimization suggestions and rewritten text are returned to the user end, and user selection information is obtained through user selection and feedback;
[0010] S4. Adjust the parameters and evaluation rules of the quality scoring model based on the user selection information to achieve self-learning optimization of the quality scoring model.
[0011] Furthermore, the present invention provides an intelligent text processing system, comprising:
[0012] The text input module is used to obtain user text data; identify the language type and pre-process the user text data;
[0013] An error detection module is used to perform error detection on the preprocessed user text data and obtain an error category list;
[0014] The quality assessment module is used to evaluate and analyze user text data to obtain a diagnosis report; quantify the diagnosis report and output a text quality score;
[0015] The optimization suggestion generation module is used to optimize and refine the diagnosis report using the pre-trained optimization suggestion generation model to obtain optimization suggestions;
[0016] The text re-creation and feedback module is used to re-create the user text data to obtain the re-creation text; return the optimization suggestions and the re-creation text to the user end, and obtain the user selection information through user selection and feedback;
[0017] The model adaptive optimization module is used to adjust the parameters and evaluation rules of the quality scoring model to achieve self-learning optimization of the quality scoring model.
[0018] The technical effects and advantages of the multilingual text quality assessment method and intelligent text processing system of the present invention are as follows:
[0019] In the technical solution provided by this application, by performing multilingual recognition and text preprocessing on user text data, user text can be identified as a consistent encoding, thereby achieving efficient processing of multilingual text. A dual-stage error detection and evaluation analysis is designed. Error detection achieves rapid evaluation of user text and location of text errors, and evaluation analysis achieves in-depth evaluation of text, which not only ensures the evaluation and analysis of error types, but also provides data support for text quality scoring. By optimizing and refining the diagnostic report and reproducing the text and returning it to the client, the optimization strategy of the text is analyzed in detail, which improves the adaptability of the text to user selection. The model is adjusted based on the regular collection of user selection data, and adaptive optimization of the model is achieved, so that the system can continuously adapt to the changing network environment and user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of a multilingual text quality assessment method according to the present invention;
[0021] Figure 2 Schematic diagram of an intelligent text processing system of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0023] Example 1
[0024] See also Figure 1 As shown, the multilingual text quality assessment method described in this embodiment includes:
[0025] S1. Obtain user text data; perform language identification and preprocessing on the user text data;
[0026] User text data includes the original natural language text uploaded by the user, text language information, text usage scenario identifier, timestamp, user editing history, user feedback information, etc. The original natural language text may be diverse, including but not limited to articles, comments, Q&A, translation results, abstracts, technical documents, etc. The text language information contains the specific language used in the content of the original text provided by the user, which may contain only one official language or be composed of two or more languages, such as text in two or more official languages and its translation. For these easy-to-identify texts, the text language information records and classifies the specific language used in the user text and its corresponding complete text sentences, facilitating the subsequent identification and preprocessing of the corresponding language. The text usage scenario identifier, timestamp, user editing history, etc. each reflect the detailed record of the user-uploaded text, providing the necessary contextual information for subsequent text quality assessment.
[0027] Specific operations for language identification and preprocessing of user text data:
[0028] Obtain user text data; perform language recognition on the user text data and classify it into single-language text and multi-language text;
[0029] This embodiment uses sophisticated language recognition technology to automatically determine the language of a text. Existing language recognition technology is highly capable of recognizing most of the world's official languages. It utilizes a model based on character n-grams and machine learning to accurately distinguish between multiple languages, including hundreds of languages such as Chinese, English, Japanese, Korean, Arabic, and Spanish. For example, Facebook's FastText provides a pre-trained model that can recognize 176 languages. After receiving user text data, the system quickly determines the primary language label and confidence level of the text. For short text or content that combines multiple languages, the system provides multiple possible language options and records the current text data as multilingual.
[0030] Perform secondary recognition on multilingual texts to identify the complete language set;
[0031] For mixed language scenarios, mixed language texts are detected and recognized, the complete text content is divided into detailed categories, and the same language is classified. The alternative language texts interspersed in them retain their original order and return to the language library in order to perform candidate language recognition again until all uncertain language content is completely identified. This operation completes the language detection and recognition of various mixed language texts composed of multiple language texts.
[0032] Perform encoding format detection and encoding conversion on single-language and multi-language texts to obtain unified encoding texts;
[0033] The text data uploaded by different users may use different official languages. These input texts have different character encoding formats in the input format, which requires text encoding detection and conversion. For different types of text content, the encoding used by the text (such as UTF-8, GBK, ISO-8859-1, etc.) is automatically detected. By analyzing byte features or BOM marks, the encoding type is identified and associated with the corresponding text. After detecting non-standard encoding, it is converted to the universal Unicode (UTF-8) encoding. After processing, all uploaded text data is converted into a unified format at the character encoding level and is represented in standard Unicode (UTF-8) encoding, effectively eliminating the garbled code problem that may be caused by encoding differences.
[0034] Perform text preprocessing on Unicode text;
[0035] Common text preprocessing methods are:
[0036] Denoising: Removes noise and irrelevant content from text, including removing unnecessary punctuation and spaces, deleting emojis and other special symbols that are irrelevant to analysis, and retaining clean natural language content.
[0037] Word segmentation: Segment and mark continuous text content. English, Arabic, and other Latin languages are segmented according to the spaces between words, converting the original text into continuous words. For language texts without spaces, such as Chinese and Japanese, the text is divided according to the semantic relevance between words, thereby converting them into phrases with actual meaning.
[0038] Case unification: Unifies the case of letters in a text. This feature is primarily targeted at languages with both uppercase and lowercase Latin letters, such as English and Spanish. Unifying the case of letters does not affect case-insensitive languages. This feature reduces the feature dimension of a language and improves matching and statistical accuracy.
[0039] These user text data have significant differences in classification, and different analysis methods are also required in subsequent text processing. The recognized language text may be a text-based language or a character-based language. For different types of languages, text-based languages maintain their original meaning in structure, while character-based languages are composed of a variety of common characters and have a considerable number of permutations and combinations. Therefore, when recognizing this type of text, you only need to prepare the dictionary content used by the language in advance, search in the order of arrangement, and insert the semantics according to the semantics applicable to the text to obtain the specific semantics of this type of user language text data. For text-based languages such as Chinese, Japanese, and Korean, which involve different meanings of each character, a new meaning is represented by combining characters with different meanings to form a new noun.
[0040] S2. Perform error detection on the preprocessed user text data to obtain an error category list; based on the error category list, use a pre-trained quality scoring model to evaluate and analyze the user text data to obtain a diagnosis report and text quality score;
[0041] In order to conduct a detailed evaluation of the pre-processed user text data, the present invention is designed to have two stages, namely the error detection stage I1 and the evaluation and analysis stage I2. The error detection stage uses a pre-trained text error recognition model to quickly scan the input text, identify potential error types, annotate the text according to the trained error category label set, and output a list of error categories. In the evaluation and analysis stage, based on the error detection results, the pre-trained quality scoring model is called to evaluate and analyze the text, and the user text data and the identified error categories are input into the quality scoring model together with a predefined instruction template to generate a diagnostic report. The diagnostic report contains the type, location and specific description of each error, and finally the text quality score is obtained through summary and quantification.
[0042] Perform error detection on the preprocessed user text data to obtain an error category list. The specific steps are as follows:
[0043] For the pre-processed user text data, the data is fed into a pre-trained text error recognition model. The text error recognition model performs preliminary and rapid error detection on the input text data, obtaining the error type and sentence location of the error type in the original text in a relatively short period of time. After obtaining the complete location of the corresponding text error, the error types of these data are counted and summarized to obtain a list of error categories, accurate to specific lines and complete text sentences. Error types in the original text include grammatical errors (detecting grammatical problems such as subject-verb inconsistency, tense misuse, and inappropriate collocation in sentences), spelling errors (identifying typos or misspellings of foreign words, and detecting possible clerical errors), punctuation and formatting errors (checking for misuse of punctuation and formatting issues), and style issues (detecting lengthy sentences, excessive passive voice, and mixed spoken and written styles). The pre-trained text error recognition model is fine-tuned based on the existing large language model. After training on the dataset, it generates a function that can quickly detect various types of text errors in the text. The results of the rapid error detection are presented in the form of an error type list, which summarizes the various types of problems in the text and their frequency of occurrence. This list provides a high-level overview and shows where the current text quality is lacking.
[0044] Specifically, given a user text T, it is fed into the text error recognition model to perform the error detection function Perform fast error classification, and the model quickly produces a set of possible error categories:
[0045] , in is the set of m error category labels detected, is the error detection function obtained by training the text error recognition model, and I1 is the prompt instruction to guide the model to locate the error. Contains several error category labels, each of which identifies a problem type with a very concise phrase. The model is fine-tuned from a pre-trained Large Language Model (LLM). After some optimization and improvement, it is more suitable for text detection in the present invention. θ represents its model parameters. The loss function to minimize the classification error rate is trained as follows:
[0046]
[0047] Where T j is the jth training sample text, y j,k is the true label (1 means sample j has the kth type error, otherwise 0), Pθ(c k ∣T j ) is the model prediction text T j The probability of belonging to the wrong class k is obtained by minimizing the cross entropy loss Lerr , the model learns to accurately identify various types of errors.
[0048] Based on the error category list, the pre-trained quality scoring model is used to evaluate and analyze the user text data to obtain a diagnosis report and text quality score;
[0049] The error category list and original user text data are input into a pre-trained quality scoring model, and a diagnostic report and text quality score are output based on the quality scoring model. A diagnostic report is generated by conducting detailed diagnostic analysis of the error types in the user text, thereby realizing an intelligent assessment of the user text quality, so as to facilitate subsequent text reproduction of the user text.
[0050] A pre-trained quality scoring model refers to a machine learning model that has been trained using a large number of annotated user text datasets before implementing an intelligent evaluation of each user text. These datasets contain text quality features of different user texts. Specifically, user texts may contain multiple types of text errors. The model accurately identifies the specific location of the error through the error category list, such as the original line of text or the complete original text sentence. After learning the patterns and rules of the training set samples, it can accurately identify the type of text errors in new user text data and accurately evaluate the text quality. During the training process, this quality scoring model uses supervised learning methods to establish a mapping relationship between the input error category list and the corresponding original user text and the target output (detailed diagnostic report or text quality score). Through repeated iterative optimization, the model can effectively capture the complex relationship between the error category list and text quality, and ultimately achieve a high scoring accuracy.
[0051] In practical applications, such pre-trained quality scoring models are typically built based on advanced natural language processing models, such as BERT, RoBERTa, and DistilBERT, using Transformer architectures. These models possess strong semantic understanding and error analysis capabilities. The training process involves several key steps: First, sufficient high-quality user text data and corresponding error type lists are collected as training samples. These samples must cover a variety of text error types and user texts of varying quality to ensure the model's generalization capabilities. Second, feature analysis is performed on the data, such as calculating the impact of each text error type on the semantics of the original text and its overall quality, and standardizing these features. Third, the data is divided into a training set and a validation set. The model is trained on the training set, while the validation set is used to evaluate model performance. Finally, an optimization algorithm (such as gradient descent) is used to continuously adjust the model parameters to minimize the evaluation error. Through this training process, the model learns the intrinsic connection between the error type list of the original text and the text quality assessment, laying the foundation for subsequent user text reproduction.
[0052] During the application phase, a pre-trained quality scoring model takes each piece of user text data and its corresponding error type list as input and outputs a text quality score after model calculation. The text quality score is typically a quantitative value representing text quality. It is a comprehensive assessment of text quality derived from a combination of multiple text errors and semantic errors, calculated by the model. Its purpose is to objectively score user texts using the correlation errors within the user text data. A high text quality score indicates high-quality text, with few or no text errors, and high-quality copywriting also positively impacts the final score. A low text quality score indicates a high number of text errors, and the detailed diagnostic report generated for that text also includes a more detailed error diagnosis. Using the model-generated text quality score, the system can intelligently evaluate each user text and further determine whether the text requires subsequent text rework. In this way, the pre-trained quality scoring model not only enables fast and efficient processing of large-scale user text data, but also significantly improves the accuracy and reliability of text error diagnosis, providing a scientific basis for subsequent text rework and user feedback.
[0053] The quality scoring model is not specifically limited here. Any machine learning model that can comprehensively analyze various types of text errors to generate a text quality score can be used. In order to implement the technical solution of the present invention, the present invention provides a specific implementation method; the calculation formula for generating the text quality score is:
[0054]
[0055] in Score the quality of the text output by the model, For a detailed diagnostic report, is the function generated by the model after training, and T is the total number of words or sentences in the text. The quality scoring model is essentially a conditional text generation model, for which a joint loss L is defined gen To train report generation and rating prediction:
[0056]
[0057] Among them S * is the manual rating of the training samples, is the model text quality score, both are normalized to the same dimension, is the language model loss for the report text, Indicates the lth word of the report, The first term uses the mean squared error to make the model prediction score close to the human rating, and the second term uses cross entropy to ensure that the generated report content matches the high-quality human-written reviews. score is a trade-off coefficient. By simultaneously optimizing these two components, the model can output reasonable scores while generating coherent and useful text descriptions. The specific data represents the error rate of the user text data in terms of textual, grammatical, and semantic consistency in the corresponding scenario. The lower the error rate, the more applicable the user text data is in that scenario, and the model has strong self-regulation capabilities. The higher the error rate, the greater the impact on the text quality score. Regarding the model's scoring mechanism, the closer the model parameters used to adjust the user text data are to the final required parameters, the more the model's output text quality score will reflect the factual quality of the text.
[0058] Due to different requirements for text errors in different scenario types, the same text may receive different scores in different scenarios. Similar situations generally occur due to the following reasons: the platforms and purposes of user uploads are different. For example, in one scenario, the user needs to upload the text to a professional translation platform or a scientific research platform in a related field. The platform will require the data uploaded by the user to be scientifically rigorous and require more written language texts, while the degree of recognition for text content that tends to be spoken will be lower than in general scenarios. At this time, the system automatically adjusts the proportion of such text errors according to the user's needs to adapt to the review requirements of different platforms.
[0059] S3. Based on the text quality score, the detailed diagnostic report is optimized and refined using a pre-trained optimization suggestion generation model to obtain optimization suggestions. Based on the optimization suggestions, the user text data is rewritten to obtain rewritten text. The optimization suggestions and rewritten text are returned to the user end, and user selection information is obtained through user selection and feedback.
[0060] The text quality score generated for each user's text is input into a pre-trained optimization suggestion generation model. The detailed diagnostic report is then optimized and refined based on the optimization suggestion generation model. Users can then make further optimizations based on the detailed diagnostic report or choose to upload a new document. In principle, the optimization suggestion generation model generates detailed optimization suggestions based on the text quality scores generated by the aforementioned quality scoring model and the detailed diagnostic report. The text is then rewritten based on the optimization suggestions, including correcting errors, polishing logically correct common text, and retaining the original text for high-quality text. After receiving the optimization suggestions, the text is rewritten using the optimization suggestion generation model and the data is returned to the user's client (e.g., mobile phone, computer, etc.). After the user selects the corresponding optimization suggestion, they can choose to fully or partially adopt the rewritten text provided by the system. The system records the portion of the rewritten text that the user chooses to accept as approved text and records any unaccepted opinions as rejected text. The system then continues to import the rewritten text into the optimization suggestion generation model for a new round of improvements, obtaining an updated rewritten text until the user selects or actively cancels the update. While optimizing new text, the model also focuses on the user's text modification history. Based on this history, it determines the user's primary modification intent (i.e., user preference). It then categorizes the user text data based on its intended use. In subsequent operations, it refines and optimizes short sentences or complete unit sentences containing incorrect text content based on the user text type. After complete updates, the model promptly returns the updated text to the user. Ultimately, the system aggregates the approval and rejection texts, as well as the update history, contained in this complete user text to generate user selection information.
[0061] The pre-trained optimization suggestion generation model is trained using a large amount of annotated text quality diagnostic data. This dataset contains detailed diagnostic information for different types of text errors and their corresponding optimization suggestions. The model learns the patterns and regularities within this data and is able to find high-quality alternative text in new user text. During training, the model generally uses supervised learning to establish a mapping relationship between the input user text features and the target output. Through repeated iterative optimization, the model learns the complex relationship between user text and optimized text, and matches the corresponding optimization content with the error list, achieving the effect of correcting and optimizing text errors in user text data.
[0062] Pre-trained optimization suggestion generation models are typically built based on advanced deep learning algorithm frameworks, such as the Transformer-based self-attention model. The training process includes the following key steps: collecting sufficient high-quality detailed diagnostic report samples that cover a variety of text error types and corresponding accurate optimization suggestions to ensure the model has good generalization capabilities; performing feature extraction and standardization on the data to clearly identify the error category and location in each detailed diagnostic report; dividing the data into a training set and a validation set, using the training set for model training and the validation set for real-time evaluation of model performance; and using an optimization algorithm to adjust model parameters to minimize the error in generating optimization suggestions. The trained model can learn the intrinsic relationship between detailed diagnostic features and optimization suggestions, thus providing a reliable foundation for generating optimization suggestions.
[0063] The pre-trained optimization suggestion generation model takes the features of each detailed diagnostic report as input and, after model calculation, outputs specific optimization suggestion text. These suggestions typically include clear text modification plans, such as spelling corrections, grammar adjustments, formatting standardization, and style simplification suggestions. Based on the optimization suggestions output by the model, the system automatically generates optimized re-edited text and returns it to the user for selection and feedback. Utilizing the optimization suggestions generated by the model not only enables the rapid and efficient processing of large amounts of text data, but also significantly improves the accuracy and reliability of text quality optimization.
[0064] The advantage of this pre-trained optimization suggestion generation model is that it can extract potential patterns from complex and diverse text error features and automatically adapt to the differences between different users or text types. For example, texts written by different users may have different styles or error characteristics, or due to changes in the text generation environment (such as different platforms and uses), detailed diagnostic data may contain certain noise or deviations. By introducing diverse diagnostic report sample data during the training phase, the model can have high robustness and generalization capabilities, and can stably output accurate and effective optimization suggestions even when the actual application environment changes. The model can be dynamically optimized and updated as new data continues to accumulate, so that it remains efficient and reliable when dealing with new types of text errors or more complex diagnostic scenarios. This dynamic optimization capability makes the optimization suggestion generation model an indispensable and important tool in the text quality assessment and optimization process, providing strong technical support for high-quality text production.
[0065] S4. Adjust the parameters and evaluation rules of the quality scoring model based on the user selection information to achieve self-learning optimization of the quality scoring model.
[0066] User feedback data is transmitted to a cloud server for analysis. The cloud server regularly aggregates user selection information and optimizes the model. The feedback data is first cleaned and annotated. This preprocessing ensures that all text data can be aligned with the same technical framework for model parameter adjustment. Optimization suggestions implemented in user feedback are selected as positive examples, while cases where users reject modifications are considered negative examples requiring improvement. The quality scoring model is incrementally trained using these positive and negative examples to continuously update the training set. Training methods can include supervised fine-tuning, preference learning optimization, and multi-round interactive learning. Supervised fine-tuning uses the user-approved revision as a high-quality reference and adjusts the model's scores for similar texts, prioritizing expected similar texts with high scores. This reduces the rate of revisions to high-scoring texts and better reflects user expectations for such texts. Preference learning optimization constructs a preference comparison between text pairs before and after user modifications, allowing the model to learn to prefer the modified text. This information is recorded and, when faced with similar questions, the model responds based on user preferences, providing more relevant recommendations. For cases where users remain unsatisfied after multiple interactions, multi-round interactive learning analyzes the reasons and incorporates similar conversational context into training to improve the model's performance in these multiple rounds of improvement. After completing the updated training, the new model parameters are adjusted to the quality scoring model in the cloud server, and the cloud-based quality scoring model is optimized to better adapt to user needs in new cases.
[0067] Taking preference learning optimization as an example, in order to inject preference learning signals, the current quality scoring model is used to automatically generate some disturbed versions of text and corresponding evaluation rules as additional training data for training. orig , let the model evaluation rules randomly introduce some grammatical errors (such as disrupted word order, wrong verb tense), replace some words with inaccurate synonyms, delete some important information to make the content incomplete, and thus generate a reduced quality version T neg , based on which it is assumed that T orig Better than T neg , thus obtaining a pair of preference samples (T orig ,T neg ), define the loss for the preferred sample:
[0068]
[0069] ,in is the score of high-quality text, is the score of low-quality text. The ideal evaluation model should give T orig High rating, give T neg Lower score, and able to locate Tneg Let the existing version of the model be labeled with the evaluation rule T neg The error category is T neg (marked as an error, low score) and T orig (no obvious errors, high scores) form the training data to fine-tune the model. In actual training, multiple groups (T orig ,T neg ) take the average, set a certain score safety interval Δ, and use the maximum interval loss Enhanced discrimination. This optimization allows the model to notice the negative examples it generates, thereby enhancing its ability to identify poor quality text. By continuously increasing the number of negative examples, the quality scoring model becomes more sensitive to subtle differences in quality during the scoring process. This process is essentially equivalent to adversarial training of the model against itself: it continuously generates difficult examples to test and train the quality scoring model. By adjusting the intensity and type of perturbations, it can generate examples of various quality levels, freeing the model from being limited to the limited patterns in the original training set, thereby improving its robustness and generalization capabilities.
[0070] In the technical solution provided by the present application, by performing multilingual recognition and text preprocessing on user text data, user text can be identified as a consistent encoding, thereby achieving efficient processing of multilingual text. A dual-stage design of rapid error detection and detailed evaluation and analysis is implemented. Rapid error detection achieves rapid evaluation of user text and location of text errors, while detailed evaluation and analysis achieves in-depth evaluation of text, which not only ensures the evaluation and analysis of error types, but also provides data support for text quality scoring. By optimizing and refining the detailed diagnostic report and reproducing the text and returning it to the client, the optimization strategy of the text is analyzed in detail, thereby improving the adaptability of the text to user selection. The model is adjusted based on the regular collection of user selection data, thereby achieving adaptive optimization of the model, enabling the system to continuously adapt to changing network environments and user needs.
[0071] Example 2
[0072] See also Figure 2 As shown, for the parts not described in detail in this embodiment, please refer to the description of Example 1. An intelligent text processing system is provided, including:
[0073] The text input module is used to obtain user text data; identify the language type and pre-process the user text data;
[0074] An error detection module is used to perform error detection on the preprocessed user text data and obtain an error category list;
[0075] The quality assessment module is used to evaluate and analyze user text data to obtain a diagnosis report; quantify the detailed diagnosis report and output a text quality score;
[0076] The optimization suggestion generation module is used to optimize and refine the diagnosis report using the pre-trained optimization suggestion generation model to obtain optimization suggestions;
[0077] The text re-creation and feedback module is used to re-create the user text data to obtain the re-creation text; return the optimization suggestions and the re-creation text to the user end, and obtain the user selection information through user selection and feedback;
[0078] The model adaptive optimization module is used to adjust the parameters and evaluation rules of the quality scoring model to achieve self-learning optimization of the quality scoring model.
[0079] The modules are connected via wired and / or wireless means to achieve data transmission between modules.
[0080] Since the electronic device described in this embodiment is an electronic device used to implement a multilingual text quality assessment method in the embodiment of this application, based on the multilingual text quality assessment method described in the embodiment of this application, those skilled in the art will be able to understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application will not be described in detail here. As long as those skilled in the art implement the electronic device used in the multilingual text quality assessment method in the embodiment of this application, it falls within the scope of protection of this application.
[0081] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0082] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for users of ordinary skill in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A multilingual text quality assessment method, characterized in that: include: S1. Obtain user text data; Perform language identification and preprocessing on user text data; S2. Perform error detection on the preprocessed user text data to obtain an error category list; based on the error category list, use a pre-trained quality scoring model to evaluate and analyze the user text data to obtain a diagnosis report and text quality score; S3. Based on the text quality score, the pre-trained optimization suggestion generation model is used to optimize and refine the diagnosis report to obtain optimization suggestions; Based on the optimization suggestions, the user text data is remade to obtain a remade text; Return the optimization suggestions and re-edited text to the user end, and obtain the user selection information through user selection and feedback; S4. Adjust the parameters and evaluation rules of the quality scoring model based on the user selection information to achieve self-learning optimization of the quality scoring model.
2. A multilingual text quality assessment method according to claim 1, characterized in that: The language identification and preprocessing of user text data includes: Obtain user text data; perform language recognition on the user text data and classify it into single-language text and multi-language text; Perform secondary recognition on multilingual texts to identify the complete language set; Perform encoding format detection and encoding conversion on single-language and multi-language texts to obtain unified encoding texts; Perform text preprocessing on Unicode text.
3. A multilingual text quality assessment method according to claim 1, characterized in that: The design consists of two phases: error detection and evaluation. In the error detection phase, a pre-trained text error recognition model is used to scan the input user text, identify the error categories in the text, annotate the original text according to the trained error categories, and output a list of error categories. In the detailed evaluation and analysis phase, based on rapid error detection, a pre-trained quality scoring model is called to conduct a detailed evaluation and analysis of the text. The original text and the identified error categories are input into the model to generate a detailed diagnostic report and a comprehensive text quality score.
4. A multilingual text quality assessment method according to claim 3, characterized in that: The error detection is performed to obtain an error category list, including: For the preprocessed user text data, the data is put into a pre-trained text error recognition model, and the input text data is detected by the text error recognition model to obtain the error type of the original text and the sentence location of the error type; after obtaining the complete location of the corresponding text error, the error type of the text data is counted and summarized to obtain a list of error categories.
5. A multilingual text quality assessment method according to claim 3, characterized in that: The evaluation and analysis of user text data to obtain a diagnostic report and text quality score includes: The error category list and original user text data are input into a pre-trained quality scoring model. A detailed diagnostic report and text quality score are output based on the quality scoring model. A detailed diagnostic report is generated by conducting detailed diagnostic analysis of the error types in the user text.
6. A multilingual text quality assessment method according to claim 1, characterized in that: Based on the text quality score, the pre-trained optimization suggestion generation model is used to optimize and refine the diagnosis report to obtain optimization suggestions, including: The text quality score generated for each user text is input into the pre-trained optimization suggestion generation model, and the detailed diagnostic report is optimized and refined based on the optimization suggestion generation model. Optimization suggestions are derived based on the text quality score generated by the aforementioned quality scoring model and the detailed diagnostic report, and the text is remade based on the optimization suggestions. The data is returned to the user's client, and the customer selects the optimization suggestion and can choose to fully or partially adopt the remade text fed back by the system. The part of the remade text that the customer chooses to accept is recorded as the approved text, and the unaccepted opinions are recorded as the rejected text. The text is continuously imported into the optimization suggestion generation model for a new round of improvement to obtain an updated remade text until the user's choice is obtained or the user actively cancels the update. The approved text, rejected text, and update history contained in this complete user text are summarized to generate user selection information.
7. A multilingual text quality assessment method according to claim 1, characterized in that: The adjusting of the parameters and evaluation rules of the quality scoring model based on the user selection information includes: User selection information is transmitted to the cloud server, and data analysis is performed in the cloud; the cloud server regularly summarizes user selection information and optimizes the model; data cleaning and annotation of user selection information are performed, model parameters are adjusted, and optimization suggestions adopted in user feedback are selected as positive samples. Cases where users refuse to modify text are selected as negative samples that need to be improved. The training set is continuously updated with positive and negative samples, and the quality scoring model is incrementally trained; training methods include supervised fine-tuning, preference learning optimization, and multi-round interactive learning methods.
8. A multilingual text quality assessment system according to claim 1, characterized in that: include: Text input module, used to obtain user text data; Perform language identification and preprocessing on user text data; The error detection module is used to quickly detect errors in the preprocessed user text data and obtain a list of error categories; The quality assessment module is used to conduct detailed evaluation and analysis of user text data and obtain a detailed diagnostic report; Quantify the detailed diagnostic report and output text quality score; The optimization suggestion generation module is used to optimize and refine the detailed diagnostic report using the pre-trained optimization suggestion generation model to obtain optimization suggestions; The text re-creation and feedback module is used to re-create the user text data to obtain the re-creation text; return the optimization suggestions and the re-creation text to the user end, and obtain the user selection information through user selection and feedback; The model adaptive optimization module is used to adjust the parameters and evaluation rules of the quality scoring model to achieve self-learning optimization of the quality scoring model.
9. An intelligent text processing system using the multilingual text quality assessment system according to claim 8.