Establishment and procedures for evaluating scientific texts using text tools and machine learning methods
The device and method for evaluating scientific texts using text tools and machine learning techniques address the limitations of existing evaluation methods by providing comprehensive qualitative and quantitative analysis, achieving high prediction accuracy through advanced machine learning and semantic technologies.
Patent Information
- Application Number
- DE102023005102
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-05
AI Technical Summary
Existing methods for evaluating scientific texts lack comprehensive qualitative and quantitative analysis capabilities, particularly in predicting the quality and relevance of scientific content.
A device and method utilizing text tools and machine learning techniques, including a database with import and export interfaces, connected to a training device for adapting data to models, and a weighting device for validation, enabling qualitative and quantitative analysis through semantic technologies and information retrieval.
The solution provides a robust framework for evaluating scientific texts, achieving high prediction accuracy by extracting and combining text features and metadata, and applying advanced machine learning methods to assess text quality and relevance effectively.
Abstract
Description
The invention relates to a device for evaluating scientific texts with text tools and machine learning methods, wherein the device for evaluating comprises a database with import interfaces and export interfaces and an import interface of the database is connected to a data input device, and a method for evaluating scientific texts with text tools and machine learning methods.The document DE 10 2018 213 021 A1 discloses a computer-implemented method and apparatus for text analysis and relates in particular to a prediction of an membership of a composite from a text in a technical field without the texts being evaluated.The invention set out in claims 1 and 5 is based on the object of providing a device and a method for qualitative and quantitative analysis and assessment of scientific texts.This object is achieved by the features listed in claims 1 and 5.The device for evaluating scientific texts with text tools and machine learning methods, wherein the device for evaluating has a database with import interfaces and export interfaces and an import interface of the database is connected to a data input device, and the method for evaluating scientific texts with text tools and machine learning methods is characterized in particular in that a qualitative and quantitative analysis and evaluation of scientific texts with methods of machine learning, the information retrieval and semantic technolgia is provided therewith.For this purpose, the database is connected via a programming interface and a data distribution device to a training device for adapting the data to at least one model and to a weighting device, wherein the data distribution device is designed such that the captured data are divided into test data, training data and validation data, wherein the training device is designed such that the model or a new model is trained using training data by means of semantic technologies of machine learning methods and methods of information retrieval, and wherein the weighting device is designed such that the model or the new model of the training device is validated using validation data. The data distribution device and the evaluation device are connected to a device for text analysis and assessment of the text quality, which, using the model or models on the test data, which have data and metadata provided with features, determines the quality of the model or models and determines decisive features of the texts, such that main components of the texts are identifiable. Furthermore, the device for text analysis and assessment of the text quality is connected to the database via the programming interface.This is done for this purposeextraction of text features and features from metadata in order to create a data base for an evaluation model, andcreating twice and three feature combinations for text features and features from metadata to develop optimized evaluation models; anda training process with a boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with formal, statistical data and metadata for analysis of formal and statistical features and / ora training process with a boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with text data, in particular word lengths, sentence lengths, frequencies and / or measures for text quality such as in particular type-to-token ratio and / or readability indices) for low-level text checking and / ora training process with high-level methods, including sentitime analysis, reduction of spam recognition methods, similarity analysis with K-nearest neighbors and / or deep learning methods for evaluation at an elevated level, anda reduction of features which provide a prediction probability of less than 0.7 for improving the quality and effectiveness of the prediction and / orsequential testing with the combination of the training methods taking into account defined threshold values for increased accuracy, anda use of an additive result formula for the analysis of formal and statistical features, for the low-level classification of texts and for the high-level classification of texts.For training with a boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with text data for low-level text checking, it is possible in particular to use word lengths, sentence lengths, frequencies and / or measures for text quality, such as in particular type-to-token ratio and / or readability indices.By means of the device and the method for evaluating scientific texts with text tools and machine learning methods, in particular a weighting is given, so that there is an acceptance or non-acceptance of publications and / or a approval or non-approval of research requests by means of text tools and machine learning methods. With the use of text tools and machine learning methods in combination with individual methods and tools, features from text data and meta-data are used to achieve weighted results and from this to achieve the best possible prediction probability. For this purpose, a test data corpus for prediction of unknown data is created corresponding to the training corpus in order to ensure reliable assessment. Text features and features from metadata are extracted to create a comprehensive database for the evaluation. Furthermore, twice and three feature combinations are created for text features and features from metadata in order to develop optimized evaluation models. Furthermore, sequential testing is carried out with the combination of methods, taking into account defined threshold values for the accuracy, in order to achieve meaningful results. This is done by training with a boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with formal, statistical data and metadata for analysis of formal and statistical features and / or training with a boosted tree classifier, a random forest classifier, a decision tree classifier, a decision tree classifier, and, A SVM classifier and / or a logistics classifier with text data for low-level text checking and / or training with high-level methods, including sential analysis, reduction of spam recognition methods, similarity analysis with K nearest neighbors and / or deep learning methods for evaluation at an elevated level. Methods for analyzing formal and statistical features, for low-level classification of texts and for high-level classification of texts are weighted by means of an additive result formula. This additive result formula is used to estimate approval or non-approval of a research request or acceptance or non-acceptance of a publication.Advantageous embodiments of the invention are listed in the following developments and embodiments. These may develop the means for evaluating scientific texts with text tools and machine learning methods, wherein the means for evaluating comprises a database with import interfaces and export interfaces and an import interface of the database is connected to a data input means, and the method for evaluating scientific texts with text tools and machine learning methods individually or in a combination.In one embodiment, the database is connected to the evaluation device via a user interface and the data distribution device, such that specific requirements are taken into account in the evaluation of the model or models.In one embodiment, the user interface is connected to a device for annotation, which is designed such that texts and metadata are annotated in order to achieve improved analysis and evaluation of the texts.The data input device is connected to the database via a device for processing features of text data and metadata relating to the scientific texts (research instructions, publications). The means for processing features of text data and metadata to the scientific texts has a metadata extractor which searches and extracts all relevant features, and a metadata collector which aggregates the features relevant to the specific task. A default setting of the device for evaluating scientific texts can thus be mapped with a feature reduction, wherein the prediction probability is greater than 70%.The additive result formula is used in one embodiment for unknown scientific text data to estimate approval of a research request or acceptance of a publication.In one embodiment, errors within the context of the classification are reduced by the application of adaptations in the weighting and combination of the models.Correlation errors and / or interpretation problems are reduced in one embodiment within the scope of classification with the use of regulation techniques. For this purpose, in particular overfitting can be used as a regulation technique.In one embodiment, classic features are used as text features for evaluating the classifier or classifiers.As classical text features, in one embodiment, word lengths, sentence lengths, frequencies of words, higher measures of text quality, and / or readability indices are used.In one embodiment, text features and normalized readability indicesthe number of characters,the number of words,the number of sets,the number of unique words,the number of words which occur only once in the text,the number of words having more than three syllables,the number of syllables,the average of syllables per word,the number of long words,the number of single-syllable words,the number of words having more than three syllables,the flesch reading ase as readability index,the Wiener text formula as readability index,the Flesch-Kincaid readability index,the SMOG readability index,the Coleman-Liau readability index,the automated readability index (ARI),the gunning fog readability index,the gulpase readability index andthe readability index LIX.An embodiment of the invention will be described in more detail below, wherein an apparatus and a method for evaluating scientific texts with text tools and machine learning methods will be described together in more detail.The means for evaluating comprises a database connected to a data input means and a data output means, wherein the data input means is connected to the database via means for editing features of text data and metadata on the scientific texts. The means for processing features of text data and metadata to the scientific texts has a metadata extractor which searches and extracts all relevant features, and a metadata collector which aggregates the features relevant to the specific task.The database is furthermore connected via a programming interface and a data distribution device to a training device for adapting the data to at least one model and to a weighting device.The data distribution device is designed such that the recorded data are divided into test data, training data and validation data. The training device is designed such that the model or a new model is trained with training data by means of semantic technologies of machine learning methods and methods of information retrieval. The evaluation device is designed such that the model or the new model of the training device is validated with validation data.The data distribution device and the evaluation device are furthermore connected to a device for text analysis and assessment of the text quality, which, using the model or models on the test data, which have data and metadata provided with features, determines the quality of the model or models and determines decisive features of the texts, so that main components of the texts are identifiable. Furthermore, the device for text analysis and assessment of the text quality is connected to the database via the programming interface.The database can be connected to the evaluation device via a user interface and the data distribution device, so that specific requirements are taken into account in the evaluation of the model or models. These specific requirements can be input via the data input device. For this purpose, texts and metadata can be annotated with an associated device for annotation in order to achieve an improved analysis and evaluation of the texts.Thus, for the evaluation of scientific texts using text tools and machine learning methods, extraction of text features and features from metadata takes place in order to create a data base for an evaluation model and creation of 2-fold and 3-fold combinations of features for text features and features from metadata in order to develop optimized evaluation models.The evaluation model or the evaluation models are included in the training devicea boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with formal, statistical data and metadata for the analysis of formal and statistical features and / ora boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with text data for low-level text checking and / ortrained with high-level methods, including sentitime analysis, reduction of spam recognition methods, similarity analysis with K-nearest neighbors and / or deep learning methods for evaluation at an elevated level.The basis for this is the test data, training data and validation data from the data distribution device. To improve the evaluation model or the evaluation models, a new distribution data set of test data, training data and validation data can be provided by means of the data distribution device in conjunction with the device for annotation. By means of a reduction of features which provide a prediction probability of less than 0.7, the quality and effectiveness of the prediction is improved and / or sequential testing with the combination of training methods the accuracy is increased taking into account defined threshold values. An additive result formula is used for the analysis of formal and statistical features, for the low-level classification of texts and for the high-level classification of texts.As text features for evaluating the classifier or classifiers, use can be made in particular of classic features. These may be word lengths, sentence lengths, frequencies of words, higher degrees of text quality, and readability indices.Text features and readability indices can be used in particularthe number of characters,the number of words,the number of sets,the number of unique words,the number of words,which occur only once in the text,the number of words having more than three syllables,the number of syllables,the average of syllables per word,the number of long words,the number of single-syllable words,the number of words having more than three syllables,the flesch reading ase as readability index,the Wiener text formula as readability index,the Flesch-Kincaid readability index,the SMOG readability index,the Coleman-Liau readability index,the automated readability index (ARI),the gunning fog readability index,the gulpase readability index andthe readability index LIX.The readability indices are based on different formulas and calculation methods, so normalization is difficult. Normalization is required, however, to provide for comparison between data sets. However, practical maximum values make it possible to normalize with a small deviation from reality. The value one stands in particular for a high text quality, it also being possible to use texts which are difficult to read. The flesch reading erase can theoretically achieve a value of 121, but in practice rarely reaches a value of 100. Since difficult texts achieve a value below 30, normalization can be done at 1-point / 100. The Wiener text formula established for German texts reaches values up to 20, which corresponds to a text that is very easy to read (point value / 20). With regard to the maximum value, the flesch kincaid index is similar to the flesch reading ase, it can theoretically be infinite, but in practice it rarely reaches values above 100. A calculation can therefore be made with a 1-point value / 100. The SMOG index takes into account multi-syllable words, and in practice academic texts achieve values above 12. The maximum value can also be set higher, for example point value / 25th The Coleman-Liau readability index has no maximum value; in practice, academic texts have a value of twelve. Here too, the maximum value can be set higher, for example point value / 30, in order to ensure the comparison. Thus, the maximum value at the automated readability index (ARI) can also be set higher, for example point value / 40. The gunning fog readability index generally reaches values up to 30, but can also be set at point value / 35. The Gulpease readability index evaluates texts on a scale of 1 to 100, with difficult and academic texts achieving values below 30. The calculation can therefore be made at 1-point value / 100. The readability index LIX reaches values of up to 100, which value applies to difficult or academic texts. A calculation can thus be made with point value / 100.Errors within the scope of the classification can be reduced in particular by using adaptations in the weighting and combination of the models. Correlation errors and / or interpretation problems can be reduced, for example, within the scope of classification by using regulation techniques.The additive result formula for unknown scientific text data is particularly useful for evaluating approval of a research request or acceptance of a publication.References included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Patent Literature citedDE 10 2018 213 021 A1
[0002]
Claims
Device for evaluating scientific texts with text tools and machine learning methods, wherein the device for evaluating has a database and the database is connected to a data input device and a data output device, characterized - in that the database is connected via a programming interface and a data distribution device to a training device for adapting the data to at least one model and a evaluation device, wherein the data distribution device is designed such that the captured data are divided into test data, training data and validation data, wherein the training device is designed such that the model or a new model is trained with training data by means of semantic technologies of machine learning methods and methods of information retrieval, and wherein the evaluation device is designed such that the model or the new model of the training device is validated with validation data, the data distribution device and the evaluation device are connected to a device for text analysis and assessment of the text quality, which, using the model or models on the test data, which have data and metadata provided with features, determines the quality of the model or models and determines decisive features of the texts, so that main components of the texts are identifiable, and the device for text analysis and assessment of the text quality is connected to the database via the programming interface.The device according to claim 1, characterized in that the database is connected to the evaluation device via a user interface and the data distribution device, so that specific requirements are taken into account in the evaluation of the model or models.Device according to at least one of claims 1 and 2, characterised in that the user interface is connected to a device for annotation which is designed such that texts and metadata are annotated in order to achieve an improved analysis and evaluation of the texts.Device according to at least one of Patent Claims 1 to 3, characterized in that the data input device is connected to the database via a device for processing features of text data and of metadata relating to the scientific texts, wherein the device for processing features of text data and of metadata relating to the scientific texts has a metadata extractor which searches and extracts all relevant features, and a metadata collector which aggregates the features relevant for the specific task configuration, such that a default setting of the device for evaluating scientific texts can be mapped with a feature reduction, wherein the prediction probability is greater than 70%.Method for evaluating scientific texts with text tools and machine learning methods, comprising - an extraction of text features and features from metadata in order to create a data base for an evaluation model and - a creation of twice and three-fold combinations of features for text features and features from metadata in order to develop optimized evaluation models, and - a training process with a boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with formals, Statistical data and metadata for analysis of formal and statistical features and / or - a training process with a boosted tree classifier, a random forest classifier, a decision tree classifier, an SVM classifier and / or a logistics classifier with text data for low-level text checking and / or - a training process with high-level methods, including sential analysis, reduction of spam recognition methods, similarity analysis with K-nearest neighbors and / or deep learning methods for evaluation at an elevated level and - a reduction of features, providing a prediction probability of less than 0.7 for improving the quality and effectiveness of the prediction and / or sequential testing with the combination of training methods taking into account defined threshold values for increased accuracy and using an additive result formula for analyzing formal and statistical features, for low-level classification of texts and for high-level classification of texts.The method of claim 5, characterized in that the additive result formula for unknown scientific text data is used to estimate approval of a research request or acceptance of a publication.Method according to at least one of Patent Claims 5 and 6, characterized in that errors within the scope of the classification are reduced by using adaptations in the weighting and combination of the models.Method according to at least one of Patent Claims 5 to 7, characterized in that correlation errors and / or interpretation problems are reduced within the scope of the classification by means of use of regulation techniques.Method according to at least one of Patent Claims 5 to 8, characterized in that classic features are used as text features for evaluating the classifier or classifiers.Method according to at least one of Patent Claims 5 to 9, characterized in that word lengths, sentence lengths, frequencies of words, higher degrees of text quality and / or readability indices are used as classical text features.Method according to at least one of Patent Claims 5 to 10, characterized in that text features and normalized readability indices - the number of characters, - the number of words, - the number of phrases, - the number of unique words, - the number of words which occur only once in the text, - the number of words having more than three syllables, - the number of syllables, - the average of syllables per word, - the number of long words, - the number of uni-syllable words, - the number of words having more than three syllables, - the Flesch reading erase as readability index, - the Wiener text formula as readability index, - the Flesch-Kincaid readability index, the SMOG readability index, the Coleman-Liau readability index, the automated readability index (ARI), the gunning fog readability index, the gulpeease readability index, and the readability index are LIX.
Citation Information
Patent Citations
Technique for document editorial quality assessment
US20060100852A1