A method and system for comparative analysis of ultra-long text sets in unsupervised augmented large models

By acquiring data from external Chinese corpora and performing continuous character combination segmentation and evaluation, and then using a pre-trained model for training and parameter adjustment in conjunction with a benchmark model, the problem of insufficient accuracy and reliability in ultra-long text analysis is solved, and efficient comparative analysis of Chinese text data is achieved.

CN120181068BActive Publication Date: 2026-01-06SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510116435.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2026-01-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively extract key information in the analysis of extremely long text datasets, resulting in insufficient accuracy and reliability of the analysis. This is especially true when processing Chinese text, where difficulties exist. Furthermore, the high cost of manual labor and the lack of sufficient Chinese datasets limit the performance of the models.

Method used

By acquiring the corpus data to be processed from an external Chinese corpus, performing continuous character combination segmentation and evaluating word probability, training a pre-trained model, and combining it with a preset benchmark model for parameter adjustment and a preset retrieval enhancement generation mechanism, the accuracy of segmentation and part-of-speech analysis is improved.

Benefits of technology

It improves the accuracy and reliability of data comparison analysis of ultra-long text datasets, meets the complex needs of Chinese text analysis, reduces manual costs, and makes full use of Chinese corpus resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181068B_ABST
    Figure CN120181068B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unsupervised enhancement large model super-long text set data comparative analysis method and system, it can be applied to text data processing technical field.The application obtains the to-be-processed corpus data from external Chinese corpus, then the to-be-processed corpus data is continuously combined and divided to obtain corpus combination data, then the possibility of each corpus combination data forming vocabulary is evaluated to obtain evaluation results, and the pre-training model is trained according to the evaluation results and corpus combination data, so that the pre-training model can make full use of Chinese corpus resources, and then improve the splitting and part-of-speech analysis accuracy of super-long text, then the to-be-processed part-of-speech screening result is input into the preset benchmark model after parameter adjustment by the preset application scenario, so that the preset retrieval enhancement generation mechanism in the preset benchmark model can be used to improve the accuracy and reliability of the data comparative analysis of super-long text dataset, and then meet the analysis demand of super-long text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text data processing technology, and in particular to a method and system for comparative analysis of unsupervised augmented large model ultra-long text sets. Background Technology

[0002] In the comparison of ultra-long text datasets, the length and large amount of information in these texts make it difficult for existing methods to effectively extract key information, thus failing to meet the requirements of high-precision analysis. Furthermore, current ultra-long text processing methods rely on external dictionary data, necessitating extensive manual annotation of dictionary data beforehand, leading to a sharp increase in labor costs. In addition, the lack of sufficient and representative Chinese datasets hinders model training, preventing the effective learning of various features and rules of the Chinese language. This significantly limits the model's performance in Chinese contexts and makes it difficult to adapt to the complexities of Chinese text analysis. Moreover, the unique linguistic system of Chinese presents numerous challenges for models in understanding and processing Chinese text, resulting in performance that falls short of expectations. Consequently, accurate sentence segmentation, word segmentation, and part-of-speech analysis are impossible, impacting the accuracy and reliability of the overall ultra-long text dataset comparison analysis.

[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0004] The main objective of this application is to propose an unsupervised augmented large model data comparison and analysis method and system for ultra-long text sets, which can meet the analysis needs of ultra-long texts and improve the accuracy and reliability of data comparison and analysis of ultra-long text datasets.

[0005] To achieve the above objectives, one aspect of this application proposes a method for comparative analysis of unsupervised augmented large model data on extremely long text sets, the method comprising the following steps:

[0006] Obtain the corpus data to be processed from an external Chinese corpus;

[0007] The corpus data to be processed is divided into continuous character combinations according to a preset corpus length range to obtain corpus combination data;

[0008] The likelihood of each combination of the corpus data forming a vocabulary is evaluated, and the evaluation results are obtained.

[0009] The pre-trained model is trained based on the evaluation results and the combined corpus data;

[0010] Obtain the extremely long text set to be analyzed;

[0011] The long text set to be analyzed is input into the pre-trained model after training for splitting and part-of-speech processing to obtain the part-of-speech filtering results.

[0012] The part-of-speech filtering results to be processed are input into a preset benchmark model for comparative analysis, and the comparative analysis results are output. The preset benchmark model is processed by a preset retrieval enhancement generation mechanism, and the parameters of the preset benchmark model are adjusted according to a preset application scenario.

[0013] In some embodiments, the step of inputting the long text set to be analyzed into the pre-trained model for splitting and part-of-speech tagging to obtain the part-of-speech tagging results includes:

[0014] The first and second ultra-long text sets in the ultra-long text set to be analyzed are respectively input into the pre-trained model after training for sentence segmentation processing to obtain the first sentence-segmented normalized text set corresponding to the first ultra-long text set and the second sentence-segmented normalized text set corresponding to the second ultra-long text set.

[0015] The first and second standardized text sets after sentence segmentation are respectively input into the pre-trained model after training for word segmentation processing to obtain the first word segmentation text set corresponding to the first standardized text set after sentence segmentation and the second word segmentation text set corresponding to the second standardized text set after sentence segmentation.

[0016] Perform word frequency statistics on the first segmented text set and the second segmented text set respectively to obtain the first dictionary set corresponding to the first segmented text set and the second dictionary set corresponding to the second segmented text set;

[0017] The first dictionary set and the second dictionary set are respectively input into the pre-trained model after training to perform part-of-speech tagging, and the part-of-speech tagging results to be processed are obtained.

[0018] In some embodiments, the step of inputting the first ultra-long text set and the second ultra-long text set from the ultra-long text set to be analyzed into the trained pre-trained model for sentence segmentation processing, to obtain the first sentence-segmented normalized text set corresponding to the first ultra-long text set and the second sentence-segmented normalized text set corresponding to the second ultra-long text set, includes:

[0019] Preprocessing is performed on the first ultra-long text set and the second ultra-long text set in the ultra-long text set to be analyzed, respectively, to obtain the first sentence segment and the first subdividable paragraph corresponding to the first ultra-long text set, and the second sentence segment and the second subdividable paragraph corresponding to the second ultra-long text set.

[0020] The first subdivided paragraph and the second subdivided paragraph are respectively input into the pre-trained model after training for sentence segmentation, to obtain the third sentence corresponding to the first subdivided paragraph and the fourth sentence corresponding to the second subdivided paragraph.

[0021] The first and third sentence segments are merged to obtain the standard text set after the first sentence segmentation;

[0022] The second and fourth sentence segments are merged to obtain the standard text set after the second sentence segment.

[0023] In some embodiments, the step of inputting the first post-segmented standardized text set and the second post-segmented standardized text set into the trained pre-trained model for word segmentation processing includes:

[0024] The first and second standardized text sets after clause segmentation are respectively input into the pre-trained model after training, so that the pre-trained model after training can perform word segmentation on the first and second standardized text sets after clause segmentation by combining the lexical formation rules, semantic associations and word collocation patterns of the Chinese language.

[0025] In some embodiments, performing word frequency statistics on the first segmented text set and the second segmented text set respectively to obtain a first dictionary set corresponding to the first segmented text set and a second dictionary set corresponding to the second segmented text set includes:

[0026] Calculate the first frequency information of each word in the first segmented text set in the first ultra-long text set;

[0027] Statistically analyze the second frequency information of each word in the second segmented text set in the second ultra-long text set;

[0028] The first segmented text set is stored in a preset dictionary based on the first frequency information to obtain the first dictionary set;

[0029] The second word segmentation text set is stored in the preset dictionary according to the second frequency information to obtain the second dictionary set.

[0030] In some embodiments, the step of inputting the first dictionary set and the second dictionary set into the trained pre-trained model for part-of-speech tagging to obtain the part-of-speech tagging result to be processed includes:

[0031] The first dictionary set is input into the pre-trained model after training, so that the pre-trained model performs part-of-speech tagging on each word in the first dictionary set to obtain the first part-of-speech tagging result.

[0032] The second dictionary set is input into the pre-trained model after training, so that the pre-trained model performs part-of-speech tagging on each word in the second dictionary set to obtain the second part-of-speech tagging result.

[0033] Based on the first part-of-speech tagging result, words in the first dictionary set are filtered through a preset filtering mechanism to obtain the first analyzable high-frequency words in the part-of-speech tagging result to be processed;

[0034] Based on the second part-of-speech tagging result, the words in the second dictionary set are filtered through the preset filtering mechanism to obtain the second analyzable high-frequency words in the part-of-speech tagging result to be processed.

[0035] In some embodiments, the evaluation of the probability of each combination of corpus data forming a vocabulary, to obtain an evaluation result, includes:

[0036] Calculate the point mutual information and left and right information entropy of the combined corpus data respectively;

[0037] The sum of the point mutual information and the left and right information entropy is calculated as the probability value of the combined corpus data forming independent word groups;

[0038] The evaluation result of the corpus combination data is generated based on the probability value.

[0039] In some embodiments, the step of inputting the part-of-speech tagging results to be processed into a preset benchmark model for comparative analysis and outputting the comparative analysis results includes:

[0040] The first analyzable high-frequency words and the second analyzable high-frequency words are merged to obtain a merged dictionary set;

[0041] The merged dictionary set is input into a preset benchmark model for comparative analysis, and the comparative analysis results are output.

[0042] To achieve the above objectives, another aspect of this application proposes an unsupervised augmented large model data comparison and analysis system for ultra-long text sets, the system comprising:

[0043] The first module is used to obtain the corpus data to be processed from an external Chinese corpus;

[0044] The second module is used to divide the corpus data to be processed into continuous character combinations according to a preset corpus length range to obtain corpus combination data.

[0045] The third module is used to evaluate the probability of each combination of the corpus data forming a word, and to obtain the evaluation result;

[0046] The fourth module is used to train the pre-trained model based on the evaluation results and the combined corpus data;

[0047] The fifth module is used to acquire the extremely long text set to be analyzed;

[0048] The sixth module is used to input the long text set to be analyzed into the pre-trained model after training for splitting and part-of-speech processing, so as to obtain the part-of-speech filtering results to be processed;

[0049] The seventh module is used to input the part-of-speech filtering results to be processed into a preset benchmark model for comparative analysis and output the comparative analysis results. The preset benchmark model is processed by a preset retrieval enhancement generation mechanism, and the parameters of the preset benchmark model are adjusted according to a preset application scenario.

[0050] To achieve the above objectives, another aspect of this application proposes an unsupervised augmented large model data comparison and analysis system for ultra-long text sets, comprising:

[0051] At least one processor;

[0052] At least one memory for storing at least one program;

[0053] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.

[0054] The embodiments of this application include at least the following beneficial effects: This application provides an unsupervised augmented large model for comparative analysis of ultra-long text datasets. This scheme obtains the corpus data to be processed from an external Chinese corpus, divides the corpus data into corpus combination data according to a preset corpus length range using continuous character combinations, and then evaluates the probability of each corpus combination forming a word. Based on the evaluation results and the corpus combination data, a pre-trained model is trained, allowing the pre-trained model to fully utilize Chinese corpus resources, thereby improving the accuracy of ultra-long text segmentation and part-of-speech analysis. The part-of-speech filtering results are then input into a preset benchmark model with parameters adjusted according to a preset application scenario. This allows the preset retrieval enhancement generation mechanism in the benchmark model to improve the accuracy and reliability of comparative analysis of ultra-long text datasets, thus meeting the analysis needs of ultra-long texts. Attached Figure Description

[0055] Figure 1 This is a flowchart of the unsupervised augmented large model ultra-long text set data comparison and analysis method provided in the embodiments of this application;

[0056] Figure 2 This is a schematic diagram of the architecture of UELBERT provided in the embodiments of this application;

[0057] Figure 3 This is a flowchart of processing ultra-long text data provided in an embodiment of this application;

[0058] Figure 4 This is a sentence segmentation flowchart provided in an embodiment of this application;

[0059] Figure 5 This is a schematic diagram of the structure of the unsupervised augmented large model ultra-long text set data comparison and analysis system provided in the embodiments of this application. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application.

[0061] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0062] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0064] In related technologies, research in the three fields of natural language processing, namely word segmentation, sentence segmentation, and part-of-speech analysis, has made certain progress under the impetus of different methods. However, in complex scenarios, the processing accuracy of these three fields of natural language processing still needs to be further improved. Whether it is traditional methods based on rules or modern means based on statistics, machine learning, and deep learning, most rely on massive data resources or heavy manual rule-making work, which undoubtedly limits their flexibility and efficiency and is difficult to fully meet the increasingly diverse application requirements in the field of natural language processing.

[0065] Compared with Chinese, Western language systems such as English and Latin usually perform better in the above research fields. On the one hand, the corpus and dataset resources related to Western language systems are more abundant and diverse, providing solid data support for model training and algorithm optimization. On the other hand, the processing difficulty of Chinese in these fields is significantly higher than that of alphabetic languages. In terms of word segmentation, a single word in English often constitutes an independent semantic unit, while Chinese words are mostly composed of multiple single characters and are prone to ambiguity. For example, in the sentence "曾想过过过过过过过的生活", the different combinations and semantic interpretations of the character "过" make word segmentation challenging even for people with a certain foundation in Chinese, let alone those without a deep accumulation of Chinese and literary heritage. In terms of sentence segmentation, English, with capitalized first letters and relatively limited and standardized punctuation usage, makes the determination of sentence boundaries relatively simple and direct. In part-of-speech analysis, English can quickly infer the part of speech by relying on rich morphological features such as prefixes and suffixes. In contrast, Chinese lacks such obvious morphological markers, and part-of-speech judgment is more complex and difficult, requiring comprehensive consideration of more semantic, grammatical, and contextual factors.

[0066] Based on this, when comparing existing technologies for ultra-long text datasets, due to the long length and large amount of information in long texts, the existing methods are unable to effectively obtain key information during text analysis, thus failing to meet the high-precision analysis requirements. Moreover, in the existing processing of ultra-long texts, external dictionary data is relied on, so a large amount of manual annotation of dictionary data needs to be carried out in advance, resulting in a sharp increase in labor costs. In addition, in the Chinese field, due to the lack of sufficient and representative Chinese datasets, model training is difficult to be fully carried out, and thus it is unable to effectively learn various features and rules of the Chinese language, greatly limiting the performance of the model in the Chinese context and making it difficult to adapt to the complex requirements of Chinese text analysis. At the same time, due to the unique Chinese language system, the model encounters great difficulties in understanding and processing Chinese texts, resulting in the model effect failing to meet expectations, and thus being unable to accurately perform operations such as sentence segmentation, word segmentation, and part-of-speech analysis, affecting the accuracy and reliability of the data comparison and analysis of the entire ultra-long text dataset.

[0067] In view of this, this application provides an unsupervised augmented large-scale model comparative analysis method and system for ultra-long text datasets. This application obtains the corpus data to be processed from an external Chinese corpus, divides the corpus data into continuous character combinations according to a preset corpus length range to obtain corpus combination data, evaluates the probability of each corpus combination forming a word, and then trains a pre-trained model based on the evaluation results and the corpus combination data. This allows the pre-trained model to fully utilize Chinese corpus resources, thereby improving the accuracy of ultra-long text segmentation and part-of-speech analysis. The part-of-speech filtering results are then input into a preset benchmark model with parameters adjusted according to a preset application scenario. This allows the preset retrieval augmentation generation mechanism in the benchmark model to improve the accuracy and reliability of comparative analysis of ultra-long text datasets, thus meeting the analysis needs of ultra-long texts.

[0068] The unsupervised augmented large model ultra-long text set data comparison and analysis method provided in this application relates to the field of text data processing technology. The unsupervised augmented large model ultra-long text set data comparison and analysis method provided in this application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, etc., but is not limited to these; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the server can also be a node server in a blockchain network; the software can be an application implementing the unsupervised augmented large model ultra-long text set data comparison and analysis method, etc., but is not limited to the above forms.

[0069] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0070] The embodiments of this application will be described in detail below with reference to the accompanying drawings:

[0071] Figure 1 This is an optional flowchart of the unsupervised augmented large model ultra-long text set data comparison and analysis method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S110 to S170:

[0072] Step S110: Obtain the corpus data to be processed from an external Chinese corpus;

[0073] Step S120: Divide the corpus data to be processed into continuous character combinations according to the preset corpus length range to obtain corpus combination data;

[0074] Step S130: Evaluate the probability of each combination of corpus data forming vocabulary, and obtain the evaluation results;

[0075] Step S140: Train the pre-trained model based on the evaluation results and the combined corpus data;

[0076] Step S150: Obtain the ultra-long text set to be analyzed;

[0077] Step S160: Input the long text set to be analyzed into the pre-trained model after training for splitting and part-of-speech processing to obtain the part-of-speech filtering results.

[0078] Step S170: Input the part-of-speech filtering results to be processed into the preset benchmark model for comparative analysis, and output the comparative analysis results. The preset benchmark model is processed by the preset retrieval enhancement generation mechanism, and the parameters of the preset benchmark model are adjusted by the preset application scenario.

[0079] It is understood that the pre-trained model in this embodiment can be a model formed using the BERT architecture as the base model (UELBERT). BERT consists of multiple Transformer layers and incorporates unsupervised data augmentation methods to enhance performance. Specifically, BERT is a pre-trained language representation model proposed by Google. Based on the Transformer architecture, it captures contextual information in text sequences through bidirectional encoder layers. BERT's core innovation lies in its bidirectional training mechanism, which allows the model to consider both the left and right contexts of a word when understanding it. Unlike unidirectional language models, BERT can more accurately capture complex relationships between words, thereby generating richer and more semantic word vector representations. This model typically consists of multiple stacked Transformer encoder layers, each containing a self-attention mechanism and a feedforward neural network to achieve deep feature extraction from the input sequence. Due to its powerful contextual understanding and language modeling capabilities, BERT has achieved significant success in various NLP tasks, including but not limited to text classification, question answering systems, and named entity recognition. BERT has demonstrated exceptional performance, particularly in handling long texts and complex sentence structures. Furthermore, BERT and its variants (such as RoBERTa, ALBERT, etc.) have become standard components of modern NLP systems, widely used in industry and academia. In this embodiment, BERT was chosen as the base model, leveraging its powerful language understanding and context-aware capabilities to enhance the performance of the processing. By combining the advantages of BERT's pre-training with task-specific fine-tuning, not only is the model's accuracy improved, but its efficiency and robustness are also ensured. In particular, this embodiment further optimizes the BERT input processing flow and introduces techniques such as unsupervised data augmentation to address the challenges of real-world application scenarios, ensuring optimal user experience and technical effectiveness.

[0080] It is understandable that, such as Figure 2 As shown, this embodiment first obtains the corpus data to be processed from a large-scale external corpus, and then divides the corpus data into continuous character combinations based on a preset corpus length range (4-gram) to obtain corpus combination data. Specifically, for each character in the input sentence, this embodiment extracts all possible continuous character combinations with a length not exceeding 4. In this way, a large number of character combinations can be generated from a single sentence.

[0081] Then, the probability of these corpus combinations forming words is evaluated. In this embodiment, the likelihood of these combinations forming words can be evaluated by calculating the correlation between them. Specifically, this embodiment calculates the point mutual information (PMI) and left-right information entropy of the corpus combinations separately, and then calculates the sum of the PMI and left-right information entropy as the probability value of the corpus combinations forming independent word phrases. The evaluation result of the corpus combinations is then generated based on the probability value. It is understood that point mutual information (PMI) is used to measure the degree of correlation between characters within the corpus combinations, and left-right information entropy is used to evaluate the relevance of the corpus combinations to their context. For example, the word "apple" has a high correlation with the word "juice" in "apple juice," while it is difficult to establish such a connection between non-word units. In this embodiment, the point mutual information is shown in Formula 1, and the left-right information entropy is shown in Formula 2.

[0082] Formula 1;

[0083] Formula 2;

[0084] In the formula, g represents the input combination, p() expresses the probability of its occurrence in the corpus, and m is the specific number of combinations. LE and RE are calculated based on a given input combination g and its left and right adjacent character sets. Taking left entropy as an example, for each left adjacent character in the left adjacent character set, g and its combination are calculated. Conditional probability under given conditions .

[0085] Finally, in this embodiment, the three values ​​are added together to obtain the probability value of each corpus combination data as an independent word group. The larger the probability value, the higher the probability of independence.

[0086] Understandably, after the pre-trained model is trained, this embodiment can input the acquired long text set to be analyzed into the trained pre-trained model for segmentation and part-of-speech tagging. For example... Figure 3 As shown, the splitting and part-of-speech tagging process includes, but is not limited to, the following steps:

[0087] The first and second ultra-long text sets in the ultra-long text set to be analyzed are respectively input into the pre-trained model after training for sentence segmentation, to obtain the first sentence-after-normalized text set corresponding to the first ultra-long text set and the second sentence-after-normalized text set corresponding to the second ultra-long text set.

[0088] The first and second standard text sets after the first clause are input into the pre-trained model after training for word segmentation, respectively, to obtain the first word segmentation text set corresponding to the first standard text set after the first clause and the second word segmentation text set corresponding to the second standard text set after the second clause.

[0089] Word frequency statistics are performed on the first segmented text set and the second segmented text set respectively to obtain the first dictionary set corresponding to the first segmented text set and the second dictionary set corresponding to the second segmented text set;

[0090] The first and second dictionary sets are input into the pre-trained model for part-of-speech tagging, and the part-of-speech tagging results are obtained.

[0091] Specifically, in this embodiment, when inputting the extremely long text set to be analyzed into the pre-trained model, the extremely long text set is first preprocessed to ensure data cleanliness and consistency. It is understood that, as... Figure 4 As shown, this embodiment cleanses the data by removing unnecessary leading and trailing spaces and meaningless special characters caused by formatting errors or garbled characters. Then, regular expressions are used to perform rule-based preliminary sentence segmentation based on standard sentence-ending symbols in Chinese and English grammar (such as periods, question marks, exclamation marks, ellipses, etc.), resulting in the first sentence segment and the first subdividable paragraph corresponding to the first ultra-long text set, and the second sentence segment and the second subdividable paragraph corresponding to the second ultra-long text set. This step reduces the amount of data processed by the model, simplifies the processing flow, and thus significantly shortens the overall processing time.

[0092] In this embodiment, for sentences and paragraphs whose meaning is ambiguous due to the use of commas, semicolons, or other punctuation marks, the information is further input into the model for precise differentiation to ensure the accuracy and completeness of the analysis. It can be understood that in this embodiment, the first and second re-segmentable paragraphs can be input into the pre-trained model for sentence segmentation, resulting in the third sentence segment corresponding to the first re-segmentable paragraph and the fourth sentence segment corresponding to the second re-segmentable paragraph. Then, the first and second sentences are merged to obtain the first segmented text set, and the second and fourth sentences are merged to obtain the second segmented text set. Specifically, in this embodiment, the process of segmenting re-segmentable paragraphs is as shown in Formula 3:

[0093] Formula 3;

[0094] In the formula, T represents the original input text set (a very long text set); P(T) represents the preprocessing operation on T to generate a cleaned text set (removing extra spaces and special characters, and escaping to UTF-8 format); then, the text set is input into the UELBERT model designed in this embodiment of the application. The model will use the built-in tokenization function Tok to convert the optimized text set into a series of token sequences that the model can understand. Subsequently, the token sequences will be segmented into sentences, and finally T' is the segmented text set (the segmented canonical document set).

[0095] Understandably, after completing the sentence segmentation process, this embodiment inputs the first and second standardized text sets after sentence segmentation into the pre-trained model, respectively. This allows the pre-trained model to perform word segmentation on the first and second standardized text sets after sentence segmentation, based on the lexical formation rules, semantic relationships, and word collocation patterns of the Chinese language. Specifically, the pre-trained model in this embodiment can utilize its word segmentation module to perform detailed word segmentation on the input text based on the lexical formation rules, semantic relationships, and common word collocation patterns of the Chinese language.

[0096] Specifically, in this embodiment, after word segmentation, the first frequency information of each word in the first segmented text set appearing in the first ultra-long text set is calculated, and the second frequency information of each word in the second segmented text set appearing in the second ultra-long text set is calculated. Based on the first frequency information, the first segmented text set is stored in a preset dictionary to obtain a first dictionary set; based on the second frequency information, the second segmented text set is stored in the preset dictionary to obtain a second dictionary set. It can be understood that this embodiment, by transforming the originally lengthy and complex ultra-long text set into a dictionary data format characterized by word frequency, greatly simplifies the data scale, reduces data redundancy, and makes the data easier to manage and analyze, while also preserving the key semantic information of the text. The process is shown in Formula 4:

[0097] Formula 4;

[0098] In the formula, w i For the i-th word in the word sequence output after model processing, Fc is a function that counts the occurrences of the word. By applying the Fc function to each word w... i The frequencies of words are statistically analyzed to obtain a word frequency statistics dictionary W, where key-value pairs represent each word element and its corresponding number of occurrences.

[0099] Understandably, in this embodiment, after completing word frequency statistics, the first dictionary set is input into the pre-trained model so that the pre-trained model performs part-of-speech tagging on each word in the first dictionary set to obtain a first part-of-speech tagging result; the second dictionary set is input into the pre-trained model so that the pre-trained model performs part-of-speech tagging on each word in the second dictionary set to obtain a second part-of-speech tagging result; based on the first part-of-speech tagging result, the words in the first dictionary set are filtered through a preset filtering mechanism to obtain a first analyzable high-frequency word in the part-of-speech tagging result to be processed; based on the second part-of-speech tagging result, the words in the second dictionary set are filtered through a preset filtering mechanism to obtain a second analyzable high-frequency word in the part-of-speech tagging result to be processed.

[0100] Specifically, the pre-trained model (UELBERT model) in this embodiment uses its built-in part-of-speech tagging algorithm, combined with language knowledge trained on a large-scale corpus, to accurately tag each word in the dictionary, such as nouns, verbs, and adjectives. After completing the part-of-speech tagging, based on in-depth text analysis results, it was found that nouns and some verbs play a key decisive role in comprehensively comparing the sentiment and other aspects of text content. Therefore, this embodiment designed an intelligent filtering mechanism to retain only words with the part of speech of nouns and some key verbs, while removing other words with less impact on the final analysis. After filtering, the dictionaries corresponding to the two text sets are merged, and the merged dictionary set is in the form of {key1, value1, key2, value2}, where key represents the retained words and value represents their corresponding word frequency information. Through this process, highly generalized dictionary data is obtained, which concisely and accurately reflects the core semantics and key information of the text, providing a high-quality data foundation for the final analysis.

[0101] In some embodiments, to verify the performance of the UELBERT model corresponding to the present application embodiments, this embodiment conducted multiple tests on Chinese sequence labeling tasks, including but not limited to word segmentation, part-of-speech tagging, and named entity recognition. These tasks were selected based on their fundamental and widespread applicability in natural language processing, effectively evaluating the model's language understanding and processing capabilities. It is worth noting that no additional independent tests were conducted for the sentence segmentation task in this embodiment, because the main purpose of sentence segmentation is to quickly and accurately break down extremely long texts into more easily processed segments, thereby improving the efficiency and accuracy of subsequent tasks.

[0102] In the tasks selected above, this embodiment chose several representative public datasets for testing. All datasets used the dataset versions from Modelscope, and their specific divisions are shown in Table 1 as model test datasets. Since some datasets do not provide a Dev set in their official documentation, 10% of the data in the training set will be randomly selected as the Dev set for actual training.

[0103] Table 1

[0104]

[0105] Based on the data in Table 1, the performance of the model in this application embodiment and different models in Chinese word segmentation, part-of-speech tagging, and named entity recognition tasks is compared, demonstrating the advantages of the UELBERT model proposed in this application embodiment over existing excellent models. In the Chinese word segmentation task, UELBERT demonstrates its powerful segmentation ability (Table 2), achieving accuracies of 97.55%, 98.64%, and 96.85% on the CTB6, MSRA, and PKU datasets, respectively. Particularly on the MSRA dataset, UELBERT outperforms the existing best model by 0.24%, showing strong adaptability to large-scale news texts. UELBERT can more accurately capture word boundaries, thus performing excellently in complex and varied Chinese contexts.

[0106] Table 2

[0107]

[0108] In the part-of-speech tagging task (Table 3), UELBERT also performed exceptionally well. It achieved an accuracy of 95.21% on the CTB6 dataset, and 95.75% and 95.71% on the UD1 and UD2 datasets, respectively. Compared to other models, UELBERT not only maintains competitiveness on general standards (such as CTB6), but also performs excellently in cross-linguistic tagging systems (such as UD1 and UD2). This is thanks to the unsupervised external data and adaptive fine-tuning introduced in the pre-training phase of this application, which enables the model to better understand the grammatical roles and contextual relationships of words, further improving the accuracy of part-of-speech tagging.

[0109] Table 3

[0110]

[0111] For the named entity recognition task (Table 4), UELBERT achieved F1 scores of 82.35% and 80.90% on the OntoNotes 4.0 and News datasets, respectively. This performance not only comprehensively outperforms BERT and ERNIE but also surpasses NFLAT and FGN models specifically optimized for named entity recognition. UELBERT's unique feature lies in its enhanced self-attention mechanism, which effectively captures complex relationships between entities. Furthermore, through fine-tuning with external data and a small amount of self-made data, the model parameters were further optimized, allowing UELBERT to maintain high efficiency and accuracy when handling long texts and complex sentence structures.

[0112] Table 4

[0113]

[0114] In summary, UELBERT has achieved excellent results on multiple benchmark datasets by introducing unsupervised augmentation with external data and the self-attention mechanism of Transformer, demonstrating its powerful performance and wide applicability in Chinese natural language processing tasks.

[0115] In other embodiments, this embodiment constructs two sub-databases as a test dataset for the overall method, based on 204 books recommended for reading by teenagers by the Ministry of Education and 460 books most frequently read by students of each grade in a certain region. These datasets cover the full text content of these books. To verify the effectiveness of this method on texts of different lengths and genres, this embodiment selects classic works from this dataset, including *Grimm's Fairy Tales (Selected)*, *Andersen's Fairy Tales (Selected)*, *Romance of the Three Kingdoms*, and *Water Margin*, and conducts separate comparative tests. These works not only represent different literary styles and complexities but also cover a variety of text lengths, from short stories to novels.

[0116] This embodiment compares three sets of data: a selection of Andersen's Fairy Tales and a selection of Grimm's Fairy Tales (100,000-word level); a high-frequency word analysis comparison of Romance of the Three Kingdoms and Water Margin (million-word level); and a comparison of books recommended by the Ministry of Education for teenagers and books actually read by teenagers in a certain region (tens of millions of words level). These data encompass texts of various lengths, demonstrating the generalization and robustness of the model in this embodiment. The first set shows a comparative analysis of texts of the same genre; the second set compares classical Chinese texts; and the third set is a comprehensive comparative analysis of books recommended by the Ministry of Education (204 books) and books actually read by teenagers (460 books). The effectiveness evaluation for text sets of different sizes verifies the efficiency and reliability of the large-model-based comparative analysis method for ultra-long text sets in this embodiment. This embodiment not only has significant advantages in language generation quality but also effectively handles text processing tasks of different scales and complexities, providing a solid foundation for future research and applications.

[0117] Reference Figure 5 This application also provides an unsupervised augmented large model data comparison and analysis system for ultra-long text sets, the system comprising:

[0118] The first module 910 is used to obtain the corpus data to be processed from an external Chinese corpus;

[0119] The second module 920 is used to divide the corpus data to be processed into continuous character combinations according to a preset corpus length range to obtain corpus combination data;

[0120] The third module 930 is used to evaluate the probability of each combination of corpus data forming a word and obtain the evaluation result;

[0121] The fourth module 940 is used to train the pre-trained model based on the evaluation results and the combined corpus data;

[0122] Module 5, 950, is used to acquire extremely long text sets to be analyzed;

[0123] The sixth module 960 is used to input the long text set to be analyzed into the pre-trained model after training for splitting and part-of-speech processing to obtain the part-of-speech filtering results to be processed;

[0124] The seventh module 970 is used to input the part-of-speech filtering results to be processed into a preset benchmark model for comparative analysis and output the comparative analysis results. The preset benchmark model is processed by a preset retrieval enhancement generation mechanism, and the parameters of the preset benchmark model are adjusted according to a preset application scenario.

[0125] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0126] This application also provides an unsupervised augmented large model data comparison and analysis system for ultra-long text sets. The system includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0127] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0128] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0129] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0130] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0131] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0132] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0134] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0135] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0136] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0137] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An unsupervised enhanced large model super-long text set data comparative analysis method, characterized in that, The method comprises: obtaining to-be-processed corpus data from an external Chinese corpus; dividing the to-be-processed corpus data according to a preset corpus length range to obtain corpus combination data; respectively calculating point mutual information and left and right information entropy of the corpus combination data; calculating a sum of the point mutual information and the left and right information entropy as a probability value of the corpus combination data forming an independent word group; and generating an evaluation result of the corpus combination data according to the probability value; training a pre-trained model according to the evaluation result and the corpus combination data; obtaining a to-be-analyzed super-long text set; inputting the to-be-analyzed super-long text set into the trained pre-trained model for splitting and part-of-speech processing to obtain a to-be-processed part-of-speech screening result, comprising: respectively preprocessing a first super-long text set and a second super-long text set in the to-be-analyzed super-long text set to obtain a first sentence break corresponding to the first super-long text set and a first re-segmentable paragraph, and a second sentence break corresponding to the second super-long text set and a second re-segmentable paragraph; inputting the first re-segmentable paragraph and the second re-segmentable paragraph into the trained pre-trained model for sentence processing to obtain a third sentence break corresponding to the first re-segmentable paragraph and a fourth sentence break corresponding to the second re-segmentable paragraph; merging the first sentence break and the third sentence break to obtain a first post-sentence regular text set corresponding to the first super-long text set; merging the second sentence break and the fourth sentence break to obtain a second post-sentence regular text set corresponding to the second super-long text set; inputting the first post-sentence regular text set and the second post-sentence regular text set into the trained pre-trained model for word processing to obtain a first worded text set corresponding to the first post-sentence regular text set and a second worded text set corresponding to the second post-sentence regular text set; respectively performing word frequency statistics on the first worded text set and the second worded text set to obtain a first dictionary set corresponding to the first worded text set and a second dictionary set corresponding to the second worded text set; inputting the first dictionary set and the second dictionary set into the trained pre-trained model for part-of-speech screening to obtain the to-be-processed part-of-speech screening result; inputting the to-be-processed part-of-speech screening result into a preset reference model for comparative analysis to output a comparative analysis result, and parameters of the preset reference model are adjusted through a preset application scenario.

2. The method of claim 1, wherein, The inputting the first post-sentence regular text set and the second post-sentence regular text set into the trained pre-trained model for word processing comprises: inputting the first post-sentence regular text set and the second post-sentence regular text set into the trained pre-trained model to enable the trained pre-trained model to combine Chinese language vocabulary composition rules, semantic associations and word collocation patterns to respectively perform word processing on the first post-sentence regular text set and the second post-sentence regular text set.

3. The method of claim 1, wherein, The word frequency statistics are performed on the first segmented text set and the second segmented text set respectively to obtain a first dictionary set corresponding to the first segmented text set and a second dictionary set corresponding to the second segmented text set, including: statistical information of the first frequency of each word in the first segmented text set appearing in the first super-long text set; statistical information of the second frequency of each word in the second segmented text set appearing in the second super-long text set; storing the first segmented text set into a preset dictionary according to the first frequency information, to obtain the first dictionary set; storing the second segmented text set into the preset dictionary according to the second frequency information, to obtain the second dictionary set.

4. The method of claim 1, wherein, The first dictionary set and the second dictionary set are respectively input into a trained pre-training model for part-of-speech filtering to obtain the processing word part-of-speech filtering result, including: inputting the first dictionary set into the trained pre-training model, so that the trained pre-training model performs part-of-speech tagging on each word in the first dictionary set to obtain a first part-of-speech tagging result; inputting the second dictionary set into the trained pre-training model, so that the trained pre-training model performs part-of-speech tagging on each word in the second dictionary set to obtain a second part-of-speech tagging result; according to the first part-of-speech tagging result, filtering the words in the first dictionary set through a preset filtering mechanism to obtain first analyzable high-frequency words in the processing word part-of-speech filtering result; according to the second part-of-speech tagging result, filtering the words in the second dictionary set through the preset filtering mechanism to obtain second analyzable high-frequency words in the processing word part-of-speech filtering result.

5. The method of claim 4, wherein, The processing word part-of-speech filtering result is input into a preset benchmark model for comparative analysis, and a comparative analysis result is output, including: merging the first analyzable high-frequency words and the second analyzable high-frequency words to obtain a merged dictionary set; inputting the merged dictionary set into the preset benchmark model for comparative analysis to output the comparative analysis result.

6. An unsupervised enhanced large model super-long text set data comparative analysis system, characterized in that, The system is used to perform the method of any one of claims 1-5, and the system includes: a first module configured to obtain processing corpus data from an external Chinese corpus; a second module configured to divide the processing corpus data into corpus combination data according to a preset corpus length range; a third module configured to evaluate the possibility of forming a vocabulary for each corpus combination data to obtain an evaluation result; a fourth module configured to train a pre-training model according to the evaluation result and the corpus combination data; a fifth module configured to obtain a super-long text set to be analyzed; a sixth module configured to input the super-long text set to be analyzed into the trained pre-training model for splitting and part-of-speech processing to obtain a processing word part-of-speech filtering result; a seventh module configured to input the processing word part-of-speech filtering result into a preset benchmark model for comparative analysis to output a comparative analysis result, wherein parameters of the preset benchmark model are adjusted through a preset application scenario.

7. An unsupervised enhanced large model super-long text set data comparative analysis system, characterized in that, including: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor is caused to implement the method recited in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for acquiring threat intelligence data model, medium and electronic equipment

    CN117114009A

  • Interactive data analysis method based on large model

    CN118886415A