Long text abstract automatic generation method based on adaptive blocking and semantic hierarchical clustering

Through the adaptive blocking and semantic hierarchical clustering methods, the problems of insufficient video memory and low computing efficiency in long text digest generation are solved, and more accurate and organized digest generation is achieved.

CN120086367APending Publication Date: 2025-06-03SUZHOU AEROSPACE INFORMATION RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411984652.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing summary generation methods are insufficient video memory, low computational efficiency, accuracy and organization when processing long text.

Method used

Using a method based on adaptive chunking and semantic hierarchical clustering, long text is collected and preprocessed through network crawlers, professional term recognition and standardization is performed, adaptive chunking, summary generation and semantic embedding are performed in parallel, and finally semantic hierarchical clustering is performed to generate a long text summary.

Benefits of technology

Improves computing efficiency, reduces video memory pressure, and generates summary more precise and organized, suitable for long text summary generation application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086367A_ABST
    Figure CN120086367A_ABST
Patent Text Reader

Abstract

The invention discloses a long text abstract automatic generation method based on adaptive partitioning and semantic hierarchical clustering, which comprises the following steps: collecting related long texts based on a web crawler technology, and carrying out terminology recognition, standardization, cleaning and normalization processing; performing self-adaptive blocking, abstract generation and semantic embedding of the long text in parallel by a pipeline, and converting the blocked text into a semantic vector; and performing semantic hierarchical clustering to obtain a highest-layer cluster abstract, namely a long text abstract. According to the method, the length and semantic coherence of the text blocks are considered, calculation and video memory pressure generated by single processing of the long text is avoided, semantic correlation in the same text block is ensured, and semantic independence of adjacent text blocks is ensured; different GPU devices can be fully utilized, the idle waiting time is shortened, the video memory pressure of a single GPU is reduced, and the overall calculation efficiency is improved; different themes and hierarchical structures in the long text can be captured, so that the finally generated abstract is more accurate and organized, and the abstract generation process is more interpretable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to text information processing technology, and specifically to an automatic long text summary generation method based on adaptive chunking and semantic hierarchical clustering. Background Art

[0002] In this era of information explosion, the amount of various text data shows an exponential growth. As a carrier for refining information and conveying key points, summaries play a crucial role in many fields such as government affairs, business, and technology. At the same time, a large number of long text information are continuously generated in various fields. These long texts may contain key information in many aspects, which poses a great challenge to users who need to quickly understand and extract the essence. Therefore, there is an urgent need for a method to accurately and efficiently generate long text summaries.

[0003] Traditional summary production mainly relies on manual reading and analysis, which not only consumes manpower and has low efficiency, but also requires the producer to have a certain level of domain knowledge, and it is difficult to guarantee the summary quality; TextRank [1,2] and other machine learning algorithms mainly identify keywords and key sentences by constructing a graph structure. Although it liberates manpower, it still has deficiencies in deeply understanding semantics.

[0004] With the development of computer hardware, the progress of artificial intelligence technology, and the accumulation of massive training data, in recent years, generative large language models (LLMs) have developed vigorously. Through pre-training on large-scale datasets, LLMs have powerful language understanding and generation capabilities [3] , and the supported context length has also increased to hundreds of K tokens, which provides a feasible basis for its application in the field of long text summary generation. However, too long a context not only brings greater pressure on hardware resources such as video memory and communication bandwidth, but also is not conducive to LLMs grasping key information, which may lead to missing key points or even generating hallucinations. Summary of the Invention

[0005] The purpose of the present invention is to propose an automatic long text summary generation method based on adaptive chunking and semantic hierarchical clustering, so as to solve the problems of insufficient video memory, low computational efficiency, lack of accuracy and coherence in existing summary generation methods.

[0006] The technical solution to achieve the purpose of the present invention is: an automatic long text summary generation method based on adaptive chunking and semantic hierarchical clustering, and the specific steps are as follows:

[0007] Collect relevant long texts based on web crawler technology, and perform professional term recognition, standardization, cleaning and normalization processing;

[0008] Execute long text adaptive chunking, summary generation and semantic embedding in pipeline parallel, and convert the chunked text into semantic vectors;

[0009] Perform semantic-level clustering to obtain the highest-level cluster summary, which is the long text summary.

[0010] Furthermore, collect relevant long texts based on web crawler technology, and perform professional term recognition, standardization, cleaning, and normalization. The specific methods are as follows:

[0011] 1) Based on web crawler technology, collect relevant texts in the field of computer science from multiple databases and websites, including Web of Science, IEEE Xplore, Wikipedia, and Github, etc.;

[0012] 2) Uniformly convert the collected long texts in different formats into plain text format through Python toolkits;

[0013] 3) Identify and standardize professional terms in the long text to reduce data redundancy and ambiguity;

[0014] 4) Clean and normalize the standardized long text, use regular expressions to match and delete abnormal characters in the text, including special symbols such as line breaks, tab characters, and emojis, and identify and delete copyright statements, reference annotations, chart numbers, and footnotes in the text.

[0015] Furthermore, pipeline parallel execution of long text adaptive chunking, summary generation, and semantic embedding, converting the chunked text into semantic vectors, where: The specific method of long text adaptive chunking is as follows:

[0016] 1) WordPiece tokenization

[0017] After preprocessing, the long text is represented as document = {s i |i = 1, 2, …, N s}, where the long text document is a set composed of sentences s, the subscript i is the sentence number, and N s is the number of sentences; after WordPiece tokenization, the sentence is composed of tokens token, represented as The subscript t is the serial number of the token in the current sentence, is the number of tokens of s i ;

[0018] 2) Adaptive chunking

[0019] Successively take two adjacent sentences s i and s i+1Concatenate into the input sequence of the chunking model. Insert the classification token ([CLS]) at the beginning of the sequence to determine whether adjacent sentences are semantically coherent, and insert the segmentation token ([SEP]) between sentences to distinguish different sentences. Then the input sequence is represented as:

[0020]

[0021] where i = 1, 2, …, N s -1;

[0022] Use the semantic coherence prediction model based on the BERT structure to calculate s i and s i+1 The probability of semantic coherence:

[0023] P i = NSP(input i ) (1)

[0024] where NSP is the semantic coherence prediction model based on the BERT structure, P i ∈[0, 1], the closer P i is to 1, the more relevant s i and s i+1 are, and the more inclined to assign the two to the same text chunk. The closer P i is to 0, the more inclined to assign the two to different text chunks;

[0025] To avoid unlimited growth of text chunks while considering semantic coherence, set an upper limit on the text chunk length and adaptively adjust the chunking threshold in combination with the current text chunk length:

[0026]

[0027] where maxlen is the upper limit of the text chunk length set manually, chunk represents the text chunk, the subscript j is the serial number of the chunk, tokennum(chunk j ) is the number of tokens in chunk j , chunk j is initially empty. First, add s i to chunk j , calculate the probability threshold R j , R j ∈[0.5, 1], then compare whether the coherence probability P i of s i+1 and s i is higher than the threshold. If P i > R j , then s i+1 is added to chunk j, and then update R according to the new token num (chunk j ) j . Compare s and s i+1 in the same way i+2 for the coherence probability P i+1 . If P i+1 >R j , then add s i+2 to chunk j . Otherwise, add s i+2 to chunk j+1 . Continue to repeat the above operations until i = N s -1. Finally, obtain the set of text chunks Chunks = {chunk j |j = 1, 2, …, N}, where N is the number of chunks obtained from the long text.

[0028] . Further, the long text adaptive chunking, summary generation, and semantic embedding are executed in parallel in a pipeline to convert the chunked text into semantic vectors, where: For summary generation, the specific method is:

[0029] Based on a large language model, generate summaries for each chunk in the set of text chunks Chunks, expressed as:

[0030] chunk_summary j = LLM(chunk j ) (3)

[0031] where LLM is the large language model for generating summaries, and chunk_sumary j represents the summary generated by chunk j , and j = 1, 2, …, N.

[0032] . Further, the long text adaptive chunking, summary generation, and semantic embedding are executed in parallel in a pipeline to convert the chunked text into semantic vectors, where: For semantic embedding, the specific method is:

[0033] Based on a semantic embedding model, perform semantic embedding on the obtained text chunk summaries chunk_summary respectively to convert the text into semantic vectors. The specific method is:

[0034] chunk_vector j = Embedding(chunk_summary j ) (4)

[0035] where Embedding is the semantic embedding model, and its function is to convert the text chunk summary chunk_summary jConvert to a high-dimensional semantic vector chunk_vector j , j = 1, 2, …, N.

[0036] Furthermore, the pipeline parallelly executes long text adaptive chunking, summary generation, and semantic embedding, converting the chunked text into semantic vectors, where: The pipeline parallel execution, the specific method is as follows:

[0037] 1) Model loading

[0038] Prepare 3 GPUs, where GPU1 loads the semantic coherence prediction model NSP, GPU2 loads the summary generation model LLM, and GPU3 loads the semantic embedding model Embedding;

[0039] 2) Pipeline parallelism

[0040] Execute long text chunking on GPU1 to obtain chunk 1 , and then transfer chunk 1 to GPU2. Meanwhile, GPU1 performs long text chunking on the remaining text, and GPU2 performs summary generation on chunk 1 to obtain chunk_summary 1 , and then transfer chunk_summary 1 to GPU3. Meanwhile, GPU2 performs summary generation on chunk 2 , and GPU3 performs semantic embedding to obtain chunk_vector 1 ;

[0041] Obtain chunk_vector 2 , chunk_vector 3 , …, chunk_vector N in the above manner.

[0042] Furthermore, perform semantic hierarchical clustering to obtain the top-level cluster summary, which is the long text summary. The specific method is as follows:

[0043] 1) First-layer clustering

[0044] Cluster the text block summary set based on the chunk_vector of the chunked text vector set, so that text block summaries with similar themes are concentrated in the same cluster. If there are multiple summaries in a cluster, use the large language model to summarize them into 1 cluster summary. The generation method of the first-layer cluster summary is:

[0045] cluster_summary 1 , k = LLM({chunk_summary j|chunk_summary j ∈cluster 1,k ) (5)

[0046] Among them, LLM is the large language model for generating the summary, and cluster 1,k represents the kth cluster in the first layer, and cluster_summary 1,k is the cluster summary of cluster 1,k ;

[0047] Then, use the semantic embedding model to perform semantic embedding on the cluster summary:

[0048] cluster_vector 1,k = Embedding(cluster_summary 1,k ) (6)

[0049] Among them, Embedding is the semantic embedding model, and cluster_vector 1, k is the cluster semantic vector of cluster_summary 1,k ;

[0050] 2) Subsequent layer clustering

[0051] Based on the cluster semantic vectors of each layer, perform subsequent layer clustering, and then generate the cluster summaries of the subsequent layers:

[0052] cluster_summary l+1,k =

[0053] LLM({cluster_summary l,m |clustre_summary l,m ∈cluster l+1,k ) (7)

[0054] Among them, clustre_summary l,m represents the mth cluster summary in the lth layer, cluster l+1,k represents the kth cluster in the (l + 1)th layer, cluster_summary l+1,k is the cluster summary of cluster l+1,k , if the number of clustering layers is L, then l = 1, 2,..., L - 1;

[0055] Then, use the semantic embedding model to perform semantic embedding on the cluster summary;

[0056] 3) Obtain the long text summary

[0057] Generate the cluster summaries of each layer layer by layer in the above manner, and the highest-layer cluster summary is the final long-text summary.

[0058] An automatic long-text summary generation system based on adaptive chunking and semantic hierarchical clustering, characterized in that the described automatic long-text summary generation method based on adaptive chunking and semantic hierarchical clustering is implemented to realize automatic long-text summary generation based on adaptive chunking and semantic hierarchical clustering. Specifically, it is divided into three modules, which respectively execute:

[0059] Collect relevant long texts based on web crawler technology and perform standardization, cleaning and normalization processing;

[0060] Execute long-text adaptive chunking, summary generation and semantic embedding in pipeline parallel, and convert the chunked text into semantic vectors;

[0061] Perform semantic hierarchical clustering to obtain the highest-layer cluster summary, which is the long-text summary.

[0062] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the described automatic long-text summary generation method based on adaptive chunking and semantic hierarchical clustering is implemented to realize automatic long-text summary generation based on adaptive chunking and semantic hierarchical clustering.

[0063] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the described automatic long-text summary generation method based on adaptive chunking and semantic hierarchical clustering is implemented to realize automatic long-text summary generation based on adaptive chunking and semantic hierarchical clustering.

[0064] Compared with the prior art, the significant advantages of the present invention are: 1) The adaptive chunking algorithm takes into account both the length of the text chunks and semantic coherence, avoiding the computational and video memory pressure caused by processing long texts at one time, and ensuring that the semantics within the same text chunk are related and the semantics of adjacent text chunks are independent; 2) Pipeline parallelism can make full use of different GPU devices, reduce the idle waiting time, reduce the video memory pressure of a single GPU, and improve the overall computational efficiency; 3) The semantic hierarchical clustering algorithm is beneficial to capturing different themes and hierarchical structures in long texts, making the finally generated summary more accurate and organized, and the summary generation process is also more interpretable. Brief Description of the Drawings

[0065] Figure 1 is the design principle of pipeline parallelism;

[0066] Figure 2 is the semantic hierarchical clustering process;

[0067] Figure 3It is the overall method framework for generating long text abstracts. Detailed implementation manners

[0068] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0069] An automatic long text abstract generation method based on adaptive chunking and semantic hierarchical clustering of the present invention is specifically as follows:

[0070] Step 1, data preprocessing

[0071] 1) Based on technologies such as web crawlers, relevant texts in the field of computer science (hereinafter referred to as "CS texts") are collected from multiple databases and websites (such as Web of Science, IEEE Xplore, Wikipedia, Github, etc.). Since CS texts involve many delicate and complex concepts and principles and have strong integrity and rigor, they are usually relatively long.

[0072] 2) CS texts in different formats (such as HTML, PDF, etc.) collected are uniformly converted into plain text format through Python toolkits such as BeautifulSoup and PyPDF2.

[0073] 3) Identify and standardize professional terms in CS texts. The field of computer science develops rapidly, and new terms constantly emerge in the field, which have high professionalism and certain ambiguity. For example, "central processing unit (CPU)", "deep learning frameworks (such as Caffe, PyTorch, TensorFlow)", "blockchain (Blockchain)", etc. all have multiple aliases. Therefore, a professional term dictionary and a term recognition model are used to locate professional terms in the text and then standardize them into a unified format. This not only reduces data redundancy and ambiguity but also facilitates subsequent semantic analysis and data processing.

[0074] 4) Clean and standardize the above text to achieve denoising. Regular expressions are used to match and delete abnormal characters in the text, including special symbols such as line breaks, tab stops, emojis, etc., and non-critical information irrelevant to the core content that often appears in CS texts, such as copyright statements, reference markings (such as [1], [2], etc.), chart numbers, footnotes, etc., to prevent interference with subsequent abstract generation.

[0075] Step 2, adaptive chunking

[0076] During the pre-training process of BERT, the Next Sentence Prediction (NSP) task enables it to predict whether two sentences are connected, effectively capturing the coherence information between sentences. CS texts often contain long and complex sentences, with a relatively rigorous syntactic logic, numerous professional concepts that are closely related to each other. Therefore, the present invention uses NSP to improve the traditional sentence splitting method, ensuring the semantic coherence of sentences within the same text block and the semantic independence of adjacent text blocks. The specific steps are as follows:

[0077] 1) WordPiece tokenization

[0078] Represent the CS text preprocessed in step 1 as:

[0079] document = {s i | i = 1, 2, …, N s}

[0080] where a CS text document is a set composed of sentences s, the subscript i is the sentence number, and N s is the number of sentences.

[0081] There are a large number of combinations of professional terms in the field of computer science. They may consist of multiple fixed words but are still a semantic whole. For example, "Machine Learning", "Turing Machine", "Data Structure", etc. When these combinations of professional terms appear in the corpus and the frequency reaches a certain threshold (e.g., 100 times), they are combined into a single token and added to the vocabulary. Combining the above method to perform WordPiece tokenization on s, it can be represented as:

[0082]

[0083] where token is the token, and the subscript t is the serial number of the token in the current sentence, is the number of tokens of s i .

[0084] 2) Adaptive chunking algorithm

[0085] Successively splice two adjacent sentences s i and s i+1 into the input sequence of the chunking model. Insert the classification token ([CLS]) at the beginning of the sequence to determine whether adjacent sentences are semantically coherent, and insert the segmentation token ([SEP]) between sentences to distinguish different sentences. Then the input sequence can be represented as:

[0086]

[0087] where \(i = 1, 2, \ldots, N\) s -1, input i The input prediction model can calculate \(s\) i and \(s\) i+1 The probability of semantic coherence:

[0088] \(P\) i = NSP(input i ) (1)

[0089] where NSP is the semantic coherence prediction model of the BERT structure, \(P\) i \(\in [0, 1]\), the closer \(P\) i is to 1, the more semantically related \(s\) i and \(s\) i+1 are, and the more inclined to assign the two to the same text block; conversely, the closer \(P\) i is to 0, the more inclined to assign the two to different text blocks.

[0090] Because there are many professional concepts in CS texts and they are closely related to each other, the context usually has a strong logical and sequential relationship. Therefore, while considering semantic coherence, it is also necessary to avoid the unlimited growth of the text block length caused by a fixed probability threshold. For this reason, according to the difference between the current text block length and the upper limit of the text block length, the probability threshold for chunking is adaptively adjusted:

[0091]

[0092] where maxlen is the upper limit of the text block length set manually, chunk represents the text block, the subscript j is the serial number of chunk, tokennum(chunk j ) is the number of tokens of chunk j , chunk j is initially empty, first add \(s_i\) to chunk j , calculate the probability threshold \(R\) j , \(R\) j \(\in [0.5, 1]\), then compare the coherence probability \(P\) i of \(s\) i+1 and \(s\) i with the threshold. If \(P\) i > \(R\) j , then add \(s\) i+1 to chunk j , then update \(R\) j according to the new tokennum(chunk j ), and use the same method to compare the coherence probability \(P\) i+1 of \(s\) i+2 and \(s\) i+1Whether it is higher than the threshold. If P i+1 >P j , then s i+2 adds chunk j , otherwise s i+2 adds chunk j+1 , and continues to repeat the above operations until i = N s -1, and finally obtains the set of text chunks:

[0093] Chunks = {chunk j |j = 1, 2, …, N}

[0094] where N is the number of chunks obtained from the CS text. For a clearer expression, the above process is abstracted into the pseudocode in Table 1.

[0095] Table 1 Pseudocode of the adaptive chunking algorithm

[0096]

[0097] Step 3, generate summary

[0098] Based on the large language model, understand and summarize the large number of professional concepts and theories contained in the CS text. Concatenate the sentences in each chunk obtained in Step 2 into a string in sequence, and then input it into the LLM to generate a summary. The CS text contains a large number of professional vocabulary, terms, and abbreviations. Through large-scale pre-training on a wide corpus, the LLM has mastered the precise meanings of these professional terms and their contextual associations. The generated summary is:

[0099] chunk_summary j = LLM(chunk j ) (3)

[0100] where LLM is the large language model for generating the summary, and chunk_summary j represents the summary generated by chunk j , j = 1, 2, …, N.

[0101] Step 4, semantic vectorization

[0102] Based on the semantic embedding model, perform semantic vectorization on the chunk_summary in Step 3. The specific method is:

[0103] chunk_vector j = Embedding(chunk_summary j ) (4)

[0104] Among them, Embedding is a semantic embedding model, and its function is to convert the text block summary chunk_summary j into a high-dimensional semantic vector chunk_vector j , where j = 1, 2, …, N.

[0105] Step 5, Steps 2-4 are executed in pipeline parallel.

[0106] Adaptive chunking needs to be performed on a per-text-block basis. This sequential dependence causes Steps 2-4 not to be fully parallelizable. Therefore, this invention draws on the multi-instruction execution principle of the CPU and adopts a pipeline parallel approach to execute Steps 2-4, reducing the video memory pressure and improving the processing efficiency. It is specifically divided into the following steps:

[0107] 1) Model loading

[0108] Prepare 3 GPUs. Among them, GPU1 loads the semantic coherence prediction model NSP of Step 2, GPU2 loads the abstract generation model LLM of Step 3, and GPU3 loads the semantic embedding model Embedding of Step 4.

[0109] 2) Pipeline parallel

[0110] The core design is as Figure 1 shown. First, execute Step 2 on GPU1 to obtain chunk 1 , then transfer chunk 1 to GPU2. At the same time, GPU1 executes Step 2 on the remaining text, and GPU2 executes Step 3 on chunk 1 to obtain chunk_summary 1 , then transfer chunk_summary 1 to GPU3. At the same time, GPU2 executes Step 3 on chunk 2 , and GPU3 executes Step 4 to obtain chunk_vector 1 ; Continuing the calculation in the above manner, chunk_vector 2 , chunk_vector 3 , …, chunk_vector N can be obtained in sequence.

[0111] Step 6, semantic hierarchical clustering

[0112] 1) The first layer of clustering

[0113] Using a clustering algorithm, cluster the chunk_summary based on the chunk_vector obtained in step 5, so that the text chunk summaries with similar themes are concentrated in the same cluster. If there are multiple summaries in a cluster, these summaries need to be concatenated into a string, and then summarized into 1 cluster summary by the LLM. The generation method of the cluster summary in the first layer is as follows:

[0114] cluster_summary 1,k = LLM({chunk_summary j |chunk_summary j ∈cluster 1,k}) (5)

[0115] where LLM is the large language model for generating summaries, and cluster 1,k represents the k-th cluster in the first layer, and cluster_summary 1,k is the cluster summary of cluster 1,k . Then, perform semantic embedding on the cluster summary:

[0116] cluster_vector 1,k = Embedding(cluster_summary 1,k ) (6)

[0117] where Embedding is the semantic embedding model, and cluster_vector 1,k is the cluster semantic vector of cluster_summary 1,k . With the cluster semantic vector of this layer, subsequent layer clustering can be performed.

[0118] 2) Subsequent layer clustering

[0119] Based on the cluster semantic vectors of each layer, perform subsequent layer clustering, and then generate the cluster summaries of the subsequent layers:

[0120] cluster_summary l+1,k = LLM({cluster_summary l,m |cluster_summary l,m ∈cluster l+1,k}) (7)

[0121] where cluster_summary l,m represents the m-th cluster summary in the l-th layer, cluster l+1,k represents the k-th cluster in the (l + 1)-th layer, and cluster_summary l+1,kIt is the cluster summary of the cluster l+1,k . If the number of clustering layers is L, then l = 1, 2, …, L - 1. l+1,k

[0122] 3) Obtain the final summary of the CS text

[0123] Generate the cluster summaries of each layer layer by layer in the above manner. The highest - layer cluster summary is the final summary of this CS text. Figure 2 It is an example of generating a text summary from 5 chunk_summaries.

[0124] In summary, combined with natural language processing technologies such as LLM, the present invention can accurately and efficiently generate summaries for CS texts; improve the semantic coherence of text chunks through adaptive chunking, optimize the overall processing efficiency through pipeline parallelism, and improve the accuracy and coherence of summaries through semantic hierarchical clustering, making the present invention applicable to the application scenarios of summary generation.

[0125] Embodiment

[0126] To verify the effectiveness of the solution of the present invention, the following experiment is carried out.

[0127] First, obtain the long text for which a summary is to be generated. This is a noisy long text collected from the Internet: "The Transformer model (literally translated as <em>‘Converter’< / em> ) is a deep - learning model that adopts an attention mechanism, and this mechanism can assign different weights according to the different importance of each part of the input data. This model is mainly used in the fields of natural language processing (NLP) and computer vision (CV). [1] ……\nPicture [image:https: / / www.xxx.cn / xxx / 640?wx_fmt=png&;from=appmsg&;tp=webp&;wxfrom=5&;wx_lazy=1&;wx_co=1]\n■ Each attention head represents the attention between different tokens, and multiple attention heads can target different <em>‘Correlation’< / em> , calculate different attention weights... References [Edit]...", and a summary needs to be generated for this long text mixed with noise.

[0128] Step 1. Data pre - processing. Denoise the above long text to obtain a high - quality text: "The Transformer model (literally translated as 'transformer') is a deep - learning model that adopts an attention mechanism, and this mechanism can assign different weights according to the different importance of each part of the input data. This model is mainly used in the fields of natural language processing (NLP) and computer vision (CV).... Each attention head represents the attention between different tokens, and multiple attention heads can target different'relevances' and calculate different attention weights...".

[0129] Step 2. Adaptive chunking. Perform adaptive chunking on the high-quality text obtained in Step 1, and set the length upper limit maxlen to 256.

[0130] Step 3. Generate summaries for each chunk obtained in Step 2.

[0131] Step 4. Semantic embedding. Perform semantic embedding on the chunk_summary obtained in Step 3 and convert it into a 1024-dimensional semantic vector.

[0132] Step 5. Execute Steps 2-4 in a pipeline parallel manner, and the obtained results are shown in Tables 2-4 respectively.

[0133] Table 2 Adaptive Chunking Results

[0134]

[0135]

[0136] Table 3 Text Chunk Summary Results

[0137]

[0138] Table 4 Semantic Embedding Results

[0139] Semantic vector Content <![CDATA[chunk_vector 1 > [-0.0058,0.0284,0.0038,...,-0.0497,0.0238,-0.0588] <![CDATA[chunk_vector 2 > [0.0508,0.0475,-0.0226,...,-0.0710,0.0133,-0.0266] <![CDATA[chunk_vector 3 > [0.0219,0.0760,-0.0137,...,-0.0222,0.0160,-0.0308] <![CDATA[chunk_vector 4 > [-0.0033,0.0065,-0.0253,...,-0.0183,-0.0056,-0.0451] …… ……

[0140] Step 6. Semantic hierarchical clustering to obtain the long text summary. After semantic hierarchical clustering, the final long text summary is: "The Transformer model, with the attention mechanism at its core, is widely used in the fields of NLP and CV. Its ability to process all inputs at once significantly reduces the training time and is particularly suitable for tasks such as translation. The model was introduced by the Google Brain team in 2017 and has promoted the development of pre-trained models with its parallelization advantages. Transformer adopts an encoder-decoder architecture, where the encoder processes the input layer by layer and the decoder generates the output. Although the memory requirement of the original model grows quadratically with the context window size, its performance is superior and it has been implemented in frameworks such as TensorFlow and PyTorch."

[0141] After testing, the present invention can significantly improve the accuracy and efficiency of long text summaries.

[0142] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0143] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for automatically generating long text summaries based on adaptive chunking and semantic hierarchical clustering, characterized in that: The specific steps are as follows: Collect relevant long texts based on web crawler technology, and perform professional terminology recognition, standardization, cleaning and normalization processing; The pipeline performs adaptive segmentation, summary generation, and semantic embedding of long texts in parallel, converting segmented texts into semantic vectors. Perform semantic hierarchical clustering to obtain the highest-level cluster summary, which is the long text summary.

2. The method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering according to claim 1, characterized in that: Based on web crawler technology, relevant long texts are collected to identify, standardize, clean and normalize professional terms. The specific methods are as follows: 1) Based on web crawler technology, collect relevant texts in the field of computer science from multiple databases and websites, including Web of Science, IEEE Xplore, Wikipedia, and Github; 2) The collected long texts in different formats are uniformly converted into plain text format through the Python toolkit; 3) Identify and standardize professional terms in long texts to reduce data redundancy and ambiguity; 4) Clean and normalize the standardized long text, use regular expressions to match and delete abnormal characters in the text, including line breaks, tabs, emojis and other special symbols, identify and delete copyright statements, reference notes, figure numbers, and footnotes in the text.

3. The method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering according to claim 1, characterized in that: The pipeline performs long text adaptive segmentation, summary generation and semantic embedding in parallel, converting segmented text into semantic vectors, among which: long text adaptive segmentation, the specific method is: 1) WordPiece lemmatization After preprocessing, the long text is represented as document={s i |i=1,2,…,N s }, where the long text document is a set of sentences s, the subscript i is the sentence number, N s is the number of sentences; after WordPiece tokenization, the sentence consists of tokens, represented as The subscript t is the token's serial number in the current sentence. For i The number of tokens; 2) Adaptive Blocking Sequentially put two adjacent sentences s i and i+1 The input sequence of the block model is concatenated. A classification token ([CLS]) is inserted at the beginning of the sequence to determine whether adjacent sentences are semantically coherent. A segmentation token ([SEP]) is inserted between sentences to distinguish different sentences. The input sequence is represented as follows: where i=1,2,…,N s -1; Using the semantic coherence prediction model based on the BERT structure, we calculate s i and i+1 Probability of semantic coherence: P i =NSP(input i ) (1) Among them, NSP is a semantic coherence prediction model based on the BERT structure. i ∈[0,1],P i The closer to 1, the more s i and i+1 The more relevant, the more likely they are to be assigned to the same text block. i The closer it is to 0, the more likely it is to assign the two to different text blocks; In order to consider semantic coherence and avoid unlimited growth of text blocks, the upper limit of text block length is specified, and the threshold of block segmentation is adaptively adjusted based on the current text block length: Among them, maxlen is the upper limit of the length of the text block set manually, chunk represents the text block, subscript j is the sequence number of the chunk, tokennum(chunk j ) is chunk j The number of tokens, chunk j Initially empty, first set s i Add chunk j , calculate the probability threshold R j , R j ∈[0.5,1], then compare s i and i+1 Coherence probability P i Is it higher than the threshold? If P i >R j , then s i+1 Add chunk j , and then according to the new tokennum(chunk j ) Update R j , using the same method to compare s i+1 and i+2 Coherence probability P i+1 Is it higher than the threshold? If P i+1 >R j , then s i+2 Add chunk j , otherwise s i+2 Add chunk j+1 , continue to repeat the above operation until i = N s -1, and finally get the set of text blocks Chunks = {chunk j |j=1,2,…,N}, where N is the number of chunks obtained for the long text.

4. The method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering according to claim 3, characterized in that: The pipeline performs adaptive segmentation, summary generation, and semantic embedding of long texts in parallel, converting segmented texts into semantic vectors, including summary generation. The specific method is: Based on the large language model, a summary is generated for each chunk in the text block set Chunks, expressed as: chunk_summary j =LLM(chunk j ) (3) Among them, LLM is the large language model for generating summaries, chunk_summary j Represents chunk j The generated summary, j=1,2,…,N.

5. The method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering according to claim 4, characterized in that: The pipeline performs adaptive segmentation, summary generation and semantic embedding of long text in parallel, converting segmented text into semantic vectors, including semantic embedding. The specific method is: Based on the semantic embedding model, the obtained text chunk summary chunk_summary is semantically embedded and the text is converted into a semantic vector. The specific method is as follows: chunk_vector j =Embedding(chunk_summary j ) (4) Among them, Embedding is a semantic embedding model, which is used to embed the text block summary chunk_summary j Converted into high-dimensional semantic vector chunk_vector j , j=1,2,…,N.

6. The method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering according to claim 5, characterized in that: The pipeline executes adaptive segmentation, summary generation and semantic embedding of long text in parallel, and converts segmented text into semantic vectors. The specific method is as follows: 1) Model loading Prepare 3 GPUs, GPU1 loads the semantic coherence prediction model NSP, GPU2 loads the summary generation model LLM, and GPU3 loads the semantic embedding model Embedding; 2) Pipeline parallelism Execute long text chunking on GPU1 to get chunk1, then transfer chunk1 to GPU2. Meanwhile, GPU1 performs long text chunking on the remaining text. GPU2 performs summary generation on chunk1 to get chunk_summary1, then transfer chunk_summary1 to GPU3. Meanwhile, GPU2 performs summary generation on chunk2, and GPU3 performs semantic embedding to get chunk_vector1. According to the above method, we can obtain chunk_vector2, chunk_vector3, …, chunk_vectorN in turn.

7. The method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering according to claim 5, characterized in that: Perform semantic hierarchical clustering to obtain the highest-level cluster summary, which is the long text summary. The specific method is: 1) Layer 1 Clustering The text block summary set is clustered based on the chunk_vector text vector set, so that the text block summaries with similar topics are concentrated in the same cluster. If there are multiple summaries in a cluster, they are summarized into one cluster summary using a large language model. The generation method of the first-level cluster summary is: cluster_summary 1,k =LLM({chunk_summary j |chunk_summary j ∈cluster 1,k }) (5) Among them, LLM is a large language model for generating summaries, cluster 1,k Indicates the kth cluster in the first layer, cluster_summary 1,k It is cluster 1,k Cluster summary of ; Then use the semantic embedding model to semantically embed the cluster summary: cluster_vector 1,k =Embedding(cluster_summary 1,k ) (6) Among them, Embedding is a semantic embedding model, cluster_vector 1,k cluster_summary 1, k's cluster semantic vector; 2) Subsequent layer clustering Based on the cluster semantic vector of each layer, clustering of subsequent layers is performed, and then the cluster summary of subsequent layers is generated: cluster_summary l+1,k =LLM({cluster_summary l,m |cluster_summary l,m ∈cluster l+1,k }) (7) Among them, cluster_summary l,m represents the summary of the mth cluster at the lth layer, cluster l+1,k Indicates the kth cluster in the l+1th layer, cluster_summary l+1,k It is cluster l+1,k Cluster summary, if the number of clustering levels is L, then l = 1, 2, ..., L-1; Then the semantic embedding model is used to semantically embed the cluster summary rows; 3) Get a long text summary The cluster summaries of each layer are generated layer by layer in the above manner, and the highest layer cluster summary is the final long text summary.

8. A long text summary automatic generation system based on adaptive chunking and semantic hierarchical clustering, characterized in that: The method for automatically generating a long text summary based on adaptive block segmentation and semantic hierarchical clustering according to any one of claims 1 to 7 is implemented to realize the automatic generation of a long text summary based on adaptive block segmentation and semantic hierarchical clustering, which is specifically divided into three modules and respectively performs: Collect relevant long texts based on web crawler technology, and perform standardization, cleaning and normalization processing; The pipeline performs adaptive segmentation, summary generation, and semantic embedding of long texts in parallel, converting segmented texts into semantic vectors. Perform semantic hierarchical clustering to obtain the highest-level cluster summary, which is the long text summary.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering as described in any one of claims 1 to 7 is implemented to realize automatic generation of long text summaries based on adaptive block segmentation and semantic hierarchical clustering.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for automatically generating long text summaries based on adaptive block segmentation and semantic hierarchical clustering according to any one of claims 1 to 7 is implemented to achieve automatic generation of long text summaries based on adaptive block segmentation and semantic hierarchical clustering.