Program generation model design method and device, storage medium and program product

By combining information about confusion, semantic similarity, syntactic structure and logical boundary detection, we optimize paragraph segmentation and outline generation of long texts, solving the problems of information fragmentation and excessive consumption of computing resources in the existing technology, and achieving efficient and accurate outline generation.

CN119962673APending Publication Date: 2025-05-09HUA DATA TECH (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510019761.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When processing long texts, the prior art relies on semantic similarity to paragraph division and outline generation, resulting in information fragmentation and excessive consumption of computing resources, affecting model performance and scalability.

Method used

A variety of information-combining methods are adopted, including confusion analysis, semantic segmentation, syntactic structure analysis and logical boundary detection, to generate comprehensive scores, optimize paragraph segmentation, and construct a multi-level outline through large language models.

Benefits of technology

It improves the efficiency and accuracy of outline generation, reduces the computational burden, enhances the robustness and interpretability of the model, and is suitable for multiple application fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962673A_ABST
    Figure CN119962673A_ABST
Patent Text Reader

Abstract

The invention provides an outline generation model design method and device, a storage medium and a program product, and relates to the technical field of text understanding and processing. The long text outline generation model design method comprises the steps that a confusion degree score is obtained through a confusion degree analysis module; semantic similarity and syntactic structure information are obtained through a semantic segmentation module; generating a comprehensive score according to the confusion score, the semantic similarity and the syntactic structure information; performing preliminary paragraph segmentation through a logic boundary detection module; optimizing the preliminary paragraph segmentation according to the comprehensive score; a multi-level outline is generated using a large language model. Compared with the prior art, the method is more efficient and accurate, the integration degree is higher, the interpretability is higher, and the method can be popularized in multiple application fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text understanding and processing, and in particular to a long text outline generation model design method, device, storage medium and program product. Background Art

[0002] The goal of text understanding and processing technology is to help users quickly grasp the rich information and complex logic contained in long texts, so as to quickly understand the core content of the text. In the process of users understanding long texts, the text outline plays an absolutely critical role. It summarizes the text while presenting the text structure. For example, academic papers that researchers need to read frequently are often difficult to understand due to their complex structure, unclear outlines, and intertwined terms and concepts; for materials such as handouts, manuals, and instructions, it is difficult to quickly locate the required content without a structured outline, which affects the effectiveness of the use of the materials; for news information involving multiple events and background information, if there is no clear outline, it is difficult for readers to quickly obtain the key points of the information.

[0003] Existing technologies directly identify semantic similarity based on large language models, and then perform paragraph division and outline generation based on semantic similarity. However, segmentation based solely on semantic similarity will lead to information fragmentation and affect users' understanding of the overall structure of the text. The reason is that in practical applications, many paragraphs with low semantic similarity should actually belong to the same topic. For example, in a news report, paragraphs describing the disaster situation, rescue operations, and government response measures of a natural disaster, although the semantic similarity between these paragraphs may not be high, they together constitute a complete reporting process, and therefore should not be divided into multiple independent paragraphs.

[0004] There is another problem with semantic paragraph segmentation and outline generation directly based on a large language model. When processing very long texts, in order to ensure coherence and logic, a large amount of computing resources are required, which can easily lead to slow model operation and even memory overflow. This not only affects the overall performance and processing efficiency of the model, but also limits the scalability and flexibility of the model in practical applications. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention proposes an outline generation model design method, device, storage medium and program product. The segmentation result obtained by the present invention not only takes into account the results of logical boundary detection, but also combines various information such as the perplexity of the language model, text semantic similarity and syntactic structure; the present invention not only generates concise and clear titles for long texts, but also constructs a hierarchical outline structure. In terms of operating effect, the present invention quickly filters out those situations that obviously do not require segmentation, reduces the computational burden, and improves overall efficiency; through a multi-level decision-making mechanism, the robustness of the model is enhanced and the errors caused by a single method are reduced.

[0006] Compared with the existing technology, the method of the present invention is more efficient and accurate, has a higher degree of integration, is more interpretable, and can be promoted in multiple application fields.

[0007] In a first aspect, the present invention provides a method for designing a long text outline generation model, comprising the following steps:

[0008] Obtain the perplexity score through the perplexity analysis module;

[0009] The semantic similarity and syntactic structure information are obtained through the semantic segmentation module;

[0010] Generate a comprehensive score based on the perplexity score, semantic similarity, and syntactic structure information;

[0011] Perform preliminary paragraph segmentation through the logical boundary detection module;

[0012] Optimize the initial paragraph segmentation based on the comprehensive score;

[0013] Generate multi-level outlines using large language models.

[0014] As a further improvement of the present invention, obtaining the perplexity score through the perplexity analysis module includes:

[0015] Deconstruct a long text into a series of sentences through text preprocessing;

[0016] Use a large language model to calculate the perplexity score of a single sentence;

[0017] Use a large language model to calculate the perplexity score of sentence combinations;

[0018] Preliminary screening of potential split points based on the perplexity score.

[0019] The present invention measures the coherence and logical relevance of sentences and sentence groups through a perplexity analysis module, and then evaluates the language complexity of text segments to assist in determining reasonable text segmentation points.

[0020] As a further improvement of the present invention, the obtaining of semantic similarity and syntactic structure information through the semantic segmentation module includes:

[0021] Use the text representation model to embed sentences and calculate the semantic similarity;

[0022] Use text analysis tools to perform syntactic structure analysis to obtain syntactic structure information.

[0023] The semantic segmentation module integrates multiple strategies such as semantic similarity calculation and syntactic structure analysis to ensure that sentences within the same topic are kept together as much as possible and generate logically clear text paragraphs.

[0024] As a further improvement of the present invention, generating a comprehensive score according to the perplexity score, semantic similarity, and syntactic structure information includes:

[0025] Integrate the perplexity score, semantic similarity, and syntactic structure information through a comprehensive score;

[0026] Set the threshold for the composite score.

[0027] The threshold of the comprehensive score can be set through experiments and experience, and can be adjusted according to specific tasks and datasets. The comprehensive score also covers the text's perplexity score, semantic similarity, and syntactic structure information.

[0028] As a further improvement of the present invention, the preliminary paragraph segmentation by the logical boundary detection module includes:

[0029] Use a large language model to evaluate the probability difference between segmentation and non-segmentation between consecutive sentences;

[0030] Set a difference threshold;

[0031] Preliminary paragraph segmentation is performed by comparing the probability difference with the difference threshold.

[0032] As a further improvement of the present invention, the optimization of the preliminary paragraph segmentation according to the comprehensive score includes:

[0033] For situations where a clear segmentation result can be obtained based on the comparison of the probability difference with the difference threshold, a decision is made directly;

[0034] In cases where it is not obvious to arrive at a segmentation result, the comprehensive score is used to verify and optimize the preliminary paragraph segmentation.

[0035] As a further improvement of the present invention, the step of generating a multi-level outline using a large language model includes:

[0036] Generate paragraph titles;

[0037] The title list is grouped and summarized in multiple levels to generate a multi-level title structure.

[0038] In a second aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.

[0039] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0040] In a fourth aspect, the present invention provides a computer program product, which implements the steps of the method described in the first aspect when the computer program is executed by a processor.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The present invention proposes a design method, device, storage medium and program product for a long text outline generation model. The segmentation result obtained by the present invention not only takes into account the result of logical boundary detection, but also combines various information such as the perplexity of the language model, the text semantic similarity and the syntactic structure; the present invention not only generates concise and clear titles for long texts, but also constructs a hierarchical outline structure. In terms of operating effect, the present invention quickly filters out those situations that obviously do not require segmentation, reduces the computational burden, and improves overall efficiency; through a multi-level decision-making mechanism, the robustness of the model is enhanced and the errors caused by a single method are reduced.

[0043] Compared with the existing technology, the method of the present invention is more efficient and accurate, has a higher degree of integration, is more interpretable, and can be promoted in multiple application fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 The present invention discloses a flow chart of a method for designing an outline generation model.

[0045] Figure 2 The present invention discloses a flow chart of a text paragraph segmentation method.

[0046] Figure 3 The invention discloses a flow chart for generating a multi-level outline for a text segmented into paragraphs. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the present invention will combine the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments, wherein steps S1, S2... in the embodiments described in the present invention do not limit the only execution steps of the present invention; the various models, simulation environments, and software described in the present invention are not the only limiting methods of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0048] In the present invention, computer devices / equipment / systems refer to related entities applied to computers, such as hardware, a combination of hardware and software, software or software in execution, etc. In detail, for example, software includes but is not limited to a process running on a processor, a processor, an object, executable software, an execution thread, a program and / or a computer. In addition, an application or script program running on a server, a server can also be software. One or more software can be in an execution process and / or thread, and the software can be localized on one computer and / or distributed between two or more computers, and can be run by various computer-readable media.

[0049] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other.

[0050] In a first aspect, the present invention provides an embodiment of a method for designing a long text outline generation model. Figure 1 As shown, the specific process can be as follows:

[0051] S1: Obtain text segmentation strategy and perplexity score through the perplexity analysis module.

[0052] Preferably, the Qwen-7B large language model is used to identify the logically closely connected parts in the text by calculating the perplexity score of each sentence or sentence combination, and a text segmentation strategy is obtained based on this.

[0053] The specific steps include:

[0054] S11: Deconstruct the long text into a series of sentences through text preprocessing. This step ensures that the subsequent analysis can be performed at the sentence level, not just based on words or characters.

[0055] S12: Calculate the perplexity score for a single sentence. The model calculates the perplexity score for each sentence one by one. By taking each sentence as input, the Qwen-7B model tries to predict the next word. The model assigns a probability to each possible next word based on its internal parameters and the knowledge gained from training. The perplexity score for the entire sentence is calculated based on these probabilities.

[0056] S13: Calculate the perplexity score for sentence combinations. In addition to calculating the perplexity score for each sentence individually, the model also considers the perplexity score of short text segments consisting of multiple consecutive sentences. This process involves merging two or more sentences into a new text unit and repeating the perplexity score calculation steps described above. The purpose of this is to find groups of sentences that maintain high linguistic coherence even when put together.

[0057] S14: Determine the split point by comparing the perplexity scores of different sentences and their combinations. The model can identify which sentence combinations have high coherence and logical relevance. If a sentence is added to the previous sentence and the overall perplexity score drops significantly, then the two sentences are likely to belong to the same logical unit; conversely, if the perplexity score increases significantly after adding, it may mean that this is a suitable split point.

[0058] S2: Semantic similarity and syntactic structure information are obtained through the semantic segmentation module.

[0059] The specific steps include:

[0060] S21: Use the text representation model to embed sentences. That is, convert each sentence into a high-dimensional vector representation for subsequent calculation of semantic similarity. Input each sentence into the model, and the model will output a fixed-dimensional vector that can capture the semantic information of the sentence.

[0061] Preferably, the BGE-M3 text representation model is used. BGE-M3 is a pre-trained language model based on the Transformer architecture, which is specifically used to generate high-quality sentence embeddings. It captures the contextual information of sentences through a bidirectional generative model, can more accurately represent the semantic meaning of sentences, and is suitable for various natural language processing tasks, such as sentence similarity calculation. The BGE-M3 model supports more than 100 languages ​​and performs well in multilingual text processing tasks. The BGE-M3 model can process up to 8192 tokens, which is suitable for processing long texts, ensuring that high-precision embedding representations can be maintained in large-scale texts.

[0062] S22: Perform semantic similarity calculation. Use sentence embedding vectors to calculate the cosine similarity of adjacent sentences to evaluate the semantic relevance between adjacent sentences. For each pair of adjacent sentences, calculate the cosine similarity between their embedding vectors. The closer the similarity value is to 1, the more similar the semantics of the two sentences are.

[0063] S23: Perform syntactic structure analysis and extract grammatical components. Use the syntactic parser Spacy to perform sentence structure analysis and extract syntactic structure information such as subject, predicate, and object to help determine the logical relationship between sentences. For Chinese text, the Spacy parser uses the zh_core_web_sm model to parse. Each sentence is input into the Spacy model, and the model outputs the sentence's dependency tree or syntactic tree. Extract syntactic structure information such as subject, predicate, and object from the syntactic tree. This information can help identify the structural characteristics of the sentence and further assist segmentation decisions.

[0064] Spacy is an efficient natural language processing library that provides a variety of pre-trained models for tasks such as word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing. zh_core_web_sm is a Chinese model provided by Spacy, specifically for processing Chinese text. This model can accurately perform Chinese word segmentation, part-of-speech tagging, and dependency parsing, and is suitable for various Chinese natural language processing tasks.

[0065] S3: Generate a comprehensive score based on the perplexity score, semantic similarity, and syntactic structure information. That is, the comprehensive score comes from three aspects:

[0066] a. Perplexity score: Get the perplexity score of each sentence from the perplexity analysis module.

[0067] b. Semantic similarity: Obtain the cosine similarity of adjacent sentences from the semantic similarity calculation step.

[0068] c. Syntactic structure: Obtain the subject, predicate, object and other grammatical components of the sentence from the syntactic structure analysis steps.

[0069] like Figure 3 As shown in Figure 2, the comprehensive score provides a basis for subsequent segmentation decisions.

[0070] The threshold of the comprehensive score can be set through experiments and experience, and can be adjusted according to specific tasks and datasets. The comprehensive score also covers the text's perplexity score, semantic similarity, and syntactic structure information.

[0071] S4: Preliminary paragraph segmentation through logical boundary detection module

[0072] S41: Using a large language model to evaluate the probability difference between segmentation and non-segmentation between consecutive sentences.

[0073] Preferably, the Qwen-7B large language model is used. Each pair of consecutive sentences is input into the Qwen-7B model. In order to better guide the model to evaluate the segmentation probability of the sentence pair, it is necessary to construct appropriate prompt words.

[0074] Preferably, the prompt word is: "Output the probabilities of the following two options (segmentation indication option and non-segmentation indication option), and calculate the difference between the two probability values:

[0075] 1. Option 1: Divide the sentence combination of sentence 1 + sentence 2 into two parts: sentence 1 and sentence 2;

[0076] 2. Option 2: Keep the sentence combination of sentence 1 + sentence 2 without segmentation. "

[0077] Input the constructed prompt words into the Qwen-7B model, and the model will output the probability value of the sentence pair segmentation / non-segmentation, and calculate the difference between the two probability values. This difference value reflects the tendency of whether the sentence pair should be segmented.

[0078] S42: Setting a difference threshold. The difference threshold is set based on experiments and experience, and can be adjusted based on specific tasks and data sets.

[0079] S43: Perform preliminary paragraph segmentation by comparing the probability difference with the difference threshold. Compare the calculated segmentation probability difference with the difference threshold. If the probability value difference is greater than the difference threshold, the sentence pair should be segmented to form a new paragraph, otherwise it remains merged.

[0080] S5: Optimize the preliminary paragraph segmentation based on the comprehensive score.

[0081] S51: By setting and filtering thresholds, a direct decision is made for very obvious segmentation or non-segmentation situations, that is, the step of using comprehensive scores for optimization is skipped, the computational burden is reduced, and the response speed and overall efficiency of the system are improved.

[0082] S52: For texts not filtered out by step S51, the preliminary paragraph segmentation is verified and optimized according to the comprehensive score to ensure the accuracy of the final result. Preferably, the comprehensive score is compared with the result of the preliminary paragraph segmentation to select the best result.

[0083] The comprehensive score itself can reflect the paragraph segmentation results obtained based on the perplexity score, semantic similarity, and syntactic structure information.

[0084] S6: Generate multi-level outlines using large language models.

[0085] Preferably, the Qwen-7B model is used. The specific steps include:

[0086] S61: Generate paragraph titles. Generate a short title for each segmented text segment to ensure that the title accurately reflects the core content of the paragraph.

[0087] Each segmented text segment is input into the large language model, and appropriate prompt words are constructed to guide the model to generate a concise and clear title. Preferably, the prompt word is "Please generate a concise and clear title for the following text paragraph, and ensure that the title can accurately reflect the core content of the paragraph:".

[0088] S62: Grouping and multi-level summarizing the title list in sequence to generate a multi-level title structure.

[0089] Construct appropriate prompt words to guide the Qwen-7B large language model to group the generated title list. Preferably, the prompt words are "Please make a preliminary grouping of the following title list and group titles with similar topics together:". Summarize all generated titles and grouping results into a list.

[0090] Construct appropriate prompt words to guide the Qwen-7B large language model to perform multi-level summaries on the grouped title list and generate a multi-level title structure. Preferably, the prompt words are "Please perform multi-level summaries on the following grouped title list to generate first-level titles, second-level titles, etc. until the required level depth is reached."

[0091] Embodiment 1:

[0092] S1: Get the perplexity score through the perplexity analysis module.

[0093] like Figure 2 As shown in the figure, the perplexity score of the combination of sentences 2 and 3 increased slightly from 1.8 to 2.1 compared with sentence 2, which may be a potential split point. The perplexity score of the combination of sentences 4 and 5 increased significantly from 2.9 to 3.5 compared with sentence 4, indicating that the coherence between the two sentences is poor and it is a suitable split point. The perplexity scores of other sentence combinations decreased or did not change much compared with single sentences, so they are not split points.

[0094] S2: Semantic similarity and syntactic structure information are obtained through the semantic segmentation module.

[0095] like Figure 3 As shown, sentence 1 and sentence 2 have different subjects and predicates, but related objects; sentence 2 and sentence 3 have different subjects and predicates, but related objects; sentence 3 and sentence 4 have different subjects and predicates, but related objects; sentence 4 and sentence 5 have different subjects and predicates, but unrelated objects. The syntactic similarity between sentences is calculated based on the above syntactic information.

[0096] S3: Generate a comprehensive score based on the perplexity score, semantic similarity, and syntactic structure information;

[0097] S4: Preliminary paragraph segmentation is performed through the logical boundary detection module;

[0098] S5: Optimize the preliminary paragraph segmentation based on the comprehensive score.

[0099] like Figure 2 As shown in the figure, the difference between the comprehensive scores output by the comprehensive evaluation and analysis module and the semantic segmentation module and the sentence segmentation probability output by the logical boundary detection module can ultimately determine that the segmentation points of the long text are between sentences 2 and 3, and between sentences 4 and 5, and organize the long text sentences into paragraphs based on these segmentation points.

[0100] S6: Generate multi-level outlines using large language models.

[0101] S61: Figure 3 As shown, according to the input text paragraph, three titles can be output according to the text paragraph: virus discovery and spread, response measures from all walks of life, and economic impact.

[0102] S62: Figure 3 As shown, the three titles output in the previous step are divided into two groups, one of which includes "Virus Discovery and Spread" and "Response Measures from All Sector", and the other includes "Economic Impact".

[0103] like Figure 3 As shown in the figure, the large language model generates a multi-level outline for the two sets of titles output in the previous step. The final outline has clear content levels and can help users quickly understand and process long texts.

[0104] In a second aspect, the present invention provides an embodiment of a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.

[0105] In a third aspect, the present invention provides an embodiment of a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processor.

[0106] In a fourth aspect, the present invention provides an embodiment of a computer program product, which implements the steps of the method described in the first aspect when the computer program is executed by a processor.

[0107] The present invention proposes a long text outline generation model design method, device, storage medium and program product, which integrates the perplexity score, semantic similarity information and syntactic structure information at the sentence scale of the text, and then obtains the optimized text segmentation result and multi-level text outline through logical boundary detection. Compared with the prior art, the method of the present invention is more efficient and accurate, has a higher degree of integration, is more interpretable, and can be promoted in multiple application fields.

Claims

1. A method for designing an outline generation model, characterized in that: The following steps are involved: Obtain the perplexity score through the perplexity analysis module; The semantic similarity and syntactic structure information are obtained through the semantic segmentation module; Generate a comprehensive score based on the perplexity score, semantic similarity, and syntactic structure information; Perform preliminary paragraph segmentation through the logical boundary detection module; Optimize the initial paragraph segmentation based on the comprehensive score; Generate multi-level outlines using large language models.

2. The method according to claim 1, characterized in that The obtaining of the perplexity score by the perplexity analysis module includes: Deconstruct a long text into a series of sentences through text preprocessing; Use a large language model to calculate the perplexity score of a single sentence; Use a large language model to calculate the perplexity score of sentence combinations; Preliminary screening of potential split points based on the perplexity score.

3. The method according to claim 1, characterized in that The semantic similarity and syntactic structure information are obtained through the semantic segmentation module, including: Use the text representation model to embed sentences and calculate the semantic similarity; Use text analysis tools to perform syntactic structure analysis to obtain syntactic structure information.

4. The method according to claim 1, characterized in that: Generating a comprehensive score according to the confusion score, semantic similarity, and syntactic structure information includes: Integrate the perplexity score, semantic similarity, and syntactic structure information through a comprehensive score; Set the threshold for the composite score.

5. The method according to claim 1, characterized in that The preliminary paragraph segmentation is performed by the logical boundary detection module, including: Use a large language model to evaluate the probability difference between segmentation and non-segmentation between consecutive sentences; Set a difference threshold; Preliminary paragraph segmentation is performed by comparing the probability difference with the difference threshold.

6. The method according to claim 1 or 5, characterized in that: The optimization of the preliminary paragraph segmentation according to the comprehensive score includes: For situations where a clear segmentation result can be obtained based on the comparison of the probability difference with the difference threshold, a decision is made directly; In cases where it is not obvious to arrive at a segmentation result, the comprehensive score is used to verify and optimize the preliminary paragraph segmentation.

7. The method according to claim 1, characterized in that The method of using a large language model to generate a multi-level outline includes: Generate paragraph titles; The title list is grouped and summarized in multiple levels to generate a multi-level title structure.

8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Pancreatic cancer prediction method and system based on local and global confusion weighted pruning

    CN121839088A