Letter text analysis method driven by multilayer dynamic aggregation large language model

By performing time stamp sequence conversion, semantic analysis and structural processing on multi-format function data, dynamic bridge diagrams are constructed, which solves the problems of information fragmentation and incoherence in complex function text analysis in the existing technology, and realizes efficient semantic correlation and context understanding, improving the accuracy and efficiency of function text analysis.

CN120387459AActive Publication Date: 2025-07-29ANHUI SHENHE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510524379.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-29
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process multi-format letter texts, especially in the case of complex discourse structure, strong context dependence, and irregular language style, it lacks the ability to model implicit associations between context, timing, and context, resulting in fragmented information extraction, incoherent response, and poor expansion capabilities.

Method used

By obtaining multi-format function data and converting it into a time-stamped text sequence, semantic analysis and structural processing are performed, a collection of text blocks containing timing and context is generated, multi-task training is performed and topic fusion is carried out, and dynamic bridge diagram is constructed. In response to user queries, the activated text block content is spliced into context Prompt and a large language model is input to obtain one-time generated answers, and the bridge diagram is updated based on the new function data.

Benefits of technology

It realizes efficient capture of complex semantic associations, improves context understanding coherence, realizes dynamic information aggregation, improves the integrity of information extraction and coherence of response, adapts to changes in business needs, and ensures the timeliness of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387459A_ABST
    Figure CN120387459A_ABST
Patent Text Reader

Abstract

The invention discloses a multilayer dynamic aggregation big language model driven letter text analysis method, which relates to the technical field of natural language processing, and comprises the following steps of: acquiring multi-format letter data and converting the multi-format letter data into a text sequence with a timestamp, performing semantic analysis and structure processing on the text sequence, generating a text block set, and storing the text block set into a text database; performing multi-task training and topic fusion on the text block set to generate a comprehensive semantic vector, constructing a dynamic bridging graph based on the comprehensive semantic vector, activating associated text blocks by the bridging graph through semantic and time sequence relevance, responding to user query, splicing the activated text block content into a context Prompt, and inputting the context Prompt into a large language model; obtaining a one-time generated answer; the method has the advantages that semantic information and time sequence information are associated through the dynamic bridging graph, accurate response of the large language model is achieved based on the context Prompt, and the method has the advantages that complex semantic association is efficiently captured, context understanding coherence is improved, and dynamic information aggregation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation. Background Art

[0002] In the context of the rapid development of information technology, enterprises and organizations need to receive and process a large number of structured or unstructured correspondence texts in their daily operations, such as business emails, project reports, approval records, complaints and suggestions, etc. These correspondence contents have a wide range of sources and diverse formats, with complex semantic levels and loose structures, often containing information associations across multiple paragraphs, sentences, or even paragraphs, and changing rapidly with business needs.

[0003] Traditional text analysis methods mostly rely on statistical learning or rule-based natural language processing techniques, usually using keyword matching, topic models, or shallow sentiment analysis. These methods are applicable when dealing with single sentences or standard format texts, but when facing actual correspondences with complex discourse structures, strong context dependencies, non-standard language styles, and flexible expressions, it is often difficult to accurately capture deep semantic relationships, especially lacking the ability to model implicit associations between context, time sequence, and context, resulting in fragmented information extraction, incoherent responses, and poor scalability.

[0004] In recent years, pre-trained large language models such as BERT and GPT have made significant progress in natural language processing tasks, with powerful context understanding and generation capabilities. However, how to efficiently apply them to loose-structured and long-form correspondence texts still faces the following key challenges: 1) The text structure is unclear and the information density is uneven, and directly inputting into the large model will result in both "redundancy + missing".

[0005] 2) There are semantic bridging relationships between contexts, but lack of structured expressions.

[0006] 3) It is difficult to achieve one-time retrieval and centralized response based on queries.

[0007] Therefore, a method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation is proposed. Summary of the Invention

[0008] In view of the above-mentioned existing technical situation, this application is proposed. Embodiments of this application provide a method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation, which has the advantages of efficiently capturing complex semantic associations, improving the coherence of context understanding, and realizing dynamic information aggregation.

[0009] According to one aspect of the present application, there is provided a method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation, including: obtaining multi-format correspondence data and converting it into a text sequence with timestamps; performing semantic analysis and structural processing on the text sequence to generate a set of text blocks containing time series and context; performing multi-task training and theme fusion on the text blocks to generate comprehensive semantic vectors; constructing a dynamic bridging graph based on the comprehensive semantic vectors, and the bridging graph activates associated text blocks through semantic and time series correlations; in response to a user query, splicing the content of the activated text blocks into a context Prompt and inputting it into the large language model to obtain a one-time generated answer; receiving new correspondence data and updating the bridging graph according to the new correspondence data.

[0010] Compared with the prior art, by using a method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to an embodiment of the present application, semantic and time series information can be associated through a dynamic bridging graph, and accurate responses of the large language model can be achieved based on the context Prompt, which has the advantages of efficiently capturing complex semantic associations, improving the coherence of context understanding, and realizing dynamic information aggregation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0012] Figure 1 It is a flowchart of a method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to the present invention.

[0013] Figure 2 It is a flowchart for updating the bridging graph of a method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0015] Application Overview In traditional existing natural language processing systems, the heterogeneity of multi-format correspondence data leads to difficulties in information integration, and it is difficult to effectively model temporal correlation and context dependence. The semantic analysis of text sequences is limited to local feature extraction and cannot establish dynamic associations across text blocks, resulting in coexistence of input redundancy and semantic loss in large language models, and information fragmentation and logical breaks when generating responses. When the system processes long-span business email chains, it cannot automatically recognize the correlation between session identifiers and timestamps, resulting in a lack of continuity in cross-correspondence query responses.

[0016] For example, a bank customer service system needs to process complaint cases containing PDF contracts, scanned receipts, and email bodies. After OCR conversion, the scanned documents have incorrect date formats, and the nested references in the email bodies result in loose structures. Existing systems treat each email as an independent text block and cannot recognize the association between the discussion of fund flows and contract terms across emails. When a user queries "the reason for the delay in the approval process", the system only retrieves the "process adjustment notice" in the latest email and misses the scanned copy of the "system upgrade report" submitted by the technical department three weeks ago, and the generated response fails to associate the approval queue blockage caused by version compatibility.

[0017] If the above problems are not solved, the semantic gap of cross-modal data will lead to incomplete business decision-making basis and compliance risks caused by missing key information. The lack of temporal correlation makes historical correspondences unable to effectively participate in the current session reasoning, and factual contradictions appear in the content generated by large language models. Insufficient dynamic context maintenance ability will reduce the stability of the system in multi-round interactions, and users need to repeatedly correct query conditions to obtain effective information, significantly increasing the human-machine collaboration cost.

[0018] Facing the above problems, this application first considers how to convert multi-format correspondence data into a unified temporal structure to eliminate the interference of heterogeneity, explores a phased enhancement processing scheme for the semantic gap of text sequences, and designs a dynamic association mechanism to capture the implicit connections across text blocks. Directly inputting into large language models by traditional methods will lead to coexistence of redundancy and loss. This application attempts to construct a set of temporalized text blocks in the preprocessing stage, generate comprehensive vectors through hierarchical semantic fusion, and then establish a bridging graph structure to achieve association activation. To avoid fragmented query responses, this application adopts a context splicing strategy based on the bridging graph, inputs the activated content into the large model to generate coherent answers. For the association changes introduced by new data, a bridging graph incremental update mechanism is designed to maintain dynamic adaptability.

[0019] Exemplary Method Figure 1Illustrated is a method for analyzing correspondence text driven by a multi-layer dynamic aggregation large language model according to an embodiment of the present application, including: obtaining multi-format correspondence data and converting it into a text sequence with timestamps; performing semantic analysis and structural processing on the text sequence to generate a set of text blocks including time series and context; performing multi-task training and topic fusion on the text blocks to generate comprehensive semantic vectors; constructing a dynamic bridging graph based on the comprehensive semantic vectors, and the bridging graph activates associated text blocks through semantic and time series relevance; in response to a user query, splicing the content of the activated text blocks into a context Prompt and inputting it into the large language model to obtain an answer generated at one time; receiving new correspondence data and updating the bridging graph according to the new correspondence data.

[0020] Among them, the multi-format correspondence data includes plain text files, document files, image files, and scanned documents. Specifically, OCR technology can be used to convert the text of images and scanned documents, and all formats are encoded uniformly and cleaned, with timestamp information added to generate a unified text sequence. Its function is to solve the problem that traditional methods are difficult to process multi-source heterogeneous data and ensure that subsequent processing has time series relevance.

[0021] Among them, semantic analysis and structural processing refer to performing word segmentation, vectorization, sentence segmentation, and semantic enhancement on the unified text sequence to generate a set of text blocks including time series and context. Specifically, a pre-trained language model and dependency syntax analysis can be used to achieve this. Its function is to eliminate semantic fragmentation caused by loose structures and improve the context association modeling ability.

[0022] Among them, the comprehensive semantic vector refers to semantic aggregation of text blocks through multi-task training and topic fusion. Specifically, it can be generated by jointly optimizing the masked language model loss, adjacent block discrimination loss, and cross-block contrast loss. Its function is to integrate multi-level semantic features and solve the problems of redundancy and missing caused by uneven information density.

[0023] Among them, the dynamic bridging graph refers to constructing a semantic and time series relevance network based on the comprehensive semantic vectors. Specifically, node connections can be established by calculating cosine similarity or timestamp difference. Its function is to explicitly express the implicit associations across text blocks and provide structural support for context retrieval.

[0024] Among them, the context Prompt refers to splicing the text blocks activated by the user query in chronological order as the input to the large language model. Specifically, it can be achieved by using similarity matching and neighbor node activation strategies. Its function is to concentrate relevant contexts at one time and avoid the problem of incoherent responses in traditional methods.

[0025] Among them, updating the bridging graph refers to dynamically adjusting the semantic and time series association edges according to the new correspondence data. Specifically, new text blocks can be repeatedly generated and their similarity with existing nodes can be calculated. Its function is to adapt to changes in business requirements and ensure the continuous expansion ability of the model.

[0026] The core innovation of this application lies in explicitly modeling the semantic and temporal relevance of correspondence texts through a dynamic bridging graph, generating a contextually coherent Prompt input in combination with a multi-layer dynamic aggregation mechanism, and using a large language model to achieve a one-time centralized response, overcoming the defects of information fragmentation and incoherent responses in traditional methods under complex discourse structures.

[0027] The working process and principle of this application are as follows: First, multi-format correspondence data is obtained and converted into a text sequence with timestamps. This step unifies different formats of correspondence data into a processable text form and attaches time information to preserve the temporal relationship. Then, semantic analysis and structural processing are performed on the text sequence to generate a set of text blocks containing time series and context. This step identifies the semantic units of the text through semantic analysis and organizes them into text blocks considering the context relationship. Then, multi-task training and theme fusion are performed on the text blocks to generate a comprehensive semantic vector. This process extracts multi-dimensional features of the text blocks through multi-task learning and fuses theme information to generate a more comprehensive semantic representation. A dynamic bridging graph is constructed based on the comprehensive semantic vector. The bridging graph activates associated text blocks through semantic and temporal relevance. The bridging graph structure explicitly represents the semantic and temporal relationships between text blocks, facilitating subsequent associated activation. When responding to a user query, the content of the activated text blocks is concatenated into a context Prompt and input into the large language model to obtain a one-time generated answer. This step activates relevant text blocks through the bridging graph, organizes them into the input of the large language model, and thus generates a coherent answer. Finally, new correspondence data is received, and the bridging graph is updated according to the new correspondence data, which ensures that the system can dynamically adapt to the newly added correspondence information and maintain the timeliness of the bridging graph.

[0028] As a preferred embodiment, the solution of the present application is specifically implemented as follows: First, the system receives multi-format correspondence data including emails, PDF documents, and scanned images. For image data, OCR processing is performed to extract text content, and all text content is uniformly encoded in UTF-8 format and undergoes cleaning processing such as removing special characters and unifying date formats. Timestamp information is added to each piece of text to form a unified text sequence with timestamps. Then, a pre-trained language model such as DistilBERT is used to encode the text sequence to obtain word-level semantic vectors. Next, the text is segmented based on punctuation and semantic similarity to generate preliminary text blocks. After that, dependency syntactic analysis and sentiment analysis are performed on each text block, and syntactic and sentiment features are fused to generate enhanced semantic vectors. Then, the enhanced semantic vectors are input into the pre-trained model to obtain block-level semantic vectors. The block-level semantic vectors are subject to topic clustering through a clustering algorithm, keywords are extracted to generate a topic distribution vector, and the topic distribution vector is concatenated with the block-level semantic vectors to obtain a comprehensive semantic vector. A dynamic bridging graph is constructed based on the similarity and temporal information of the comprehensive semantic vectors, where nodes represent text blocks and edges represent semantic or temporal associations. When a user query is received, the query is encoded as a vector and matched with the comprehensive semantic vectors for similarity. The most relevant text block is selected as the core block, and the adjacent nodes of the core block are activated through the bridging graph structure. The content of the activated text blocks is concatenated in chronological order to form a context Prompt, and the Prompt is input into a large language model such as GPT-3 to generate a one-time coherent answer. When new correspondence data is input, the above processing flow is repeated to generate new text blocks and comprehensive semantic vectors, and the bridging graph structure is updated according to semantic similarity and temporal relationships.

[0029] Through the above solution, the present application realizes the unified processing and semantic analysis of multi-format correspondence data, overcoming the limitations of traditional methods in dealing with complex discourse structures and context dependencies. The dynamic bridging graph structure effectively captures the semantic and temporal associations between text blocks, improving the integrity and coherence of information extraction. The context Prompt generation strategy based on activated text blocks fully utilizes the understanding and generation capabilities of large language models, while avoiding problems such as information redundancy and missing caused by direct input. In addition, the dynamic update mechanism of the bridging graph ensures that the system can continuously adapt to newly added correspondence data and maintain the timeliness of analysis results. This method significantly improves the accuracy of information extraction and the coherence of responses when dealing with loose-structured and long-form correspondence texts, providing more efficient and reliable technical support for the daily correspondence processing of enterprises and organizations.

[0030] In some of the above solutions of the present application, when obtaining multi-format correspondence data and converting it into a unified text sequence, non-text format files have problems such as high recognition error rates, chaotic character encodings, and missing temporal information, making it difficult for subsequent processing links to accurately capture semantic associations.

[0031] The present application further proposes a method for obtaining multi-format correspondence data and converting it into a timestamped text sequence, including: obtaining multi-format correspondence data, where the multi-format correspondence data includes plain text files, document files, image files, and scanned documents; performing grayscale conversion, noise reduction, and text region detection on the image files and scanned documents, and generating editable text through OCR technology; performing encoding unification and cleaning on the plain text files, document files, and editable text, including correcting OCR misrecognized characters, unifying the date format, and attaching timestamp information, to generate a timestamped unified text sequence.

[0032] Among them, for image format processing, grayscale conversion and noise reduction technologies are used to reduce optical interference, morphological filtering is used to eliminate noise points, and a text region localization algorithm based on connected component analysis is used to extract effective text regions. In the OCR process, a deep learning model is used to achieve character-level recognition, and a character confidence threshold is set in the output layer to filter low-confidence results. Encoding unification uses the UTF-8 standard to convert texts in different encoding formats, and in the cleaning stage, the date format is matched and replaced through regular expressions, and at the same time, a confusion matrix is used to replace and correct error-prone characters. The timestamp information is generated by parsing the creation time field in the file metadata or extracting the time information according to the file naming rule.

[0033] Specifically, when processing a scanned document, first convert the RGB image to a grayscale image, use non-local means filtering to eliminate scanning noise, and then determine the boundary of the text region through an edge detection algorithm. The OCR engine performs line-by-line recognition on the delimited region, and marks the position coordinates and confidence values of each character when outputting the text. For characters with a confidence lower than the set threshold, dynamic correction is performed through the semantic relevance of adjacent characters. In the encoding conversion stage, the character encoding type of the original file is detected and uniformly converted to the UTF-8 format. The date is uniformly recognized in multiple formats such as "YYYY-MM-DD" and "YYYY / MM / DD" through regular expressions and uniformly converted to the form of "YYYYMMDD". The timestamp is attached to generate a time tag accurate to the second based on the modification time recorded by the file system. When the metadata is missing, the timestamp is completed through the date information embedded in the file name. The resulting text sequence not only retains the original semantic information, but also establishes cross-document time relevance, providing a data basis for subsequent time series analysis.

[0034] In some of the above solutions of the present application, directly performing multi-task training or bridge graph construction on the original text sequence will lead to problems such as insufficient semantic understanding, loss of time series and context.

[0035] The present application further proposes to perform word segmentation on a text sequence to generate a sub-word sequence, input the sub-word sequence into a pre-trained language model to generate sub-word vectors, perform sentence segmentation on the text sequence based on punctuation marks to obtain a number of candidate sentences, perform semantic enhancement processing on the candidate sentences to generate enhanced semantic vectors, input the enhanced semantic vectors into the pre-trained language model to generate sentence-level semantic vectors, and merge or split the candidate sentences according to the semantic similarity threshold and length rule of the sentence-level semantic vectors to obtain a text block set.

[0036] As a preferred embodiment, the solution of the present application is specifically implemented as follows: Perform word segmentation on the text sequence to generate a sub-word sequence. Specifically, input the text sequence into the tokenizer of DistilBERT and set the features of the tokenizer to tokenize the text.

[0037] Input the sub-word sequence into a pre-trained language model to generate sub-word vectors. Among them, the pre-trained language model uses the DistilBERT model.

[0038] Perform sentence segmentation on the text sequence based on punctuation marks to obtain a number of candidate sentences. Further, use regular expressions to match punctuation marks such as full stops, question marks, and exclamation marks as the segmentation basis.

[0039] Perform semantic enhancement processing on the candidate sentences to generate enhanced semantic vectors. Specifically, first call Stanza for dependency parsing to obtain a dependency type vector; then concatenate the dependency type vector with the sub-word vector and input it into a two-layer fully connected neural network to obtain a syntactic enhancement vector; then use a lightweight BiLSTM model to perform sentiment analysis on the candidate sentences to obtain a sentiment distribution vector; finally, concatenate the syntactic enhancement vector with the sentiment distribution vector to obtain an enhanced semantic vector.

[0040] Input the enhanced semantic vector into a pre-trained language model to generate a sentence-level semantic vector. Thus, use the DistilBERT model to encode the enhanced semantic vector to obtain a sentence-level semantic vector.

[0041] Merge or split the candidate sentences according to the semantic similarity threshold and length rule of the sentence-level semantic vectors to obtain a text block set. For example, set the semantic similarity threshold to 0.8, the maximum length threshold to 512 characters, and the minimum length threshold to 50 characters. Calculate the cosine similarity for adjacent candidate sentences, and merge them if it exceeds 0.8; split them if the length exceeds 512 after merging; and merge them with adjacent blocks if the length is less than 50 after splitting.

[0042] Through the above technical solutions, the present application realizes the automated semantic analysis and structured processing of complex correspondence texts. Through multi-level semantic representations and flexible text block division rules, the semantic associations between sentences are effectively captured while maintaining an appropriate granularity. This method can better retain context information, providing structured input for subsequent multi-task learning and semantic retrieval, thereby improving the accuracy and efficiency of correspondence text analysis.

[0043] In some of the above solutions of the present application, it is proposed to directly merge or split candidate sentences obtained by segmenting a text sequence. However, candidate sentences may have low subsequent processing efficiency due to excessive character counts or semantic incoherence, and poor syntactic independence may lead to inaccurate sentence segmentation, affecting the quality of the text block set.

[0044] The present application further proposes to perform character-level verification before merging or splitting candidate sentences. If the character count exceeds a preset character threshold or the semantics are incoherent, it is further segmented into sub-sentences, and the syntactic independence of the sub-sentences is verified through dependency syntactic analysis. If the preset independence conditions are met, the sentence segmentation position is adjusted.

[0045] Among them, character-level verification is performed by counting the character count of the candidate sentence and comparing it with the preset threshold. If it exceeds, a segmentation operation is triggered; semantic coherence is judged by detecting whether there is a lack of logical conjunctions or a sudden change in the theme in the sentence; dependency syntactic analysis extracts the subject-predicate-object structure and modification relationships of the sub-sentence. If the sub-sentence contains complete syntactic components, it is determined to be independent; adjusting the segmentation position is based on redefining the start and end points of the sentence based on the syntactic component boundaries.

[0046] Specifically, character-level verification uses a sliding window algorithm to scan the candidate sentence character by character. When the cumulative character count exceeds the threshold, a temporary segmentation point is generated, and it is detected whether there is a semantic break before and after the segmentation point. If there is, the segmentation point is moved backward to the nearest punctuation mark or conjunction position; dependency syntactic analysis determines the syntactic independence of the sub-sentence by parsing the dependency tree structure of the sub-sentence. If the sub-sentence contains a core predicate and the subject-predicate relationship is complete, it is determined to have syntactic independence; when adjusting the segmentation position, if the sub-sentence meets the independence conditions, the segmentation point of the original candidate sentence is corrected to the start or end position of the sub-sentence. For example, a candidate sentence has 250 characters, exceeding the preset threshold of 200 characters. Through sliding window detection, a comma is found at the 180th character and the semantics are coherent before and after, so it is segmented into two sub-sentences; further dependency analysis reveals that the second sub-sentence lacks a subject, so the segmentation point is adjusted to the full stop position at the 150th character to ensure the syntactic integrity of the sub-sentence. The resulting text block set has higher semantic coherence and structural normativity, providing reliable input for subsequent multi-task training.

[0047] In some of the above - mentioned solutions of the present application, a solution is proposed to merge or split candidate sentences to obtain a set of text blocks. However, before merging or splitting, the candidate sentences may have a situation where the number of characters is too large or the semantics is incoherent, resulting in the length of the merged text block exceeding the processing capacity or the syntax of the split sub - clauses being not independent, affecting the accuracy of subsequent semantic vector generation and associated text blocks.

[0048] The present application further proposes to perform character - level verification on each candidate sentence before merging or splitting. If the number of characters exceeds the preset character threshold or the semantics is incoherent, it is further segmented into sub - clauses, and the syntactic independence of the sub - clauses is verified through dependency parsing. If the preset independence condition is met, the sentence segmentation position is adjusted.

[0049] As a preferred embodiment, the solution of the present application is specifically implemented as follows: Perform character - level verification on each candidate sentence. If the number of characters exceeds the preset character threshold or the semantics is incoherent, it is further segmented into sub - clauses. Specifically, the preset character threshold is set to 100 characters. For candidate sentences with more than 100 characters, the sliding window method is used for segmentation. The sliding window size is set to 50 characters, and the step size is 25 characters. Calculate the semantic coherence score for each text segment within the window, and the position with a score lower than the threshold is used as the segmentation point.

[0050] Furthermore, verify the syntactic independence of the sub - clauses through dependency parsing. If the preset independence condition is met, adjust the sentence segmentation position. Among them, Stanza is used for dependency parsing. Generate a dependency tree for each sub - clause and calculate the integrity score of the subtree. If the score is higher than the preset threshold, it is considered that the sub - clause has syntactic independence. For sub - clauses that do not meet the independence condition, merge them with adjacent sub - clauses or further segment them until all sub - clauses meet the syntactic independence requirements.

[0051] Through the above - mentioned technical solutions, the present application can effectively process long sentences and complex sentence patterns, improve the semantic integrity and independence of text blocks. Thereby improving the accuracy of subsequent semantic analysis and information extraction, and laying a foundation for constructing a high - quality set of text blocks. Furthermore, through syntactic independence verification, the semantic fragmentation caused by mechanical segmentation is avoided, ensuring the syntactic structure integrity of text blocks. This refined processing based on linguistic knowledge significantly improves the robustness and interpretability of text analysis.

[0052] In some of the above - mentioned solutions of the present application, when merging or splitting candidate sentences, due to only relying on the similarity of sentence - level semantic vectors and character - length rules, without fully combining the syntactic structure and sentiment - tendency information of sentences, there may be problems such as syntactic - logical confusion or inconsistent sentiment expression in the merged text blocks, thus affecting the accuracy of subsequent semantic clustering.

[0053] The present application further proposes a specific method for semantic enhancement processing of candidate sentences.

[0054] Among them, semantic enhancement processing generates a dependency type vector through dependency syntactic analysis, concatenates it with the sub-word vector, and inputs the result into a preset neural network to generate a syntactic enhancement vector; at the same time, a preset sentiment classification model is used to generate a sentiment distribution vector, and the syntactic enhancement vector and the sentiment distribution vector are concatenated to obtain an enhanced semantic vector. Specifically, dependency syntactic analysis identifies the grammatical relationships of each component in a sentence and outputs a dependency type vector to represent the syntactic structure; the sub-word vector comes from the tokenization result of a pre-trained language model; after the two are concatenated, the neural network is used to further extract syntactic and semantic features. The sentiment classification model predicts the probability distribution of the candidate sentence based on a preset sentiment label system and generates a sentiment distribution vector. The two vectors respectively reflect grammatical and sentiment information, and multi-dimensional semantic features are integrated through the concatenation operation.

[0055] As a preferred embodiment, the solution of the present application is specifically implemented as follows: Semantic enhancement processing includes the following steps: First, perform dependency syntactic analysis on the candidate sentence to generate a dependency type vector. Specifically, use Stanza to perform dependency syntactic analysis on the candidate sentence to obtain the dependency relationships between words and encode these dependency relationships as vector representations.

[0056] Second, concatenate the dependency type vector and the sub-word vector and input them into a preset neural network to generate a syntactic enhancement vector. Further, a multi-layer perceptron is used as the preset neural network, the dimension of the input layer is the sum of the dimensions of the dependency type vector and the sub-word vector, the hidden layer uses the ReLU activation function, and the dimension of the output layer is the same as that of the input layer.

[0057] Third, perform sentiment analysis on the candidate sentence through a preset sentiment classification model to generate a sentiment distribution vector. Thus, a lightweight BiLSTM model is used to input the candidate sentence into the model to obtain the probability distributions of three sentiments: positive, neutral, and negative.

[0058] Finally, concatenate the syntactic enhancement vector and the sentiment distribution vector to obtain an enhanced semantic vector. Specifically, the above two vectors are concatenated in the last dimension to form the final enhanced semantic vector.

[0059] Through the above technical solutions, the present application realizes the multi-dimensional enhancement of the semantics of candidate sentences. By capturing the structural information of sentences through dependency syntactic analysis and combining the semantic information provided by subword vectors, a richer syntactic representation is generated. At the same time, the results of sentiment analysis are introduced to further supplement the sentiment semantic information of sentences. This multi-angle semantic enhancement method can more comprehensively characterize the semantic features of candidate sentences, providing richer and more accurate inputs for subsequent text block generation and semantic vector calculation, thereby improving the performance and accuracy of the entire correspondence text analysis system.

[0060] In some of the above solutions of the present application, when merging candidate sentences based on semantic similarity, only semantic relevance is considered, and no constraint is imposed on the length of the merged text block, which may result in the generated text block being too long or too short, thereby affecting the efficiency and accuracy of subsequent semantic vector aggregation.

[0061] The present application further proposes merging or splitting, including: calculating the cosine similarity of the sentence-level semantic vectors of candidate sentences, and if the similarity exceeds a preset first similarity threshold, merging them into the same text block; if the number of characters in the merged text block exceeds a preset maximum length threshold, splitting it into sub-blocks; if the number of characters in the split sub-block is less than a preset minimum length threshold, merging it with the adjacent block.

[0062] Among them, the merging operation is triggered based on the semantic similarity threshold to ensure the semantic consistency of related sentences; when splitting sub-blocks, the maximum length threshold is used to avoid information redundancy, and the minimum length threshold is used to prevent information fragmentation; merging adjacent blocks further optimizes the sub-block length and maintains context coherence. The preset maximum length threshold can be set to 500 characters, the minimum length threshold can be set to 100 characters, and the first similarity threshold can be set to 0.85.

[0063] Specifically, after calculating the cosine similarity of candidate sentences, the sentences exceeding the threshold are merged to form an initial text block. When the number of characters in the merged text exceeds 500, it is split into multiple sub-blocks according to sentence boundaries. If the number of characters in the split sub-block is less than 100, it is merged with the adjacent sub-block to ensure that the sub-block length is within a reasonable range. For example, if the number of characters in the merged text block is 600 and it needs to be split into two sub-blocks, and if the number of characters in one of the sub-blocks is 80, it is merged with the adjacent sub-block to form a final sub-block with 180 characters. This process dynamically adjusts the text block length, not only retaining semantic relevance but also avoiding the reduction of subsequent processing efficiency or the loss of semantic information caused by too long or too short text blocks.

[0064] As a preferred embodiment, the solution of the present application is specifically implemented as follows: Calculate the cosine similarity of the sentence-level semantic vectors of candidate sentences. If the similarity exceeds a preset first similarity threshold, they are merged into the same text block. Specifically, the first similarity threshold can be set to 0.85. For any two candidate sentences A and B, calculate the cosine similarity of their sentence-level semantic vectors. If the similarity is greater than 0.85, then merge A and B into the same text block.

[0065] If the number of characters in the text block exceeds the preset maximum length threshold after merging, it is split into sub-blocks. Further, the maximum length threshold can be set to 500 characters. When the merged text block exceeds 500 characters, it is truncated at the 500th character, and the excess part is used as a new sub-block.

[0066] If the number of characters in the sub-block is less than the preset minimum length threshold after splitting, it is merged with the adjacent block. Among them, the minimum length threshold can be set to 100 characters. When the number of characters in the split sub-block is less than 100, the sub-block is merged with the adjacent text block. This can avoid generating overly short text blocks.

[0067] Through the above technical solutions, the present application realizes the intelligent merging and splitting of candidate sentences. By controlling the merging granularity through the cosine similarity threshold, it avoids the incorrect merging of semantically unrelated sentences. At the same time, by controlling the text block size through the maximum and minimum length thresholds, it not only prevents the decrease in processing efficiency caused by an overly long single text block but also avoids the semantic fragmentation caused by overly short text blocks. This flexible text block division method improves the accuracy and efficiency of subsequent semantic analysis and information retrieval.

[0068] In some of the above solutions of the present application, although the enhanced semantic vectors of the words within the text block can reflect local semantic information, they lack effective aggregation of the global relationships among multiple words within the block, resulting in fragmented block-level semantic expressions; at the same time, the existing methods do not integrate the topic distribution with the block-level semantics, making it difficult to accurately capture the topic relevance across text blocks during the subsequent dynamic bridge graph construction process, affecting the integrity and relevance expression of the comprehensive semantic vector.

[0069] The present application further proposes to aggregate the enhanced semantic vectors of all words within the text block to generate a block-level input vector; input the block-level input vector into a preset block encoder, and through the joint optimization of the masked language model loss, adjacent block discrimination loss, and cross-block contrast loss, generate a block-level semantic vector; perform topic clustering on the block-level semantic vector based on a clustering algorithm, and generate a topic distribution vector through a keyword extraction method; splice the topic distribution vector and the block-level semantic vector to generate a comprehensive semantic vector.

[0070] Among them, the block-level input vector aggregates the enhanced semantic vectors of each word in the text block through average pooling or maximum pooling to form an initial representation reflecting the global semantics within the block; the block encoder adopts a multi-layer Transformer structure, and the masked language model loss optimizes the model's robustness to local semantic loss by randomly masking some dimensions of the input vector and predicting the original value. The neighboring block discrimination loss distinguishes adjacent blocks from non-adjacent blocks through contrastive learning, strengthening the temporal context relevance. The cross-block contrast loss improves semantic discrimination by shortening the distance between semantically similar blocks and pushing away irrelevant blocks; topic clustering adopts K-means or hierarchical clustering algorithm to perform unsupervised grouping of block-level semantic vectors, extract high-frequency words or key phrases from each cluster to generate a topic distribution vector, and the topic distribution vector calculates the weight of each topic by weighting the word frequency-inverse document frequency to form a multi-dimensional probability distribution; the splicing operation uses vector concatenation or weighted summation to fuse the topic distribution vector and the block-level semantic vector to generate a comprehensive semantic vector containing topic and semantic information.

[0071] Specifically, the enhanced semantic vectors of each word in a text block are aggregated into a block-level input vector through a pooling operation and then input into a block encoder for multi-task joint training. The masked language model loss randomly masks the values of some dimensions in the input vector, forcing the model to restore the masked content based on contextual information, thereby enhancing the model's adaptability to local semantic loss. The neighboring block discriminant loss marks text blocks with adjacent timestamps as positive samples and non-adjacent blocks as negative samples, optimizing the block encoder's ability to capture temporal proximity relationships through contrastive learning. The cross-block contrastive loss calculates the cosine similarity between block-level semantic vectors, treating block pairs with similarity above a preset threshold as positive samples and those below the threshold as negative samples. The contrastive loss function adjusts the vector space distribution to improve the aggregation of semantically similar blocks. After multi-task joint optimization, the block-level semantic vector simultaneously contains local semantic robustness, temporal correlation, and semantic discrimination features. On this basis, a clustering algorithm is used to thematically group block-level semantic vectors. For example, K-means is used to divide semantic vectors into a preset number of clusters. The top N words with the highest TF-IDF values in each cluster are extracted as topic keywords to form a distribution vector reflecting the core topic of each cluster. Finally, the topic distribution vector and the block-level semantic vector are concatenated by dimension, so that the comprehensive semantic vector simultaneously encodes the semantic features of the text block and the probability distribution of its topic. For example, when the block-level semantic vector dimension is 768 and the topic distribution vector dimension is 50, the comprehensive semantic vector is expanded to 818 dimensions, of which the first 768 dimensions carry semantic information and the last 50 dimensions record topic weights. This approach enables the dynamic bridge graph construction process to not only match related blocks by semantic similarity but also strengthen cross-block connections based on topic consistency, improving the accuracy and completeness of context splicing in subsequent query responses.

[0072] As a preferred embodiment, the solution of the present application is specifically implemented as follows: Aggregate the enhanced semantic vectors of the words within the text block to generate a block-level input vector. Specifically, use the attention mechanism to perform weighted summation on the word vectors to obtain a block-level representation of a fixed dimension.

[0073] Input the block-level input vector into a preset block encoder, and generate a block-level semantic vector by jointly optimizing the masked language model loss, adjacent block discrimination loss, and cross-block contrast loss. Among them, the block encoder adopts a Transformer structure. The masked language model task predicts the randomly masked words, the adjacent block discrimination task distinguishes adjacent and non-adjacent text blocks, and the cross-block contrast task maximizes the mutual information of different block representations within the same document.

[0074] Perform topic clustering on the block-level semantic vectors based on a clustering algorithm, and generate a topic distribution vector through a keyword extraction method. Further, use the K-means algorithm for clustering, use the TF-IDF method to extract keywords for each category, and generate a topic-word distribution matrix.

[0075] Concatenate the topic distribution vector and the block-level semantic vector to generate a comprehensive semantic vector. Thus, the comprehensive semantic vector contains both block-level semantic information and topic distribution information.

[0076] Through the above technical solutions, the present application realizes multi-task joint training and topic fusion of text blocks, effectively captures the word relationships within the blocks, the context associations between blocks, and the global topic distribution. The comprehensive semantic vector integrates multi-level semantic information, improving the accuracy of subsequent text analysis and retrieval. At the same time, multi-task training enhances the generalization ability of the model, enabling it to adapt to different types and styles of correspondence texts. Topic fusion helps to understand the text content from a macroscopic perspective and provides support for subsequent text organization and summary generation.

[0077] In some of the above solutions of the present application, a dynamic bridging graph needs to be constructed after generating the comprehensive semantic vector. However, the existing methods do not clarify the specific activation rules for semantic association and temporal association, resulting in the bridging graph being unable to effectively reflect the multi-dimensional dynamic relationships between text blocks, affecting the accuracy of subsequent query responses.

[0078] The present application further proposes that constructing a dynamic bridging graph includes: if the cosine similarity of the comprehensive semantic vectors of two text blocks exceeds a preset second similarity threshold, then add a semantic association edge; if the time stamp difference between two text blocks is less than a preset number of days or they share a session identifier, then add a temporal association edge.

[0079] Among them, the generation of semantic association edges is based on the quantitative calculation of the semantic similarity between text blocks. The second similarity threshold is set in the range of 0.85 - 0.95, and the quantitative determination of cross-text block semantic association is realized by calculating the cosine value of the comprehensive semantic vector. The generation of temporal association edges adopts a dual-path conditional trigger mechanism. For path one, the timestamp difference threshold is set to 3 - 7 natural days, and for path two, the consistency of session identifiers is detected. The two conditions have a logical OR relationship. The semantic association edges and the temporal association edges are stored with different weight coefficients. Among them, the weight of the semantic edge is the cosine similarity value, and the weight of the temporal edge is the reciprocal of the time difference.

[0080] Specifically, during the construction of the bridging graph, the system traverses the comprehensive semantic vectors of all text blocks and calculates the cosine similarity of each pair of vectors in a sliding window manner. When it is detected that the similarity of a certain pair exceeds the second preset threshold, a weighted semantic edge is automatically generated and stored in the graph database. At the same time, the system reads the timestamp metadata of each text block. If the time difference between two blocks is less than the preset number of days or they belong to the same session chain, a temporal edge is established in the graph structure. The superposition of the semantic edge and the temporal edge forms a hybrid association network, where the semantic edge supports cross-period topic association mining, and the temporal edge ensures session continuity. When new correspondence data is input, the system only needs to incrementally calculate the association conditions between the new block and the existing blocks and dynamically expand the bridging graph structure. Through this dual association mechanism, the bridging graph can synchronously capture the semantic relevance and temporal continuity of text blocks and provide multi-dimensional association path support for subsequent query activation.

[0081] As a preferred embodiment, the solution of this application is specifically implemented as follows: When constructing a dynamic bridging graph, first calculate the cosine similarity of the comprehensive semantic vectors of text blocks. When the vector cosine similarity of two text blocks reaches 0.85, a semantic association edge is established in the bridging graph. For text blocks with a timestamp difference of less than 5 days or text blocks with the same session identifier, a temporal association edge is further established. The semantic association edges are stored through a bidirectional graph structure, and the temporal association edges are recorded in the form of a singly linked list in chronological order. Specifically, during the process of establishing the association edges, the semantic similarity calculation adopts the dot product operation after vector normalization, the timestamp difference calculation takes the absolute value in days, and the session identifier matching adopts the method of exact string comparison.

[0082] Through the above technical solutions, this application effectively solves the problem that it is difficult to synchronously model semantic associations and temporal continuity in correspondence text analysis. Through the dual association mechanism of the dynamic bridging graph, both the semantic similarity features between text blocks are retained, and the temporal logic relationships in the business scenario are captured. This structure enables the large language model to accurately activate context fragments with close semantic associations and temporal correlations when processing user queries, avoiding the problem of response fragmentation caused by the isolated analysis of text blocks in traditional methods. At the same time, the dynamic update mechanism ensures the adaptive ability of the bridging graph to new correspondence data, ensuring the timeliness of the association relationship.

[0083] In some of the above solutions of this application, when constructing the dynamic bridging graph, only the similarity of the comprehensive semantic vectors is relied on to establish the association edges, and the continuous relationships in the time dimension or session context cannot be effectively identified, resulting in the inability to form effective associations between text blocks of the same topic across time periods or sessions, affecting the integrity of context splicing during subsequent query responses.

[0084] This application further proposes to add semantic association edges when the cosine similarity of the comprehensive semantic vectors of two text blocks exceeds a preset second similarity threshold, and at the same time add temporal association edges when the timestamp difference between two text blocks is less than a preset number of days or they share a session identifier.

[0085] Among them, the establishment condition of the semantic association edge is set as a cosine similarity threshold, which is determined through multiple experimental verifications. For example, it is set to 0.85; the determination condition of the temporal association edge includes two independent logical branches, the preset number of days is dynamically adjusted according to the business scenario, and the typical value is 7 days, and the shared session identifier is obtained by analyzing the email header information or document metadata. The semantic association edge and the temporal association edge adopt different weight coefficients and are visually distinguished by solid lines and dashed lines respectively in the bridging graph, and the two together constitute a composite association network.

[0086] Specifically, during the process of generating the dynamic bridging graph, first traverse the comprehensive semantic vector matrix of all text blocks, and batch calculate the cosine similarity through matrix operations. When text block pairs with similarity exceeding the threshold are detected, semantic edges with weights are automatically established in the graph, and the weight values are linearly mapped to the similarity value range. At the same time, traverse the timestamp information of all text blocks, establish temporal edges for adjacent text blocks with a time interval within the set number of days, and force the establishment of temporal edges for text blocks with the same session identifier. For example, when processing project progress emails for three consecutive days, if the topic similarity of two emails reaches 0.87 and the time interval is two days, both semantic edges and temporal edges are established. When the user's query involves the project progress timeline, the bridging graph activates historical emails and the latest email through the dual association mechanism, ensuring that the context received by the large language model contains complete time series information and semantic association content.

[0087] Through the above technical solutions, the present application effectively solves the problem of broken context association in the processing of long texts by large language models. By means of a dynamic activation mechanism, it accurately captures semantic associations across documents, ensuring that the generated answers cover both the core query content and maintain the complete logical chain of business events. While maintaining the generation ability of the large model, this solution avoids the risk of key information being overwhelmed caused by directly inputting a large amount of text, and improves the system response accuracy to a level that can support actual business decisions.

[0088] In some of the above solutions of the present application, the existing bridging graph cannot be dynamically updated according to new correspondence data, resulting in the need to regenerate the entire bridging graph structure when processing new correspondences, causing waste of computing resources and response delays.

[0089] The present application further proposes to repeat the foregoing process according to new correspondence data to generate new text blocks and new comprehensive semantic vectors; calculate the similarity between the new comprehensive semantic vectors and the existing comprehensive semantic vectors; add semantic association edges if the similarity exceeds a preset second similarity threshold; and add temporal edges if the timestamp of the new text block and the existing text block is less than a preset number of days or they share a session identifier.

[0090] Among them, the same encoder and clustering algorithm as those used in the original data processing are adopted when generating new text blocks to ensure the consistency of the vector space. The cosine similarity algorithm is used when calculating the similarity, and the preset second similarity threshold is set as a value in the range of 0.85 - 0.95. The preset number of days is set to 3 days or 7 days when adding temporal edges, and the specific value is adjusted according to the business scenario. The shared session identifier is implemented by parsing the session ID field in the correspondence header metadata.

[0091] Specifically, when a new correspondence is received, the system generates a new comprehensive semantic vector through OCR conversion, semantic analysis, and topic fusion. This vector is batch-compared with the existing vectors in the bridging graph, and if the detected similarity is higher than the threshold, semantic association edges are automatically established. At the same time, the system extracts the timestamp and session identifier of the new text block and performs temporal matching with the nodes in the bridging graph: text blocks with a time difference within the preset number of days establish temporal association edges, and text blocks sharing the same session ID establish strong temporal associations. This process adopts an incremental update mechanism, only performing local calculations on the new data and associated nodes, avoiding full graph reconstruction. Through the dual association of semantic edges and temporal edges, the new text blocks are accurately embedded in the original graph structure, ensuring that the context relationship formed by the latest data can be invoked during subsequent query responses.

[0092] As a preferred embodiment, the solution of the present application is specifically implemented as follows: when the system receives a new PDF-format letter containing a customer complaint email, it is converted into a timestamped text sequence through OCR recognition technology. After semantic analysis and structural processing of this text sequence, text blocks containing the complaint date and product model description are generated, and a comprehensive semantic vector containing the product failure characteristics is generated. The similarity between this comprehensive semantic vector and the text block vectors in the historical work order records is calculated through the cosine similarity algorithm, and the similarity value with the text block of the repair record of the same type of product three months ago reaches 0.87. According to the preset similarity threshold of 0.85, a semantic association edge is established. At the same time, it is detected that the time difference between the new text block and the historical text block is 92 days, exceeding the preset 90-day association period, so a temporal association edge is not established. However, it is detected that they share the session identifier of the same customer number, and based on this, a temporal association edge is established to complete the update of the bridging graph.

[0093] Through the above technical solution, the present application realizes the efficient integration ability of the letter analysis system for new data, automatically identifies and establishes semantic associations and business session continuity between text blocks through a dynamic update mechanism, solves the problem of associated breaks existing in traditional methods when processing incremental data, ensures the collaborative analysis ability of historical data and new data, and thus improves the context integrity when responding to user queries and the effectiveness of business decision support.

[0094] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation, characterized in that, including: Obtain multi-format correspondence data and convert it into a timestamped text sequence; Perform semantic analysis and structural processing on the text sequence to generate a set of text blocks containing time series and context; Perform multi-task training and topic fusion on the set of text blocks to generate a comprehensive semantic vector; Construct a dynamic bridging graph based on the comprehensive semantic vector, and the bridging graph activates associated text blocks through semantic and time series relevance; In response to a user query, splice the content of the activated text blocks into a context Prompt and input it into a large language model to obtain a one-time generated answer; Receive new correspondence data and update the bridging graph according to the new correspondence data.

2. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 1, characterized in that, The obtaining multi-format correspondence data and converting it into a timestamped text sequence includes: Obtain multi-format correspondence data, and the multi-format correspondence data includes plain text files, document files, image files, and scans; Perform grayscale conversion, noise reduction, and text region detection on the image files and scans, and generate editable text through OCR technology; Perform encoding unification and cleaning on plain text files, document files, and the editable text, including correcting OCR misrecognized characters, unifying date formats, and adding timestamp information, to generate a timestamped unified text sequence.

3. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 1, characterized in that The performing semantic analysis and structural processing on the text sequence to generate a set of text blocks containing time series and context includes: Perform word segmentation on the text sequence to generate a sub-word sequence; Input the sub-word sequence into a pre-trained language model to generate sub-word vectors; Perform sentence segmentation on the text sequence based on punctuation marks to obtain a number of candidate sentences; Perform semantic enhancement processing on the candidate sentences to generate enhanced semantic vectors; Input the enhanced semantic vectors into a pre-trained language model to generate sentence-level semantic vectors; Merge or split the candidate sentences according to the semantic similarity threshold and length rule of the sentence-level semantic vectors to obtain a set of text blocks.

4. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 3, characterized in that, Before the merging or splitting of the candidate sentences, it further includes: Perform character-level verification on each candidate sentence. If the number of characters exceeds a preset character threshold or the semantics is incoherent, further split it into sub-sentences; Verify the syntactic independence of the sub-sentences through dependency syntactic analysis. If the preset independence conditions are met, adjust the sentence segmentation position.

5. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 4, characterized in that, The semantic enhancement processing includes: Perform dependency syntactic analysis on the candidate sentences to generate dependency type vectors; Concatenate the dependency type vectors and sub-word vectors and input them into a preset neural network to generate syntactic enhancement vectors; Perform sentiment analysis on the candidate sentences through a preset sentiment classification model to generate sentiment distribution vectors; Concatenate the syntactic enhancement vectors and sentiment distribution vectors to obtain enhanced semantic vectors.

6. The method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 5, wherein The merging or splitting includes: Calculate the cosine similarity of the sentence-level semantic vectors of the candidate sentences. If the similarity exceeds a preset first similarity threshold, merge them into the same text block; If the number of characters in the merged text block exceeds a preset maximum length threshold, split it into sub-blocks; If the number of characters in the split sub-block is less than a preset minimum length threshold, merge it with the adjacent block.

7. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 6, characterized in that The multi-task training and topic fusion includes: Aggregate the enhanced semantic vectors of all words in the text block to generate block-level input vectors; Input the block-level input vector into a preset block encoder, and generate a block-level semantic vector through the joint optimization of masked language model loss, neighboring block discrimination loss, and cross-block contrast loss; Perform topic clustering on the block-level semantic vectors based on a clustering algorithm, and generate a topic distribution vector through a keyword extraction method; Concatenate the topic distribution vector and the block-level semantic vector to generate a comprehensive semantic vector.

8. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 1, characterized in that, The construction of the dynamic bridging graph includes: If the cosine similarity of the comprehensive semantic vectors of two text blocks exceeds a preset second similarity threshold, add a semantic association edge; If the time stamp difference between two text blocks is less than a preset number of days or they share a session identifier, add a temporal association edge.

9. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 1, characterized in that The response to the user query, concatenating the content of the activated text blocks into a context Prompt and inputting it into a large language model, and obtaining the one-time generated answer includes: Encode the user query into a query vector; Perform cosine similarity matching between the query vector and the comprehensive semantic vector, and select the top several text blocks with the highest cosine similarity as the core associated blocks; Mark the core associated blocks as activated, and activate their neighbor node text blocks through the semantic association edges and temporal association edges in the bridging graph; Concatenate the content of all activated text blocks in chronological order to generate a context Prompt; Input the context Prompt into a large language model to obtain a one-time generated answer.

10. A method for analyzing correspondence text driven by a large language model with multi-layer dynamic aggregation according to claim 1, characterized in that, The update of the bridging graph includes: Repeat the foregoing process according to the new letter data to generate new text blocks and new comprehensive semantic vectors; Calculate the similarity between the new comprehensive semantic vector and the existing comprehensive semantic vectors; If the similarity exceeds the preset second similarity threshold, add a semantic association edge; If the time stamp of the new text block and the existing text block is less than a preset number of days or they share a session identifier, add a temporal edge.

Citation Information

Patent Citations

  • Multimedia opinion oriented intelligence tool

    CA2326110A1

  • Public opinion text sentiment analysis method based on semantic dependency relationship fusion features

    CN115098634A

  • View angle level text sentiment classification system based on double graph convolutional neural network

    CN115858788A

  • Tibetan language pre-training language model training method, system and device and medium

    CN119026608A

  • Industrial part surface defect detection method and device

    CN119379583A

Cited By

  • Workflow processing method, system and equipment based on routing algorithm and medium

    CN120746254A

  • Cloud platform auditing method based on collaborative auditing

    CN121073058A

  • Multi-user mixed interactive data processing method and system based on large model

    CN121118904A

  • Large model-based multi-user mixed interaction data processing method and system

    CN121118904B

  • Text information compression method and compression device for index distributed database

    CN121144271A