Long text semantic classification method based on deep learning
The deep learning-based method for long text semantic classification addresses semantic discontinuity and class imbalance by structuring text, segmenting semantics, and using a syntax-driven aggregation and flow-enhanced classifier to improve accuracy and robustness.
Patent Information
- Application Number
- CN202510764050.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional text classification methods are prone to semantic incoherence and loss of context information when processing long texts, especially when there are structural semantic mutations such as significant topic transformation, turning relationship or conclusion closing, and cannot accurately capture the semantic changes brought about by these mutations; in response to the problem of category imbalance in long texts, traditional classifiers tend to overfit most class samples and ignore the recognition of a few classes.
A deep learning-based method is adopted to generate segment-level semantic vector sequences through structured cleaning and sentence boundary recognition, and a semantic tension-driven aggregation algorithm is used to generate global semantic representation vectors, and a semantic flow enhancement classifier is designed, combining semantic flow network and loss function optimization to improve classification accuracy and robustness.
Effectively capture semantic features in long texts, ensure the context continuity and semantic consistency of texts, and improve the accuracy and adaptability of long text classification, especially in the case of category imbalance and confrontational samples, showing stronger adaptability and anti-interference ability.
Smart Images

Figure CN120316255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and natural language processing, and particularly to a long text semantic classification method based on deep learning. Background Art
[0002] Long text semantic classification is a challenging task in the field of natural language processing. With the explosive growth of information, long texts not only contain a large amount of information but also have a high degree of complexity in semantic expression. Especially in texts, there are often structural semantic mutations such as topic transitions, turning relationships between paragraphs, and conclusion endings, which make text understanding more difficult. Long texts usually involve multiple sub-topics, and the semantic connections between each sub-topic and the dependencies between them need to be effectively modeled to correctly understand the overall semantics.
[0003] Existing text classification methods mostly focus on short texts. For long texts, especially those with complex semantic structures, traditional models often cannot fully capture the deep semantics therein. Many methods cannot handle the complex turns and structural changes in the text, resulting in poor performance in long text classification. Therefore, how to efficiently extract semantic features from long texts, especially accurate classification in the face of complex situations such as structural mutations and semantic fluctuations, has always been a difficult point in text classification research.
[0004] However, in the process of implementing the inventive technical solution in the embodiments of the present application, the inventors of the present application found that the above technologies have the following technical problems: Traditional text classification methods have problems of semantic continuity and loss of context information when dealing with long texts. Especially when there are significant structural semantic mutations such as topic transitions, turning relationships, or conclusion endings in the text, traditional models often cannot accurately capture the semantic changes brought about by these mutations; for the problem of class imbalance in long texts, traditional classifiers tend to overfit to majority class samples and ignore the recognition of minority classes. Summary of the Invention
[0005] The present invention provides a long text semantic classification method based on deep learning, which solves the problems that traditional text classification methods are prone to semantic incoherence and loss of context information when dealing with long texts. Especially when there are significant structural semantic mutations such as topic transitions, turning relationships, or conclusion endings in the text, traditional models often cannot accurately capture the semantic changes brought about by these mutations; for the problem of class imbalance in long texts, traditional classifiers tend to overfit to majority class samples and ignore the recognition of minority classes.
[0006] A long text semantic classification method based on deep learning according to the present invention specifically includes the following technical solutions: A long text semantic classification method based on deep learning includes the following steps: S1. Perform structured cleaning and sentence boundary recognition on the original text to obtain a sentence sequence; based on the sentence sequence, construct semantic fragments; perform encoding processing on the semantic fragments to generate a sequence of segment-level semantic vectors; S2. Based on the sequence of segment-level semantic vectors, use the semantic tension-driven aggregation algorithm to generate a global semantic representation vector; S3. Based on the global semantic representation vector, design a semantic flow enhancement classifier to complete text classification.
[0007] Preferably, S1 specifically includes: Perform structured cleaning on the original text through regular expressions to obtain the text after structured cleaning; perform sentence boundary recognition on the text after structured cleaning through natural language processing tools to obtain a sentence sequence; introduce a sliding window strategy, and based on the sentence sequence, construct semantic fragments; add a [CLS] token at the beginning of the entire text after structured cleaning, and use [SEP] tokens to separate between each semantic fragment.
[0008] Preferably, S1 specifically includes: Perform sub-word level tokenization on the semantic fragments to obtain sub-word units, and obtain token IDs, segment identifiers, and position identifiers; input the sequence of token IDs, the sequence of segment identifiers, and the sequence of position identifiers into a pre-trained language model to perform context modeling on the semantic fragments and generate semantic vectors; perform weighted fusion on the output vector corresponding to the [CLS] token and the average value of the semantic vectors of all sub-word units within the semantic fragment to generate a segment-level semantic vector, and perform batch normalization on all segment-level semantic vectors to obtain a sequence of segment-level semantic vectors.
[0009] Preferably, S2 specifically includes: In the implementation process of the semantic tension-driven aggregation algorithm, based on the segment-level semantic vectors, introduce the Euclidean distance to calculate the tension value between adjacent fragments of the semantic fragments.
[0010] Preferably, S2 specifically includes: Based on the tension value between adjacent fragments of the semantic fragments, construct an aggregation weight coefficient to adjust the influence degree of the semantic fragments on the global semantic representation.
[0011] Preferably, S2 specifically includes: Based on the aggregation weight coefficient, perform weighted summation on the segment-level semantic vectors to generate a global semantic representation vector.
[0012] Preferably, S3 specifically includes: In the specific implementation process of the semantic flow enhanced classifier, predefined categories are regarded as nodes, the semantic representations of the category nodes are initialized through the global semantic representation vectors, and the connection weights between the category nodes are calculated by combining the semantic similarities of the category nodes to construct a semantic flow network.
[0013] Preferably, the S3 specifically includes: Through the graph propagation algorithm, based on the connection weights between the category nodes, the semantic representations of the category nodes are adjusted to obtain the updated semantic representations of the category nodes; the updated semantic representations of the category nodes are input into the Softmax layer, the probability distribution of each category is calculated, and the prediction probability of each category is output, and the category with the highest prediction probability is selected as the final prediction result.
[0014] Preferably, the S3 specifically includes: In the specific implementation process of the semantic flow enhanced classifier, a loss function is constructed by combining the standard cross-entropy loss, focal loss, and adversarial loss to optimize the semantic flow enhanced classifier.
[0015] The beneficial effects of the technical solution of the present invention are: 1. Through structured cleaning, semantic segment division, and encoding processing, a unified and standardized sequence of segment-level semantic vectors is generated, converting the original long text into processable semantic segments, ensuring the context continuity of the text, and using a pre-trained language model to extract fine-grained semantic information; introducing a sliding window strategy and batch normalization to further enhance the semantic consistency between semantic segments; the generated sequence of segment-level semantic vectors provides high-quality semantic representations for subsequent semantic aggregation and classification tasks.
[0016] 2. Through the semantic tension-driven aggregation algorithm, based on the sequence of segment-level semantic vectors, the aggregation weight coefficients of each semantic segment are dynamically adjusted to effectively capture semantic mutations in the text. By calculating the tension values between adjacent segments of the semantic segments, the focus on the information complexity and structural mutation parts is strengthened; the semantic tension-driven aggregation algorithm does not rely on training parameters and can automatically focus on important semantic structures only through structural tension modeling, improving the modeling ability of long texts and ensuring that the generated global semantic representation better reflects the overall semantic characteristics of the text.
[0017] 3. Design a semantic flow enhanced classifier, combining a semantic flow network, Focal Loss, and adversarial loss, to effectively capture the semantic flow and relationships between category nodes, improving the accuracy and robustness of text classification. The semantic flow network simulates the semantic interaction between category nodes and dynamically adjusts the semantic representations of various category nodes, enabling the semantic flow enhanced classifier to adaptively focus on the dependency relationships between category nodes. Especially in the case of class imbalance and adversarial samples, it shows stronger adaptability and anti-interference ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Flowchart of a long text semantic classification method based on deep learning according to the present invention. Detailed implementation manners
[0019] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0021] The following specifically describes the specific solution of a long text semantic classification method provided by the present invention in conjunction with the accompanying drawings.
[0022] Referring to the attached Figure 1 , which shows a flowchart of a long text semantic classification method provided by an embodiment of the present invention. The method includes the following steps: S1. Structurally clean the original text and identify sentence boundaries to obtain a sentence sequence; based on the sentence sequence, construct semantic segments; perform encoding processing on the semantic segments to generate a segment-level semantic vector sequence; First, structurally clean the original text and divide semantic segments; use existing regular expressions to remove HTML tags, special symbols, and redundant spaces to obtain the structurally cleaned text; use natural language processing tools (such as spaCy, Jieba) to identify sentence boundaries in the structurally cleaned text and divide the structurally cleaned text into an ordered sentence sequence , represents the th sentence in the division, represents the length of the sentence sequence. Considering the limitation of the input length of pre-trained language models such as BERT, let the maximum allowed number of Tokens be . On the premise of retaining special marker bits, traverse the sentence sequence and accumulate the Token quantity to construct semantic segments with a length not exceeding . At the same time, to enhance the context continuity between semantic segments, a sliding window strategy is introduced when generating the next semantic segment, and some overlapping sentences are retained between adjacent semantic segments; Each semantic segment is used as an independent input unit and fed into a pre-trained language model for encoding. Before inputting into the pre-trained language model, a [CLS] token is added at the beginning of the entire structured and cleaned text, and [SEP] tokens are used to separate semantic segments. The semantic segments are tokenized at the sub-word level using the tokenizer corresponding to the selected pre-trained language model, splitting the semantic segments into finer-grained sub-word units (Tokens). Each sub-word unit is mapped to a unique Token ID (token ID) through a vocabulary, and a Segment ID (segment identifier) for distinguishing multi-sentence input and a Position ID (position identifier) for position encoding are generated to meet the input format requirements of the pre-trained language model; the vocabulary is obtained through the analysis and tokenization of a large amount of text data during the pre-training phase.
[0023] The sequence of token IDs, segment identifiers, and position identifiers is input into a pre-trained language model such as BERT. The pre-trained language model performs context modeling on the semantic segments through a multi-layer Transformer structure and outputs the semantic vectors corresponding to each sub-word unit. The output vector corresponding to the [CLS] token represents the global semantic information of the entire structured and cleaned text. To enhance the fine-grained ability of semantic representation, the output vector corresponding to the [CLS] token is weighted and fused with the average of the semantic vectors corresponding to all sub-word units within the semantic segment to generate a segment-level semantic vector. The specific formula is: , where, represents the segment-level semantic vector of the th semantic segment, , is the total number of semantic segments; represents the output vector corresponding to the [CLS] token; represents the average of the semantic vectors corresponding to all sub-word units within the th semantic segment; is the fusion weight, obtained through experiments, and the recommended value is 0.7.
[0024] To improve the consistency of the representations of each semantic segment, batch normalization is performed on all segment-level semantic vectors to make them have a unified numerical distribution in the feature space. Finally, the structured and cleaned text is converted into a sequence of segment-level semantic vectors , where is the total number of semantic segments.
[0025] S2. Based on the sequence of segment-level semantic vectors, a global semantic representation vector is generated through a semantic tension-driven aggregation algorithm; Considering that there are often structural semantic mutations such as topic transitions, turning relationships, or conclusion summaries in long texts, a semantic tension-driven aggregation algorithm is designed. Taking the sequence of segment-level semantic vectors obtained by S1 as input, by modeling the degree of semantic change between semantic segments, it guides the weight assignment of segment-level semantic vectors during the aggregation process, thereby generating a more structure-sensitive global semantic representation vector.
[0026] The specific implementation process of the semantic tension-driven aggregation algorithm is as follows: First, calculate the semantic tension value between each semantic segment and its previous semantic segment, that is, the Euclidean distance between adjacent segment-level semantic vectors, to represent the semantic fluctuation degree of adjacent semantic segments: , where, represents the semantic tension value between the th semantic segment and its previous semantic segment, that is, the tension value between adjacent segments of the th semantic segment; represents the L2 norm operation; is the total number of semantic segments; represents the segment-level semantic vector of the th semantic segment; represents the segment-level semantic vector of the th semantic segment; when , define , indicating that the first paragraph of the text does not have a forward tension relationship. The tension value between adjacent segments of a semantic segment reflects the degree of difference in semantic expression between adjacent semantic segments. The larger the tension value between adjacent segments of a semantic segment, the more drastic the semantic mutation, and higher weights should be given in the subsequent semantic aggregation process; Next, construct an aggregation weight coefficient based on the tension value between adjacent segments of each semantic segment to adjust the influence degree of each semantic segment on the global semantic representation: , where, represents the aggregation weight coefficient of the th semantic segment; is a smoothing factor, used to prevent the denominator from being zero when the tension value between adjacent segments of all semantic segments is zero, and usually takes ; is the total number of semantic segments; represents the The tension value between adjacent segments of a semantic segment. The calculation method of the aggregation weight coefficient can ensure that semantic segments with drastic changes in semantic structure (i.e., semantic segments with large tension values between adjacent segments) have a higher contribution degree in global semantic modeling, thereby reflecting their dominant role in the overall structure; After obtaining the aggregation weight coefficients of all semantic segments, perform weighted summation on all segment-level semantic vectors according to the above aggregation weight coefficients, that is, multiply the segment-level semantic vector of each semantic segment by its corresponding aggregation weight coefficient and add all the weighted segment-level semantic vectors to generate a global semantic representation vector , where denotes the real number space of dimension. The global semantic representation vector synthesizes the information of all semantic segments, and the semantic segments with significant structural mutations dominate in the semantic expression direction, thus better reflecting the overall semantic features of the text.
[0027] Compared with traditional mean pooling or attention-based weighting methods, the semantic tension-driven aggregation algorithm does not rely on training parameters and can automatically focus on important semantic structures only through structural tension modeling, with stronger structural adaptability and semantic interpretability. It is especially suitable for long text modeling and can handle complex text structures and information-intensive scenarios.
[0028] S3. Design a semantic flow enhancement classifier based on the global semantic representation vector to complete text classification.
[0029] Based on the global semantic representation vector, design a semantic flow enhancement classifier to effectively capture the semantic flow and association between different categories in the text, thereby improving the classification accuracy; the specific implementation process of the semantic flow enhancement classifier is as follows: First, construct a semantic flow network based on the global semantic representation vector to simulate the relationship between different categories in the text after structured cleaning; regard each predefined category as a node of the semantic flow network, and use the global semantic representation vector as the initial feature input to the semantic flow network to initialize the semantic representation of the category node, and calculate the connection (i.e., edge) weight between category nodes based on the semantic similarity of the category nodes to generate a weighted directed graph to represent the mutual relationship between categories; the semantic similarity of the category nodes can be calculated using the existing cosine similarity; Next, through a graph propagation algorithm (such as a graph convolutional network), the semantic representation of the category nodes is adjusted according to the connection weights between the category nodes, gradually fusing the semantic representations from adjacent category nodes, thereby simulating the flow and interaction of information between category nodes, and updating the semantic representation of each category node; the updated semantic representation of the category nodes is input into the Softmax layer, the probability distribution of each category is calculated, and the predicted probability of each category is output; Taking the standard cross-entropy loss as the core, a loss function is constructed to optimize the semantic flow enhanced classifier; the standard cross-entropy loss is responsible for measuring the difference between the probability distribution output by the semantic flow enhanced classifier and the manually annotated true labels in the semantic flow enhanced classifier; at the same time, in order to improve the robustness of the semantic flow enhanced classifier under class imbalance and adversarial samples, the Focal Loss is used to increase the penalty for difficult-to-classify samples (especially minority class samples), reduce the loss weight for easy-to-classify samples, thereby enhancing the adaptability of the semantic flow enhanced classifier to class-imbalanced data; further, adversarial samples are generated through methods such as FGSM (a gradient-based adversarial sample generation algorithm), and the adversarial samples are input into the semantic flow network to enhance the resistance of the semantic flow enhanced classifier to perturbed data, enabling it to still make accurate predictions in the face of interference; therefore, the final loss function is a weighted combination of the standard cross-entropy loss, the focal loss, and the adversarial loss, ensuring that the semantic flow enhanced classifier not only focuses on accuracy during optimization but also improves its adaptability to complex data; the standard cross-entropy loss, the focal loss, and the adversarial loss are all existing technologies and will not be elaborated here.
[0030] Through the semantic flow enhanced classifier, it is able to adaptively focus on the dependency relationships between categories, thereby improving the classification accuracy. In the text classification task, the semantic flow enhanced classifier selects the category with the highest predicted probability as the final prediction result based on the calculated predicted probability of each category.
[0031] In summary, a long text semantic classification method based on deep learning is completed.
[0032] The sequence of the invention embodiments is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0033] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between each embodiment, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments.
[0034] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A long text semantic classification method based on deep learning, characterized in that It includes the following steps: S1. Perform structured cleaning and sentence boundary recognition on the original text to obtain a sentence sequence; based on the sentence sequence, construct semantic fragments; perform encoding processing on the semantic fragments to generate a sequence of segment-level semantic vectors; S2. Based on the sequence of segment-level semantic vectors, use the semantic tension-driven aggregation algorithm to generate a global semantic representation vector; S3. Based on the global semantic representation vector, design a semantic flow enhanced classifier to complete text classification.
2. The method for long text semantic classification based on deep learning according to claim 1, characterized in that The specific content of S1 includes: Perform structured cleaning on the original text through regular expressions to obtain the text after structured cleaning; perform sentence boundary recognition on the text after structured cleaning through natural language processing tools to obtain a sentence sequence; introduce a sliding window strategy and construct semantic fragments based on the sentence sequence; add a [CLS] token at the beginning of the entire text after structured cleaning, and use [SEP] tokens to separate each semantic fragment.
3. A long text semantic classification method based on deep learning according to claim 2, characterized in that, The specific content of S1 includes: Perform sub-word level tokenization on the semantic fragments to obtain sub-word units, and obtain token IDs, segment identifiers, and position identifiers; input the token ID sequence, segment identifier sequence, and position identifier sequence into a pre-trained language model to perform context modeling on the semantic fragments and generate semantic vectors; perform weighted fusion on the output vector corresponding to the [CLS] token and the average value of the semantic vectors of all sub-word units within the semantic fragment to generate a segment-level semantic vector, and perform batch normalization on all segment-level semantic vectors to obtain a sequence of segment-level semantic vectors.
4. A long text semantic classification method based on deep learning according to claim 1, characterized in that The specific content of S2 includes: During the implementation of the semantic tension-driven aggregation algorithm, introduce the Euclidean distance based on the segment-level semantic vectors to calculate the tension value between adjacent fragments of the semantic fragments.
5. A long text semantic classification method based on deep learning according to claim 4, characterized in that, The specific content of S2 includes: Construct an aggregation weight coefficient based on the tension value between adjacent fragments of the semantic fragments to adjust the influence degree of the semantic fragments on the global semantic representation.
6. A method for long text semantic classification based on deep learning according to claim 5, characterized in that The specific content of S2 includes: Perform weighted summation on the segment-level semantic vectors based on the aggregation weight coefficient to generate a global semantic representation vector.
7. A method for long text semantic classification based on deep learning according to claim 1, characterized in that The specific content of S3 includes: During the specific implementation of the semantic flow enhanced classifier, regard the predefined categories as nodes, initialize the semantic representation of the category nodes through the global semantic representation vector, and calculate the connection weights between the category nodes by combining the semantic similarity of the category nodes to construct a semantic flow network.
8. A long text semantic classification method based on deep learning according to claim 7, characterized in that The specific content of S3 includes: Through the graph propagation algorithm, adjust the semantic representation of the category nodes based on the connection weights between the category nodes to obtain the updated semantic representation of the category nodes; input the updated semantic representation of the category nodes into the Softmax layer, calculate the probability distribution of each category, and output the prediction probability of each category, and select the category with the highest prediction probability as the final prediction result.
9. A method for long text semantic classification based on deep learning according to claim 8, characterized in that, The specific content of S3 includes: During the specific implementation of the semantic flow enhanced classifier, combine the standard cross-entropy loss, focal loss, and adversarial loss to construct a loss function to optimize the semantic flow enhanced classifier.
Citation Information
Patent Citations
Aspect-level sentiment analysis method based on BERT neural network and multi-semantic learning
CN114579707A
Multi-round dialogue reply generation method based on dual-channel semantic enhancement and terminal equipment
CN115495552A
Aspect sentiment classification method fusing local and global contexts
CN117407759A
A multi-channel short text matching method integrating external knowledge and syntactic structure
CN119782512A
Semantic selection method based on LongformerBERT model
CN119849507A
Cited By
Malicious URL detection method and system based on semantic feature fusion and enhancement
CN122528913A
Malicious URL detection method and system based on semantic feature fusion and enhancement
CN122528913B