Text enhancement system, method and equipment for multi-scale context fusion and medium

By using a multi-scale context fusion text enhancement system, the problems of context information decay and high computational complexity in long text processing of the transformer model are solved, the sensitivity and generalization ability of the model are improved, and efficient text understanding is achieved.

CN120995390APending Publication Date: 2025-11-21DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511109896.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing transformer models suffer from problems such as contextual information decay, high computational complexity, insufficient adaptability, fixed positional encoding methods, and inflexible loss function design when processing long texts, making it difficult to effectively capture multi-scale contextual information.

Method used

A text enhancement system employing multi-scale context fusion is proposed, which includes data preprocessing, dynamic positional encoding, multi-scale feature extraction, memory enhancement network, and hierarchical semantic aggregation network. By hierarchically fusing rotational positional encoding, local-global attention fusion, and adaptive loss function, the system improves the model's sensitivity and generalization ability to long texts.

Benefits of technology

It significantly enhances the model's sensitivity and generalization ability to long texts, reduces computational overhead and the number of training parameters, and achieves an efficient and accurate end-to-end text understanding framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995390A_ABST
    Figure CN120995390A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale context fusion text enhancement system, method and device and a medium, and relates to the technical field of natural language process.The method comprises the steps that cleaning processing, word segmentation processing and vectorization processing are conducted on a to-be-processed text, and an initial vector is obtained; performing position coding processing on the initial vector based on a hierarchical fusion rotation position coding mechanism to obtain a feature vector with position information; performing fusion processing on the feature vectors with the position information through a local-global attention fusion layer to obtain a multi-scale feature map; processing the multi-scale feature map based on a memory enhancement network and a hierarchical semantic aggregation network to obtain enhanced context representation; and predicting a task prediction result corresponding to the to-be-processed text based on the enhanced context representation. According to the method, the training parameter quantity and the tuning cost are remarkably reduced, and an end-to-end efficient, accurate and high-generalization-ability text understanding framework is integrally realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, and in particular to a text enhancement system and method based on multi-scale context fusion, a device and a medium. BACKGROUND

[0002] With the development of deep learning technology, neural network models represented by transformers have made breakthrough progress in the field of natural language processing. The existing technology mainly has the following problems:

[0003] 1. The traditional transformer model has a context information decay problem when processing long text;

[0004] 2. The standard attention mechanism has high computational complexity and is difficult to process ultra-long sequences;

[0005] 3. The existing model is not adaptive enough to different domain texts;

[0006] 4. The position encoding method is fixed, which limits the model's understanding ability of sequence structure;

[0007] 5. The loss function design is not flexible enough to optimize multiple text understanding goals at the same time.

[0008] Therefore, there is an urgent need for a text processing system that can effectively capture multi-scale context information while having self-adaptive learning ability. SUMMARY

[0009] The technical problem to be solved by the present application is to overcome the shortcomings of the prior art, and specifically to provide a text enhancement system and method based on multi-scale context fusion, a device and a medium, as follows:

[0010] 1) In a first aspect, the present application provides a text enhancement system based on multi-scale context fusion, and the specific technical solutions are as follows:

[0011] The data preprocessing module is used for cleaning, word segmentation and vectorization of the text to be processed, and an initial vector is obtained;

[0012] The dynamic position encoding module is used for position encoding processing of the initial vector based on a hierarchical fusion rotary position encoding mechanism to obtain a feature vector with position information;

[0013] The multi-scale feature extraction module is used for fusion processing of the feature vector with position information through a local-global attention fusion layer to obtain a multi-scale feature map;

[0014] The context enhancement module is used for processing the multi-scale feature map based on a memory enhancement network and a hierarchical semantic aggregation network to obtain an enhanced context representation;

[0015] The task adaptation module is used for predicting a task prediction result corresponding to the text to be processed based on the enhanced context representation.

[0016] The text enhancement system with multi-scale context fusion provided by the application has the following beneficial effects:

[0017] The data preprocessing module is used for precise cleaning, word segmentation and vectorization of the text, so as to ensure the input quality and provide a high-purity initial vector for subsequent processing; the dynamic position encoding module uses a hierarchical fusion rotating position encoding mechanism to inject global and local relative position information in the embedding stage, so as to significantly enhance the sensitivity of the model to long-range dependence and word order changes; the multi-scale feature extraction module uses a local-global attention fusion layer to capture fine-grained phrase features and document-level semantics in the same representation space, so as to form a multi-scale feature map with details and overview, and reduce information loss; the context enhancement module dynamically stores and retrieves historical knowledge through a memory enhancement network, and combines a hierarchical semantic aggregation network to extract multi-layer semantics such as themes, events and sentiments step by step, so as to simultaneously strengthen the context representation in depth and breadth, and improve the modeling capability for ambiguity, implicit semantics and cross-sentence logic; the task adaptation module directly reuses the high-quality context representation, so as to quickly migrate among various downstream tasks such as classification, generation and extraction without additional structure adjustment, significantly reduce the training parameter quantity and optimization cost, and realize an end-to-end text understanding framework with high efficiency, precision and strong generalization capability.

[0018] On the basis of the above-mentioned scheme, the application can also be improved as follows.

[0019] Further, the determination process of the feature vector with position information is specifically as follows:

[0020] Based on the token feature matrix corresponding to the initial vector, the position encoding parameter calculation mechanism and the rotating position encoding calculation mechanism in the hierarchical fusion rotating position encoding mechanism are combined to determine the feature vector with position information.

[0021] The beneficial effects of the above-mentioned further scheme are as follows: by deeply coupling the token feature matrix with the position encoding parameter calculation and the rotating position encoding calculation in the hierarchical fusion rotating position encoding mechanism, each token can obtain its semantic representation while being fused with multi-level and multi-granularity relative position information in a rotating form, so as to avoid the extrapolation error of absolute position encoding, significantly compress the parameter quantity through parameter sharing and hierarchical fusion, and greatly improve the perception capability of the model for long-range dependence, word order changes and cross-sentence structure in long texts on the premise of maintaining linear complexity and parallelization advantages, so that the finally generated feature vector with position information has fine-grained local order sensitivity and global structure perception, and lays a high-robustness and low-redundancy position perception foundation for subsequent multi-scale feature extraction and context enhancement.

[0022] Further, the determination process of the multi-scale feature map is specifically:

[0023] The feature vector with position information is sequentially subjected to dynamic convolution processing, gate fusion processing, QKV construction processing, and weighted fusion processing to obtain the multi-scale feature map.

[0024] The above further scheme corresponds to the beneficial effects:

[0025] Through the dynamic convolution processing, the convolution kernel weight is first adaptively adjusted according to the input content, so that the model captures local n-gram and phrase-level patterns in fine granularity; the gate fusion processing uses a learnable gating mechanism to selectively retain and filter the convolution output and the original features, suppresses noise and strengthens key information; the QKV construction processing maps the multi-scale representation after the gating to a query-key-value triple, providing a semantic alignment basis for subsequent attention calculation; the weighted fusion processing uses attention weights to weight and aggregate features of different scales and different channels, realizing the explicit unification of local details and global semantics. The overall process completes local pattern extraction, noise suppression, semantic alignment, and multi-scale fusion in the same feature space, without the need to stack a large number of convolution or attention layers to generate a multi-scale feature map with high information density and low redundancy, which not only reduces the computational overhead, but also provides a high-quality input for the subsequent context enhancement module with fine-grained local description and cross-scale global association, significantly improving the modeling ability and generalization performance of the downstream task for complex language phenomena.

[0026] Further, the determination process of the enhanced context representation is specifically:

[0027] The multi-scale feature map is subjected to read memory processing and write memory processing by the memory enhancement network, and the enhanced representation is spliced based on the read memory processing result and the write memory processing result;

[0028] The enhanced representation is subjected to three-layer granularity aggregation processing by the hierarchical semantic aggregation network to obtain the enhanced context representation.

[0029] The above further scheme corresponds to the beneficial effects:

[0030] The explicit storage-retrieval of read-write separation is performed on the multi-scale feature map through the memory enhancement network, high-frequency, key or cross-sentence dependent information is written into the external memory in real time and dynamically read, the long-range forgetting of the Transformer is relieved and knowledge is shared across samples; the enhanced representation obtained by splicing the read memory result and the write memory result carries the current context and the accumulated priori; the hierarchical semantic aggregation network aggregates from bottom to top in three layers (word-phrase-chapter) to capture fine-grained dependencies, refine event-level structures and condense theme-level semantics, forming a cross-level consistent context representation, so that the model has stronger analysis, memory and generalization ability for implicit reference, long-distance reasoning and multi-theme interwoven text, and significantly reduces the parameter quantity and training time of fine-tuning.

[0031] 2) In a second aspect, the present application also provides a multi-scale context fusion text enhancement method, and the specific technical solutions are as follows:

[0032] The to-be-processed text is subjected to cleaning processing, word segmentation processing and vectorization processing, and an initial vector is obtained;

[0033] The initial vector is subjected to position coding processing based on a hierarchical fusion rotary position coding mechanism, and a feature vector with position information is obtained;

[0034] The feature vector with position information is subjected to fusion processing through a local-global attention fusion layer, and a multi-scale feature map is obtained;

[0035] The multi-scale feature map is processed based on a memory enhancement network and a hierarchical semantic aggregation network to obtain an enhanced context representation;

[0036] The enhanced context representation is used to predict a task prediction result corresponding to the to-be-processed text.

[0037] 3) In a third aspect, the present application also provides an electronic device, which comprises a processor and a memory, the memory is coupled with the processor, and the memory stores at least one computer program, the at least one computer program is loaded and executed by the processor, so that the electronic device realizes any one of the above methods.

[0038] 4) In a fourth aspect, the present application also provides a computer readable storage medium, which stores at least one computer program, the at least one computer program is loaded and executed by a processor, so that the computer realizes any one of the above methods.

[0039] It should be noted that the technical solutions of the second aspect to the fourth aspect of the present application and the corresponding possible implementation manners have the beneficial effects as described above for the first aspect and the corresponding possible implementation manners, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0040] Other features, objects, and advantages of the application will become more apparent from the following detailed description when read in conjunction with the accompanying drawings:

[0041] Figure 1 A flowchart of a multi-scale context fusion text enhancement method according to an embodiment of the application;

[0042] Figure 2 A structural framework diagram of an electronic device according to an embodiment of the application. DETAILED DESCRIPTION

[0043] In order to make the purposes, technical solutions and advantages of the application clearer, the embodiments of the application will be further described in detail below with reference to the drawings.

[0044] As shown in Figure 1 A multi-scale context fusion text enhancement system according to an embodiment of the application includes the following steps:

[0045] The data preprocessing module is configured to perform cleaning processing, word segmentation processing and vectorization processing on the text to be processed, and obtain an initial vector.

[0046] The dynamic position encoding module is configured to perform position encoding processing on the initial vector based on a hierarchical fusion rotation position encoding mechanism, and obtain a feature vector with position information.

[0047] The multi-scale feature extraction module is configured to perform fusion processing on the feature vector with position information through a local-global attention fusion layer, and obtain a multi-scale feature map.

[0048] The context enhancement module is configured to perform processing on the multi-scale feature map based on a memory enhancement network and a hierarchical semantic aggregation network, and obtain an enhanced context representation.

[0049] The task adaptation module is configured to predict a task prediction result corresponding to the text to be processed based on the enhanced context representation.

[0050] The multi-scale context fusion text enhancement system provided by the application has the following beneficial effects:

[0051] The data preprocessing module achieves precise cleaning, tokenization, and vectorization of text, ensuring the input quality and providing high-purity initial vectors for subsequent processing. The dynamic position encoding module injects global and local relative position information at the embedding stage using a hierarchical fusion rotation position encoding mechanism, significantly enhancing the model's sensitivity to long-range dependencies and word order changes. The multi-scale feature extraction module captures fine-grained phrase features and discourse-level semantics within the same representation space through a local-global attention fusion layer, forming a multi-scale feature map that combines details and overviews, reducing information loss. The context enhancement module dynamically stores and retrieves historical knowledge through a memory-augmented network, and combines a hierarchical semantic aggregation network to gradually refine multi-layer semantics such as themes, events, and emotions, so that the context representation is strengthened both in depth and breadth, improving the modeling ability for ambiguity, implicit semantics, and cross-sentence logic. The task adaptation module directly reuses the above high-quality context representation, and can quickly migrate between various downstream tasks such as classification, generation, and extraction without additional structural adjustment, significantly reducing the number of training parameters and tuning costs, and overall realizing an end-to-end text understanding framework with high efficiency, precision, and strong generalization ability.

[0052] In another embodiment of this solution, the specific processing process of the data preprocessing module includes:

[0053] For example, if the text to be processed is "The weather in Beijing is really nice!\n\n", after removing HTML tags, control characters, converting full-width to half-width, and unifying whitespace to single spaces for the text to be processed, we get "The weather in Beijing is really nice!".

[0054] Tokenize the processed text through a pre-trained vocabulary (BERT-BPE), and the result is: ['北', '京', '的', '天', '气', '真', '不', '错', '!'].

[0055] Further perform length processing, set the maximum length to 128. If it is too long: truncate the tail; if it is insufficient: <pad>Padding process, get a list of length 9, and further look up table mapping: [679, 1744, 4637, 1290, 7481, 1343, 6821, 722, 8021]; At the same time, add [CLS] at the head and [SEP] at the tail to generate the final ID sequence: [101, 679, 1744, 4637, 1290, 7481, 1343, 6821, 722, 8021, 102, 0, 0, …] (total 128);

[0056] Construct mask and segment ID

[0057] attention mask: the first 11 are 1, and the rest are 0;

[0058] token_type_ids: all 0 (single sentence);

[0059] Convert ID to 512-dimensional vector with pre-trained Embedding table, X_raw∈R^{128×512} (including pad vector).

[0060] In another embodiment of the present scheme, the dynamic position encoding module is specifically used for:

[0061] Determine the token feature matrix according to the initial vector, the process is: X=[x1,x2,...,xn]∈Rn×d

[0062] Where xi is the initial vector (dimension d) of the i-th token, and n is the sequence length.

[0063] The position encoding parameter computer mechanism includes:

[0064] 1. Basic parameters

[0065] Rotation angle frequency:

[0066] θ j =10000 -2j / d

[0067] Where j∈[0,d / 2) is the vector dimension index.

[0068] 2. Dynamic position offset function

[0069] Multi-scale position information:

[0070] g(i)=∑ k w k ·(i mod 2 k )

[0071] Different granularity position features (such as word level, phrase level) are aggregated through learnable weights wk.

[0072] Hiearchical relative position offset:

[0073] δ(l) = β · 2 l

[0074] where l is the current network layer index, and β is a learnable parameter.

[0075] 3. Transformed position index

[0076] Dynamic position function:

[0077] f(i) = i + a g(i)

[0078] where a is a learnable parameter, controlling the weight of multi-scale information.

[0079] Rotated position encoding computer mechanism (HF-RoPE) includes:

[0080] Perform a rotation operation for each position i and dimension 2j / 2j+1:

[0081] 1. Even dimension

[0082]

[0083] 2. Odd dimension

[0084]

[0085] * indicates the cross-position information obtained by the hierarchical offset, which is directly integrated into the rotation matrix.

[0086] Output:

[0087] Feature matrix with position information:

[0088]

[0089] where each PE(x i ) has fused multi-scale position information and hierarchical dynamic offset, and is directly input into the subsequent module.

[0090] In another embodiment of the present scheme, the multi-scale feature extraction module is specifically used for:

[0091] In the adaptive perception layer (APL), first, the dynamic convolution kernel generation process is defined: where the global average pooling: Two-layer MLP + Sigmoid: where m is the number of convolution kernels, and k is the kernel size (3, 5, 7…, etc. Multi-scale candidates). Secondly, the dynamic convolution operation process: for each position i and kernel j: In the adaptive gating fusion process, the gating value: Gating output:

[0092] X APL =G⊙Z conv +(1-G)⊙X pos

[0093] Furthermore, the Local-Global Attention Fusion Layer (LGAFL) is specifically as follows:

[0094] In local attention, the window size w is fixed at M. local (i, j) = 1 if |ij| ≤ w else 0, and calculate the local attention according to the following formula.

[0095] In global sparse attention, the k largest attention scores are retained for each row, and the rest are set to 0 → resulting in a sparse matrix Ssparse; and the global sparse attention is calculated according to the following formula.

[0096] The adaptive weight fusion process is as follows: the weights are calculated using the following formula. And the fusion process is performed using the following formula: Z=λ·A local +(1-λ)·A global .

[0097] In another embodiment of this solution, the context enhancement module is specifically used for:

[0098] Memory Enhancement Network (MEN) includes:

[0099] 1. Reading and memorizing process

[0100] The external memory matrix M ∈ R^{k×d} (randomly initialized and updated during training);

[0101] Calculate the attention weights A = softmax(ZM) T / √d)∈R^{n×k};

[0102] Read out the memory R = AM ∈ R^{n×d} (R contains long-range information).

[0103] 2. Write down the memory process

[0104] Update gate U = σ(W_u[Z;R]+b_u)∈R^{n×d};

[0105] New Memories

[0106] ( It is a k×d matrix averaged along n dimensions, then broadcast.

[0107] 3. Enhanced representation

[0108]

[0109] The E is output to a hierarchical semantic aggregation network HSAN.

[0110] In the hierarchical semantic aggregation network HSAN, S_word∈R^{n×d} is obtained by using Multi-Head(Q, K, V) in the word level phrase level, and S_phrase∈R^{n×d} is obtained by using a lightweight DynamicConv(kernel width 3, 5); in the phrase level sentence level, the sentence is regarded as a fully connected graph, the node feature is S_phrase, and S_sent∈R^{n×d} is obtained by using one layer of GNN(mean aggregation); and after fusion, the following is obtained

[0111] In another embodiment of the present scheme, the task adaptation module is specifically configured to: receive a context representation H∈R^{64×512}, route according to a task_type, obtain logits∈R^{num_labels} by using a linear layer on H[0] in a cls branch, and then output a category by using a softmax; project H into 64×num_tags by using a linear layer in an ner branch, and then decode a sequence label by using a CRF; generate a summary token sequence by using a lightweight Transformer decoder autoregressively until in a sum branch; and first locate entity pairs by using the ner in a rel branch, splice a first token vector of the entity, and then obtain a relationship logits by using a linear layer.

[0112] Further, the process of determining the feature vector with position information is specifically as follows:

[0113] Based on the token feature matrix corresponding to the initial vector, the feature vector with position information is determined by combining a position encoding parameter calculation mechanism in the hierarchical fusion rotation position encoding mechanism and a rotation position encoding calculation mechanism.

[0114] Further, the process of determining the multi-scale feature map is specifically as follows:

[0115] The feature vector with position information is sequentially subjected to dynamic convolution processing, gate fusion processing, QKV construction processing, and weighted fusion processing, to obtain a multi-scale feature map.

[0116] Further, the process of determining the enhanced context representation is specifically as follows:

[0117] The multi-scale feature map is subjected to read memory processing and write memory processing by using the memory enhancement network, and the enhanced representation is spliced and determined based on the read memory processing result and the write memory processing result.

[0118] The enhanced representation is subjected to three-layer granularity aggregation processing by using the hierarchical semantic aggregation network, to obtain the enhanced context representation.

[0119] As Figure 1 shown, the application also provides a multi-scale context fusion text enhancement method, and the specific technical solutions are as follows:

[0120] S1, performing cleaning processing, word segmentation processing and vectorization processing on the text to be processed, and obtaining an initial vector;

[0121] S2, performing position encoding processing on the initial vector based on a hierarchical fused rotary position encoding mechanism to obtain a feature vector with position information;

[0122] S3, performing fusion processing on the feature vector with position information through a local-global attention fusion layer to obtain a multi-scale feature map;

[0123] S4, performing processing on the multi-scale feature map based on a memory enhancement network and a hierarchical semantic aggregation network to obtain an enhanced context representation;

[0124] S5, predicting a task prediction result corresponding to the text to be processed based on the enhanced context representation.

[0125] In embodiment 1, the adaptive transformer text enhancement system based on multi-scale context fusion of the application has an overall architecture including a data preprocessing module, a dynamic position encoding module, a multi-scale feature extraction module, a context enhancement module, an adaptive loss function module and a task adaptation module.

[0126] In the dynamic position encoding module, the application proposes an innovative hierarchical fused rotary position encoding (HF-RoPE) mechanism. Unlike traditional RoPE position encoding, HF-RoPE realizes more accurate modeling of sequence structure by introducing multi-level position information fusion. Specifically, the calculation formula of HF-RoPE is as follows:

[0127] PE(x,i,2j)=x_i*cos(θ_j*f(i))-x_{i+δ(l)}*sin(θ_j*f(i))

[0128] PE(x,i,2j+1)=x_i*sin(θ_j*f(i))+x_{i+δ(l)}*cos(θ_j*f(i))

[0129] where: x_i represents the feature vector at position i; θ_j = 10000^(-2j / d), d is the model dimension; f(i) = i + α*g(i) is the transformation function of position index; g(i) = Σ(w_k*(i mod 2^k)) is the multi-scale position information; δ(l) = β*2^l is the hierarchical related position offset, l is the current layer index; α and β are learnable parameters.

[0130] Compared with traditional RoPE, HF-RoPE has the following innovations:

[0131] 1. Multi-scale position information g(i) is introduced, which enables the model to perceive position relationships of different granularities at the same time;

[0132] 2. Through hierarchical related position offset δ(l), differential processing of position information by different layers is realized;

[0133] 3. Position encoding parameters α and β are learnable, enabling the model to adaptively adjust the position encoding strength.

[0134] In the multi-scale feature extraction module, the module innovatively introduces an adaptive perception layer (APL) and a local-global attention fusion layer (LGAFL).

[0135] Adaptive perception layer:

[0136] The core of APL is to generate convolution kernels of different scales adaptively according to the input text features through a dynamic convolution kernel generation mechanism, so as to capture multi-granularity semantic information. Its calculation process is as follows:

[0137] 1. First, generate dynamic convolution kernel parameters according to input features X ∈ R^{n×d}:

[0138] W_k = σ(MLP(AvgPool(X))) ∈ R^{m×k×d}

[0139] where m is the number of convolution kernels, k is the size of the convolution kernel, and d is the feature dimension.

[0140] 2. Apply dynamic convolution operation:

[0141]

[0142] 3. Introduce adaptive gating mechanism:

[0143] G = σ(W_g*X + b_g) Y = G ⊙ Z + (1-G) ⊙ X

[0144] denotes element-wise multiplication, and W_g and b_g are learnable parameters.

[0145] Local-global attention fusion layer:

[0146] The LGAFL divides the attention mechanism into local attention and global attention branches, and then fuses them through adaptive weights.

[0147] 1. Local attention computation:

[0148] A_{local}(Q,K,V)=softmax((QK^T / √d)⊙M_{local})V

[0149] where M_{local} is the local attention mask, which restricts the attention range within a fixed window.

[0150] 2. Global attention computation:

[0151] A_{global}(Q,K,V)=softmax((QK^T / √d)*S_{sparse})V

[0152] where S_{sparse} is the sparsity matrix, which retains important global attention connections through Top-k selection.

[0153] 3. Adaptive fusion:

[0154] λ=sigmoid(w^T[X_{cls};X_{avg}]+b)

[0155] A_{fusion}=λ*A_{local}+(1-λ)*A_{global}

[0156] X_{cls} is the classification label feature, X_{avg} is the sequence average feature, and w and b are learnable parameters.

[0157] In the context enhancement module, this module contains the innovative Memory-Enhanced Network (MEN) and Hierarchical Semantic Aggregation Network (HSAN).

[0158] Memory-Enhanced Network:

[0159] MEN introduces an external memory matrix and a dynamic update mechanism to enhance the model's long-range dependency modeling ability.

[0160] 1. Memory query operation:

[0161] A_t = softmax((X_t M^T) / sqrt(d))

[0162] R_t = A_t * M

[0163] where M ∈ R^{k×d} is the memory matrix, containing k memory cells.

[0164] 2. Memory update operation:

[0165] U_t = sigma(W_u[X_t; R_t] + b_u)

[0166] M_{new} = (1-U_t)⊙M + U_t⊙(W_m X_t + b_m)

[0167] 3. Enhanced representation generation:

[0168] E_t = W_e[X_t; R_t] + b_e

[0169] Hierarchical Semantic Aggregation Network:

[0170] HSAN achieves fine-grained modeling of text structure through multi-level semantic aggregation.

[0171] 1. Word-level semantic aggregation:

[0172] S_w = MultiHead(Q_w, K_w, V_w)

[0173] 2. Phrase-level semantic aggregation:

[0174] S_p = DynamicConv(S_w)

[0175] 3. Sentence-level semantic aggregation:

[0176] S_s = Graph(S_p)

[0177] where Graph represents the graph neural network operation, used to model the relationship between sentences.

[0178] 4. Multi-level fusion:

[0179] S_{fusion} = W_f[S_w; S_p; S_s] + b_f

[0180] In the adaptive loss function module, the invention proposes an innovative adaptive loss function design, including contrastive learning loss (CLL) and semantic consistency loss (SCL).

[0181] Contrastive learning loss:

[0182] CLL enhances the representation ability of the model by pulling the representations of semantically similar samples closer and pushing the representations of semantically different samples further apart.

[0183] L_{CLL} = -log(exp(sim(z_i, z_i^+) / τ) / Σ(exp(sim(z_i, z_j) / τ)))

[0184] where:

[0185] z_i is the representation vector of sample i; z_i^+ is the positive sample representation of sample i; sim(u, v) is the cosine similarity function; τ is the temperature parameter; N is the batch size.

[0186] The innovation lies in the generation method of positive samples:

[0187] z_i^+ = α*z_i + (1-α)*z_j, where j is the most similar sample to sample i, and α is a dynamically calculated mixing coefficient: α = sigmoid(w_α^T[z_i; z_j; |z_i-z_j|; z_i⊙z_j] + b_α).

[0188] Semantic consistency loss:

[0189] SCL ensures that the model has consistent understanding of the same text from different perspectives:

[0190] L_{SCL} = Σ(β_l*D_{KL}(p_l||p_L))

[0191] where: p_l is the prediction distribution of the l-th layer; p_L is the prediction distribution of the last layer; D_{KL} is the KL divergence; β_l is the inter-layer weight, calculated as: β_l = (l / L)*(1+γ*(entropy(p_l) / entropy(p_L))), γ is a learnable parameter used to adjust the influence of entropy ratio.

[0192] Total loss function:

[0193] The final loss function is the weighted sum of multiple losses:

[0194] L_{total} = λ_1*L_{task} + λ_2*L_{CLL} + λ_3*L_{SCL}

[0195] where L_{task} is the loss of a specific task (such as cross-entropy loss), λ_1, λ_2, λ_3 are dynamically adjusted weights:

[0196] [λ_1, λ_2, λ_3] = softmax(W_λ*h_{cls} + b_λ)

[0197] h_{cls} is the final representation of the classification label.

[0198] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, which is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0199] It should be noted that the beneficial effects of the multi-scale context fusion text enhancement method provided in the above embodiments are the same as those of the multi-scale context fusion text enhancement system described above, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0200] like Figure 2 As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320 coupled to a memory 310. The memory 310 stores at least one computer program 330, which is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above-mentioned methods. Specifically:

[0201] The electronic device 300 can vary considerably due to differences in configuration or performance. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. The one or more memories 310 store at least one computer program 330, which is loaded and executed by the one or more processors 320 to enable the electronic device 300 to implement the multi-scale context fusion text enhancement system provided in the above embodiments. Of course, the electronic device 300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The electronic device 300 may also include other components for implementing device functions, which will not be elaborated here.

[0202] An embodiment of the present invention provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above-described methods.

[0203] Optionally, the computer readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0204] In an example embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the electronic device performs any of the above methods.

[0205] It should be noted that the terms "first", "second" in the specification and claims of the present application are used to distinguish similar objects, and represent the limitation of a specific order or sequence. The order of use of similar objects can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described.

[0206] Those skilled in the art understand that the present application can be implemented as a system, a method or a computer program product, therefore, the present disclosure can be specifically implemented as follows: it can be a complete hardware, a complete software (including firmware, resident software, microcode, etc.), and also a combination of hardware and software, which is generally referred to as "circuit", "module" or "system" herein. In addition, in some embodiments, the present application can also be implemented as a computer program product in one or more computer readable media, which includes computer readable program code.

[0207] ​Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0208] Although the embodiments of the present application have been shown and described above, it should be understood by those skilled in the art that the above embodiments are exemplary, and are not to be understood as limiting the present application, and that changes, modifications, substitutions and variations of the above embodiments can be made by those skilled in the art within the scope of the present application.< / pad>

Claims

1. A multi-scale context fusion-based text enhancement system, characterized in that, The application relates to a text task prediction method and device. The data preprocessing module is used for cleaning, word segmentation and vectorization of the text to be processed, and an initial vector is obtained; The dynamic position coding module is used for position coding processing of the initial vector based on a hierarchical fusion rotary position coding mechanism to obtain a feature vector with position information; The multi-scale feature extraction module is used for fusion processing of the feature vector with position information through a local-global attention fusion layer to obtain a multi-scale feature map; The context enhancement module is used for processing the multi-scale feature map based on a memory enhancement network and a hierarchical semantic aggregation network to obtain an enhanced context representation; The task adaptation module is used for predicting a task prediction result corresponding to the text to be processed based on the enhanced context representation.

2. The multi-scale context fused text enhancement system of claim 1, wherein, The determination process of the feature vector with position information is specifically as follows: Based on a token feature matrix corresponding to the initial vector, a position coding parameter calculation mechanism and a rotary position coding calculation mechanism in the hierarchical fusion rotary position coding mechanism are combined to determine the feature vector with position information.

3. The multi-scale context fused text enhancement system of claim 1, wherein, The determination process of the multi-scale feature map is specifically as follows: The feature vector with position information is sequentially subjected to dynamic convolution processing, gate fusion processing, QKV construction processing and weighted fusion processing to obtain the multi-scale feature map.

4. The multi-scale context fused text enhancement system of claim 1, wherein, The determination process of the enhanced context representation is specifically as follows: The multi-scale feature map is subjected to read memory processing and write memory processing through the memory enhancement network, and the enhanced representation is spliced and determined based on the read memory processing result and the write memory processing result; The enhanced representation is subjected to three-layer granularity aggregation processing through the hierarchical semantic aggregation network to obtain the enhanced context representation.

5. A method of multi-scale context fusion based text enhancement, characterized in that, The data preprocessing module is used for cleaning, word segmentation and vectorization of the text to be processed, and an initial vector is obtained; The dynamic position coding module is used for position coding processing of the initial vector based on a hierarchical fusion rotary position coding mechanism to obtain a feature vector with position information; The multi-scale feature extraction module is used for fusion processing of the feature vector with position information through a local-global attention fusion layer to obtain a multi-scale feature map; The context enhancement module is used for processing the multi-scale feature map based on a memory enhancement network and a hierarchical semantic aggregation network to obtain an enhanced context representation; The task adaptation module is used for predicting a task prediction result corresponding to the text to be processed based on the enhanced context representation. The determination process of the feature vector with position information is specifically as follows:

6. The method of claim 5, wherein, Based on a token feature matrix corresponding to the initial vector, a position coding parameter calculation mechanism and a rotary position coding calculation mechanism in the hierarchical fusion rotary position coding mechanism are combined to determine the feature vector with position information. The determination process of the multi-scale feature map is specifically as follows:

7. The method of claim 5, wherein, The feature vector with position information is sequentially subjected to dynamic convolution processing, gate fusion processing, QKV construction processing and weighted fusion processing to obtain the multi-scale feature map. The determination process of the enhanced context representation is specifically as follows:

8. The method of claim 5, wherein, The multi-scale feature map is subjected to read memory processing and write memory processing through the memory enhancement network, and the enhanced representation is spliced and determined based on the read memory processing result and the write memory processing result; ​ The enhanced representation is subjected to three-layer granularity aggregation processing by a hierarchical semantic aggregation network to obtain the enhanced context representation.

9. An electronic device, comprising: The electronic device includes a processor coupled with a memory, and the memory has at least one computer program stored therein, and the at least one computer program is loaded and executed by the processor, so that the electronic device implements the method according to any one of claims 5 to 8.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium has at least one computer program stored therein, and the at least one computer program is loaded and executed by the processor, so that the computer implements the method according to any one of claims 5 to 8.