Long text matching method and device

Through the three-level progressive sliding window attention and feature enhancer model, the shortcomings of the long text matching model in deep semantic modeling are solved, the semantic association from local grammar to global text is realized, and the accuracy of long text matching is improved.

CN120670869APending Publication Date: 2025-09-19CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510845357.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing long text matching models find it difficult to achieve semantic modeling from local grammar to global text, resulting in insufficient ability to capture deep semantic associations, which in turn affects the accuracy of long text matching.

Method used

A three-level progressive sliding window attention mechanism is adopted, combined with word-level, sentence-level and paragraph-level feature sub-models. Sliding window attention is used to perform local grammatical parsing, cross-sentence semantic association and global intent analysis on long texts, and a long text matching model is constructed. Word-level and sentence-level feature enhancement sub-models are introduced for feature screening and reconstruction, and a word-sentence cross-fusion model is used for semantic association.

Benefits of technology

It improves the ability to capture deep semantic associations in the long text matching process and improves the accuracy of long text matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670869A_ABST
    Figure CN120670869A_ABST
Patent Text Reader

Abstract

The invention provides a long text matching method and a long text matching device, relates to the technical field of natural language processing, and aims to relieve the problem that the long text matching precision is not high due to the fact that a long text matching model in the prior art is insufficient in capture capability for deep semantic association. The method comprises the steps of obtaining a to-be-matched long text; and inputting each long text pair into a trained long text matching model with a dynamic complementary relationship from a local grammar to a global chapter feature so as to improve the capture capability of deep semantic association and the accuracy of long text matching in a long text matching process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a long text matching method and device. Background Art

[0002] Text matching, a core task in NLP (Natural Language Processing), provides critical support for applications such as information retrieval, intelligent question-answering, and sentiment analysis by calculating the semantic relevance between texts. When matching long texts, this technology faces unique challenges due to the need to parse multi-paragraph structures, understand paragraph logic, and process redundant information. For example, existing long text matching models often struggle to achieve semantic modeling from local grammar to global text. As a result, existing long text matching models are unable to capture deep semantic connections during long text matching, resulting in low long text matching accuracy. Summary of the Invention

[0003] In view of this, an object of the present invention is to provide a long text matching method and apparatus to improve the ability to capture deep semantic associations during long text matching and the accuracy of long text matching.

[0004] In a first aspect, the present invention provides a long text matching method, comprising: Obtaining a long text combination to be matched; wherein the long text combination includes at least two long texts; Each long text is input into the long text matching model to obtain the long text matching results between each long text; wherein, the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model and a paragraph-level feature sub-model connected in sequence, and a matching sub-model connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model respectively; the long text matching model uses the first sliding window attention to perform local grammatical analysis on each long text through the word-level feature sub-model to obtain the word-level features of each long text, uses the second sliding window attention to perform cross-sentence semantic association on the word-level features of each long text through the sentence-level feature sub-model to obtain the sentence-level features of each text, uses the third sliding window attention to perform global intention analysis on the sentence-level features of each long text through the paragraph-level feature sub-model to obtain the paragraph-level features of each long text, and uses the matching sub-model to match each long text based on the word-level features, sentence-level features and paragraph-level features of each long text to obtain the long text matching results; the semantic perception scopes of the first sliding window attention, the second sliding window attention and the third sliding window attention are in a progressive relationship.

[0005] Optionally, the long text matching model further includes a word-level feature enhancement sub-model and a sentence-level feature enhancement sub-model, wherein the word-level feature sub-model is connected to the matching sub-model via the word-level feature enhancement sub-model, and the sentence-level feature sub-model is connected to the matching sub-model via the sentence-level feature enhancement sub-model; the long text matching method further includes: The word-level features of each long text are filtered for relational features through the word-level feature enhancer model to obtain the first enhanced word-level features of each long text; The sentence-level features of each long text are semantically reconstructed through the sentence-level feature enhancement sub-model to obtain the first enhanced sentence-level features; The matching sub-model is used to match each long text based on the first enhanced word-level features, the first enhanced sentence-level features and the paragraph-level features of each long text to obtain a long text matching result.

[0006] Optionally, the word-level feature enhancer model includes a skip-connected denoising unit and a word-level feature fusion unit; the word-level feature enhancer model is used to extract key features from the word-level features of each long text to obtain a first enhanced word-level feature of each long text, including: After calculating the attention scores of the word-level features of each long text through the denoising unit, based on the attention scores of the word-level features of each long text, the word-level features of each long text are screened for strong correlation features to obtain the intermediate enhanced word-level features of each long text; The word-level features and intermediate enhanced word-level features of each long text are subjected to multimodal feature splicing by a word-level feature fusion unit to obtain a spliced ​​word-level feature of each long text, and the intermediate enhanced word-level features of each long text are subjected to element-level feature fusion to obtain a fused word-level feature of each long text. Then, a first enhanced word-level feature of each long text is determined based on the spliced ​​word-level features and the fused word-level features of each long text; The sentence-level feature enhancer model includes a deep separable unit with skip connections and a sentence-level feature fusion unit. The sentence-level feature enhancer model reconstructs the semantic association of the sentence-level features of each long text to obtain the first enhanced sentence-level feature, including: The local temporal relationship features of the sentence-level features of each long text are extracted in different feature channels through the deep separable unit to obtain the local temporal relationship features of the different feature channels corresponding to each long text. Then, the local temporal relationship features of the different feature channels corresponding to each long text are fused across channels to obtain the fused sentence-level features of each long text. The sentence-level feature fusion unit is used to fuse the sentence-level features and fused sentence-level features of each long text to obtain the first enhanced sentence-level features of each long text.

[0007] Optionally, the attention score of the word-level features of each long text is calculated by the denoising unit, including: Calculate the initial attention score of the word-level features of each long text based on the attention score matrix; Based on the differentiable Gumbel distribution threshold gating, each initial attention score is filtered to obtain the attention score of the word-level features of each long text.

[0008] Optionally, the sentence-level features of each long text and the fused sentence-level features are fused by a sentence-level feature fusion unit to obtain a first enhanced sentence-level feature of each long text, including: After the sentence-level features and fused sentence-level features of each long text are spliced ​​together to obtain the spliced ​​sentence-level features of each long text, the spliced ​​sentence-level features are normalized and regularized to obtain the first enhanced sentence-level features of each long text.

[0009] Optionally, the long text matching model further includes a word-sentence cross-fusion sub-model, the noise reduction unit and the sentence-level feature fusion unit are respectively connected to the word-sentence cross-fusion sub-model; the word-sentence cross-fusion sub-model is connected to the matching sub-model; and the long text matching method further includes: The word-sentence cross-fusion sub-model uses bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced word-level features of each long text to obtain the word-sentence semantic association features of each long text; The matching sub-model is used to match each long text based on the first enhanced word-level features, first enhanced sentence-level features, word-sentence semantic association features and paragraph-level features of each long text to obtain a long text matching result.

[0010] Optionally, the word-sentence cross-fusion sub-model includes a skip-connected interactive correction unit and a feature fusion unit; the word-sentence cross-fusion sub-model uses bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced sentence-level features of each long text to obtain word-level semantic association features of each long text, including: Mapping the intermediate enhanced word-level features of each long text to the query space and mapping the first enhanced sentence-level features of each long text to the key-value space through the interactive correction unit to obtain the word-to-sentence mapping features of each long text, and using the first enhanced sentence-level features of each long text to correct the intermediate enhanced word-level features to obtain the sentence-to-word mapping features of each long text, and then performing nonlinear transformation on the word-to-sentence mapping features and the sentence-to-word mapping features of each long text to obtain the word-to-sentence mapping features of each long text; The first enhanced sentence-level features and word-sentence mapping features of each long text are fused through a feature fusion unit to obtain the word-sentence semantic association features of each long text.

[0011] Optionally, the word-level feature sub-model uses the first sliding window attention to perform local grammatical parsing on each long text to obtain the word-level features of each long text, including: By performing word segmentation on each long text, the word element features of each long text are obtained; Through the first sliding window and local-to-global collaborative attention, each long text is locally parsed to obtain the word-level features of each long text.

[0012] Optionally, the sentence-level features of each long text are obtained by performing cross-sentence semantic association on the word-level features of each long text using the second sliding window attention through the sentence-level feature sub-model, including: The word-level features of each long text are aggregated through mean pooling to obtain the sentence-level vector of each long text; The sentence-level features of each long text are obtained by performing cross-sentence semantic association on the sentence-level vectors of each long text through the second sliding window and cross-document attention.

[0013] Optionally, the segment-level feature sub-model uses the third sliding window attention to perform global intent analysis on the sentence-level features of each long text to obtain the segment-level features of each long text, including: Extract the sentence-level features of each long text through maximum pooling to obtain the initial sentence-level features of each long text; The initial sentence-level features of each long text are analyzed globally through the third sliding window and sparse global attention to obtain the paragraph-level features of each long text.

[0014] Optionally, each long text is matched based on the word-level features, sentence-level features, and paragraph-level features of each long text by a matching sub-model to obtain a long text matching result, including: The first enhanced word-level features, first enhanced sentence-level features, word-sentence semantic association features, and paragraph-level features of each long text are integrated and nonlinearly mapped through the cascaded fully connected layer to obtain the long text matching results.

[0015] In a second aspect, the present invention provides a long text matching device, comprising: A data acquisition module is used to acquire a long text combination to be matched; wherein the long text combination includes at least two long texts; A text matching module is used to input each long text into a long text matching model to obtain long text matching results between each long text; wherein, the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model and a paragraph-level feature sub-model connected in sequence, and a matching sub-model connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model respectively; the word-level feature sub-model of the long text matching model uses a first sliding window attention to perform local grammatical analysis on each long text to obtain the word-level features of each long text, uses a second sliding window attention to perform cross-sentence semantic association on the word-level features of each long text through the sentence-level feature sub-model to obtain the sentence-level features of each text, uses a third sliding window attention to perform global intention analysis on the sentence-level features of each long text through the paragraph-level feature sub-model to obtain the paragraph-level features of each long text, and uses the matching sub-model to match each long text based on the word-level features, sentence-level features and paragraph-level features of each long text to obtain the long text matching results; the semantic perception scopes of the first sliding window attention, the second sliding window attention and the third sliding window attention are in a progressive relationship.

[0016] In a third aspect, the present invention further provides an electronic device comprising: a processor and a memory; the memory stores a computer program that can be run on the processor, and the processor implements the above-mentioned long text matching method when executing the computer program.

[0017] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and the computer program implements the above-mentioned long text matching method when executed by a processor.

[0018] The embodiment of the present invention provides a long text matching method and device, which obtains a long text combination to be matched; wherein the long text combination includes at least two long texts; each long text is input into a long text matching model to obtain a long text matching result between each long text; wherein the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model and a paragraph-level feature sub-model connected in sequence, and a matching sub-model connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model respectively; the word-level feature sub-model of the long text matching model uses a first sliding window attention to perform local language analysis on each long text The method is used to parse the word-level features of each long text, and the sentence-level feature sub-model adopts the second sliding window attention to perform cross-sentence semantic association on the word-level features of each long text to obtain the sentence-level features of each text. The paragraph-level feature sub-model adopts the third sliding window attention to perform global intent analysis on the sentence-level features of each long text to obtain the paragraph-level features of each long text. The matching sub-model is used to match each long text based on the word-level features, sentence-level features and paragraph-level features of each long text to obtain the long text matching results, so as to improve the ability to capture deep semantic associations and the accuracy of long text matching in the long text matching process.

[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 A schematic diagram of the process of a long text matching task provided by an embodiment of the present invention is shown; Figure 2 A flowchart of a long text matching model training method provided by an embodiment of the present invention is shown; Figure 3 A schematic diagram showing a flow chart of a long text matching method provided by an embodiment of the present invention is shown; Figure 4 A schematic diagram of the structure of a long text matching model provided by an embodiment of the present invention is shown; Figure 5 A schematic diagram of the structure of a word-level feature enhancement submodule provided by an embodiment of the present invention is shown; Figure 6 A schematic diagram showing the structure of a sentence-level feature enhancement submodule provided in an embodiment of the present invention is shown; Figure 7 A schematic diagram of the structure of the word and sentence cross-fusion sub-model provided in an embodiment of the present invention is shown; Figure 8 A schematic structural diagram of a long text matching device provided by an embodiment of the present invention is shown; Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0023] Text matching is a core task of NLP. It provides key support for applications such as information retrieval, intelligent question answering, and sentiment analysis by calculating the semantic relevance between texts. Figure 1 As shown, given the training set:

[0024] in, and There are two paragraphs of text, and the label , the goal of text matching is to learn a function:

[0025] Enables the model to recognize and Whether it matches.

[0026] Current text matching technology frameworks are primarily based on deep learning models and can be categorized into three types: representational, interactive, and hybrid. Representational methods use neural network encoders (such as DSSM, C-DSSM, RNN-LSTM, etc.) to map text pairs into high-dimensional semantic vectors, and then calculate vector similarity to achieve matching. Their advantage lies in their ability to efficiently process large amounts of text. Interactive models (such as ARC-II, MatchPyramid, and Match-SRNN) are used to capture local semantic information between texts, construct matching matrices using word-level or phrase-level information, and excel at mining the association of word-level features. Hybrid methods fuse representational and interactive methods through a hybrid strategy, and introduce an attention mechanism to dynamically screen key semantic units. For long text scenarios, the SMASH framework systematically defines long text matching tasks for the first time, while the concept interaction graph implements a structured matching method based on semantic concepts by modeling the global association of events or narratives between documents.

[0027] This application embodiment provides a long text matching model training method, see Figure 2As shown, the overview process of a long text matching model training method provided in an embodiment of the present application is as follows: Step 110: Obtain a training sample data set; wherein the training sample data set includes a plurality of training sample data; each training sample data includes a text pair and a standard matching result corresponding to the long text; Step 120: Select target training sample data from the training sample data set; Step 130: Input the text pairs in the target sample data into a long text matching model, wherein the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model, and a paragraph-level feature sub-model connected in sequence, and a matching sub-model connected to the word-level feature sub-model, the sentence-level feature sub-model, and the paragraph-level feature sub-model respectively. The long text matching model uses the word-level feature sub-model to perform local grammatical analysis on each long text using a first sliding window attention to obtain word-level features of each long text, uses the sentence-level feature sub-model to perform cross-sentence semantic association on the word-level features of each long text to obtain sentence-level features of each text, uses the paragraph-level feature sub-model to perform global intent analysis on the sentence-level features of each long text using a third sliding window attention to obtain paragraph-level features of each long text, and uses the matching sub-model to match each long text based on the word-level features, sentence-level features, and paragraph-level features of each long text to obtain a long text prediction matching result. Step 140: Based on the predicted matching results and the standard matching results in the target training sample data, a loss function is used to update various model parameters of the long text matching model.

[0028] Step 150, determine whether the iterative training termination condition is met; if so, execute step 160; if not, return to step 120; wherein, the iterative training termination condition is that the number of iterations is not less than the number threshold, or the loss value is not higher than the loss value threshold.

[0029] Step 160: Update the model parameters of the long text matching model based on the last iteration training operation.

[0030] Based on the above embodiment, the present application embodiment provides a flow chart of a long text matching method, see Figure 3 As shown, the overview process of a long text matching method provided in an embodiment of the present application is as follows: Step 210: Obtain a long text combination to be matched; wherein the long text combination includes at least two long texts.

[0031] Step 220: Input each long text into a long text matching model to obtain long text matching results between each long text; wherein, the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model and a paragraph-level feature sub-model connected in sequence, and a matching sub-model connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model respectively; the word-level feature sub-model of the long text matching model uses a first sliding window attention to perform local grammatical analysis on each long text to obtain the word-level features of each long text, uses a second sliding window attention to perform cross-sentence semantic association on the word-level features of each long text through the sentence-level feature sub-model to obtain the sentence-level features of each text, uses a third sliding window attention to perform global intent analysis on the sentence-level features of each long text through the paragraph-level feature sub-model to obtain the paragraph-level features of each long text, and uses the matching sub-model to match each long text based on the word-level features, sentence-level features and paragraph-level features of each long text to obtain the long text matching results; the semantic perception scopes of the first sliding window attention, the second sliding window attention and the third sliding window attention are in a progressive relationship.

[0032] In order to achieve deep semantic modeling of each long text and optimize calculation, in the embodiment of this application, long text matching is based on a long document pre-training model, a three-layer encoding framework is constructed, and semantic modeling from vocabulary to paragraphs is achieved through the collaborative mechanism of global attention and hierarchical sliding window attention. Figure 4 As shown, the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model, a paragraph-level feature sub-model and a matching sub-model, wherein the word-level feature sub-model is connected to the paragraph-level feature sub-model through the sentence-level feature sub-model, and the matching sub-model is connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model respectively; wherein the word-level feature sub-model adopts the first sliding window attention to perform local grammatical analysis on each long text to obtain the word-level features of each long text; the sentence-level feature sub-model adopts the second sliding window attention to perform cross-sentence semantic association on the word-level features of each long text to obtain the sentence-level features of each text; the paragraph-level feature sub-model adopts the third sliding window attention to perform global intent analysis on the sentence-level features of each long text to obtain the paragraph-level features of each long text; the matching sub-model matches each long text based on the word-level features, sentence-level features and paragraph-level features of each long text to obtain the long text matching results.

[0033] Furthermore, each long text is locally parsed using the first sliding window attention through the word-level feature sub-model to obtain the word-level features of each long text, including: obtaining the word-unit features of each long text by performing word segmentation on each long text; and obtaining the word-level features of each long text by performing local grammatical parsing on each long text through the first sliding window and local-to-global collaborative attention.

[0034] Specifically, the word-level feature sub-model adopts a local-to-global collaborative attention mechanism to model the local context of long texts:

[0035] in, and Represents two long texts, and Represents word features, represents word embedding, represents the set of real numbers, and represents the sequence length, represents the embedding dimension, Represents word-level features, Indicates the window size The first sliding window attention and the sliding step size is .

[0036] The word-level feature sub-model ensures local context continuity through the first sliding window overlap. The local to global attention range of each word unit covers the 32 word units before and after, which reduces the computational complexity while retaining the local grammatical structure features. The first sliding window overlapping design can effectively avoid semantic faults and ensure cross-window coherence.

[0037] Furthermore, the sentence-level feature sub-model uses the second sliding window attention to perform cross-sentence semantic association on the word-level features of each long text to obtain the sentence-level features of each long text, including: aggregating the word-level features of each long text through mean pooling to obtain the sentence-level vector of each long text; and performing cross-sentence semantic association on the sentence-level vector of each long text through the second sliding window and cross-document attention to obtain the sentence-level features of each long text.

[0038] Specifically, the sentence-level feature sub-model constructs sentence-level representations based on word-level features and introduces a cross-document attention mechanism to implement inter-sentence relationship reasoning:

[0039] in, and Represents sentence-level features, MeanPool represents mean pooling, Indicates the number of sentences, Indicates the window size The second sliding window and global attention head .

[0040] Aggregate word vectors into sentence vectors through mean pooling, and Expand to , in order to obtain the logical connection features between sentences. Secondly, through the global attention head (i.e., cross-document attention), global attention without position constraints is imposed on key sentences such as the first sentence of a paragraph and sentences containing entity words, which reduces the computational overhead while retaining important cross-document correlation signals.

[0041] Furthermore, the paragraph-level feature sub-model adopts the third sliding window attention to perform global intent analysis on the sentence-level features of each long text to obtain the paragraph-level features of each long text, including: extracting the sentence-level features of each long text through maximum pooling to obtain the initial sentence-level features of each long text; performing global intent analysis on the initial sentence-level features of each long text through the third sliding window and sparse global attention to obtain the paragraph-level features of each long text.

[0042] Specifically, the segment-level feature sub-model adopts diluted global attention to solve the problem of difficulty in capturing the semantics of long texts:

[0043] in, and Represents segment-level features, MaxPool represents maximum pooling, represents the sparsity rate of the sparse global attention of the second sliding window .

[0044] By extracting sentence-level features through maximum pooling and using sparse global attention to obtain cross-sentence associations, it is possible to perform sparse screening of attention weights, retaining the most discriminative core events across documents, and capturing segment-level features under relatively complex conditions.

[0045] Through the long text matching model from local grammar to global chapter, a three-level progressive sliding window attention is used to process each long text to obtain the matching results of each long text, so as to improve the ability to capture deep semantic associations in the long text matching process and the accuracy of long text matching.

[0046] In order to achieve dynamic suppression of noise signals and directional enhancement of key matching signals, in an embodiment of the present application, the long text matching model also includes a word-level feature enhancer model and a sentence-level feature enhancer model. The word-level feature submodel is connected to the matching submodel through the word-level feature enhancer model, and the sentence-level feature submodel is connected to the matching submodel through the sentence-level feature enhancer model; wherein, the word-level feature enhancer model is used to perform relationship feature screening on the word-level features of each long text to obtain the first enhanced word-level features of each long text; the sentence-level feature enhancer model is used to perform semantic association reconstruction on the sentence-level features of each long text to obtain the first enhanced sentence-level features; and the matching submodel is used to match each long text based on the first enhanced word-level features, the first enhanced sentence-level features and the paragraph-level features of each long text to obtain the long text matching results.

[0047] Among them, the word-level feature enhancer model includes a jump-connected denoising unit and a word-level feature fusion unit; after calculating the attention score of the word-level features of each long text through the denoising unit, based on the attention score of the word-level features of each long text, the word-level features of each long text are strongly correlated and screened to obtain the intermediate enhanced word-level features of each long text; through the word-level feature fusion unit, the word-level features and the intermediate enhanced word-level features of each long text are multimodally spliced ​​to obtain the spliced ​​word-level features of each long text, and the intermediate enhanced word-level features of each long text are element-level fused to obtain the fused word-level features of each long text, and then the first enhanced word-level features of each long text are determined based on the spliced ​​word-level features and the fused word-level features of each long text.

[0048] Furthermore, the attention scores of the word-level features of each long text are calculated through the denoising unit, including: calculating the initial attention scores of the word-level features of each long text based on the attention score matrix; filtering each initial attention score based on the differentiable Gumbel distribution threshold gating to obtain the attention scores of the word-level features of each long text.

[0049] like Figure 5 As shown in the figure, the word-level feature enhancer model is used to filter the word-level features of each long text to obtain the first enhanced word-level features of each long text. The specific process is as follows: Given two long texts and Word-level features and ; First, the initial attention scores of the word-level features of each long text are calculated through the attention score matrix of self-attention, for example, For example, generate attention triples based on the attention score matrix:

[0050] Where, 、 and is the weight matrix, is the query matrix in triples, is the bond matrix in the triples, is the value matrix in the triple; The initial attention score of word-level features can be obtained by querying and key Normalize the dot product of calculate:

[0051] in, represents the initial attention score, represents the matrix transpose operation, yes The dimension is used to scale the dot product result to prevent the gradient from disappearing or exploding.

[0052] In order to reduce noise, differentiable Gumbel (Extreme Value Distribution) threshold gating is introduced to achieve dynamic noise filtering:

[0053] in, is the filtered attention score, represents the attention score, Gumbel-Sigmoid represents the function, represents the temperature coefficient, Indicates the exclusive OR operation, Indicates the number of training rounds.

[0054] Furthermore, the temperature coefficient Exponential decay from 1.0 to 0.1; the Gumbel-Sigmoid function achieves differentiable binarization through reparameterization techniques:

[0055] in, for noise, for function, Represents input data.

[0056] Furthermore, dynamic noise filtering is achieved by differentiable Gumbel threshold gating to maintain soft threshold screening in the early stages of training. , gradually transition to hard threshold mode , achieving end-to-end noisy edge learning.

[0057] Using filtered attention scores , the updated intermediate enhanced word-level features are calculated as:

[0058] Intermediate enhancer level features and Can better capture important information in the text.

[0059] In order to further enhance the feature representation, the intermediate enhanced word level features and Splice and calculate the Hadamard product Obtain the first enhanced word-level features in a nonlinear interaction mode :

[0060] in, Represents the features after splicing, Indicates passing The activation function enhances the concatenated features. , represents the projection matrix, represents element-level fusion features, Indicates normalization.

[0061] The noise of word-level features is reduced by attention score and threshold filtering, and the word-level feature representation is further enhanced by combining feature splicing and element-level multiplication, achieving dynamic suppression of noise signals and directional enhancement of key matching signals.

[0062] In an optional embodiment, the sentence-level feature enhancement sub-model includes a deep separable unit with jump connections and a sentence-level feature fusion unit; after extracting local temporal relationship features in different feature channels of the sentence-level features of each long text through the deep separable unit to obtain the local temporal relationship features of different feature channels corresponding to each long text, the local temporal relationship features of the different feature channels corresponding to each long text are subjected to cross-channel feature fusion to obtain the fused sentence-level features of each long text; the sentence-level features and the fused sentence-level features of each long text are subjected to feature fusion through the sentence-level feature fusion unit to obtain the first enhanced sentence-level features of each long text.

[0063] Furthermore, the sentence-level features and the fused sentence-level features of each long text are fused through a sentence-level feature fusion unit to obtain the first enhanced sentence-level features of each long text, including: splicing the sentence-level features and the fused sentence-level features of each long text to obtain the spliced ​​sentence-level features of each long text, and then normalizing and regularizing the spliced ​​sentence-level features to obtain the first enhanced sentence-level features of each long text.

[0064] like Figure 6 As shown in the figure, the sentence-level feature enhancement sub-model is used to reconstruct the semantic association of the sentence-level features of each long text to obtain the first enhanced sentence-level feature. The specific process is as follows: Given sentence-level features ,in, Indicates the number of sentences, Represents feature dimension; First, the standard convolution is decomposed into channel-independent depthwise convolution and channel-mixed pointwise convolution:

[0065] Where, Represents sentence-level features of local temporal relations, represents depthwise separable one-dimensional convolution, represents the convolution kernel, represents the depth-wise convolution kernel, Indicates the fusion of sentence-level features, represents one-dimensional convolution, express Convolution kernel.

[0066] The sentence-level feature enhancer model uses each feature channel independently in the deep convolution stage. dimensional convolution kernel (i.e., the window size is ), sliding along the sentence sequence direction to capture the local temporal sequence between adjacent sentences to obtain the sentence-level features of the local temporal relationship, which can reduce the number of parameters from the standard convolution to Reduce to , reduce the complexity of the model; in the point-by-point convolution stage, use for Convolution kernel, for depth convolution output The linear combination of channels can reconstruct the semantic association between channel dimensions and the parameter quantity can be changed from Transformed into , to focus on cross-channel information fusion.

[0067] Second, the original feature information flow is maintained through skip connections:

[0068] in, represents the processed intermediate enhancement word-level features, Representation layer normalization.

[0069] By convolution output With input The addition alleviates the gradient vanishing problem of deep networks and normalizes the fusion features layer by layer to accelerate model convergence.

[0070] Finally, to prevent overfitting of specific sentences, a random neuron inactivation Dropout with a probability of 10% is applied to the normalized features:

[0071] in, represents the first enhanced sentence-level feature, represents the probability of random dropout.

[0072] To address the semantic bias caused by unidirectional feature interaction, in an embodiment of the present application, the long text matching model further includes a word-sentence cross-fusion sub-model, to which the noise reduction unit and the sentence-level feature fusion unit are respectively connected; the word-sentence cross-fusion sub-model is connected to the matching sub-model; and the word-sentence cross-fusion sub-model employs bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced word-level features of each long text to obtain the word-sentence semantic association features of each long text. The long text matching model achieves complementary enhancement of local details and macro-semantics through bidirectional cross-attention of word-to-sentence retrieval and sentence-to-word constraints.

[0073] In an optional embodiment, the word-sentence cross-fusion sub-model includes a skip-connected interactive correction unit and a feature fusion unit; the word-sentence cross-fusion sub-model uses bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced sentence-level features of each long text to obtain word-level semantic association features of each long text, including: Through the interactive correction unit, the intermediate enhanced word-level features of each long text are mapped to the query space and the first enhanced sentence-level features of each long text are mapped to the key-value space to obtain the word-to-sentence mapping features of each long text. After the intermediate enhanced word-level features are corrected using the first enhanced sentence-level features of each long text to obtain the sentence-to-word mapping features of each long text, the word-to-sentence mapping features and the sentence-to-word mapping features of each long text are nonlinearly transformed to obtain the word-sentence mapping features of each long text; through the feature fusion unit, the first enhanced sentence-level features and the word-sentence mapping features of each long text are feature fused to obtain the word-sentence semantic association features of each long text.

[0074] like Figure 7As shown in the figure, the word-sentence cross-fusion sub-model uses bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced word-level features of each long text to obtain the word-sentence semantic association features of each long text. The specific process is as follows: Set word-to-sentence retrieval, use word-level features as query Q to retrieve sentence-level features, and use sentence-level features as query Q to retrieve word-level features, and add residual connection + nonlinear activation after bidirectional interaction; First, the intermediate enhanced word-level features and the first enhanced sentence-level features are processed through the word-sentence cross-attention and sentence-word cross-attention of the bidirectional attention fusion module to obtain the word-to-sentence mapping features and sentence-to-word mapping features of each long text:

[0075] Where, Represents word-to-sentence mapping features, is the enhanced word-level feature, that is, the intermediate enhanced word-level feature. Represents the projection matrix.

[0076] in, is the enhanced word-level feature, carrying local keyword information; the intermediate enhanced word-level feature can be mapped to the query space through the projection matrix, and the first enhanced sentence-level feature can be mapped to the key value space; the scaling factor Suppress the dot product value range to prevent Softmax gradient saturation; Based on sentence-to-word semantic constraints, the word-level representation is modified using sentence-level global features:

[0077] Where, Sentence-to-word mapping features; in, Decoupled from the forward retrieval parameters to avoid semantic confusion, the attention matrix dimension is ,in, is the number of sentences.

[0078] Then, based on the word-to-sentence mapping features and sentence-to-word mapping features of each long text, a nonlinear transformation is performed to obtain the word-to-sentence mapping features of each long text:

[0079] in, Represents the word-sentence mapping feature, Indicates feature compression, and GeLU indicates activation enhancement nonlinear relationship.

[0080] Finally, the first enhanced sentence-level features and word-sentence mapping features of each long text are fused to obtain the word-sentence semantic association features of each long text:

[0081] Where, represents the gating coefficient, represents the Sigmoid function, represents the gating weight matrix, Represents the semantic association features of words and sentences.

[0082] Among them, the gating coefficient pass Learning; Sigmoid function Constrain the gate value to , Implement feature-level soft selection and residual connection Preserve original sentence-level features to prevent information loss.

[0083] Through the bidirectional attention coordination of word-to-sentence fine-grained retrieval and sentence-to-word global constraints, a dynamic complementary mechanism between local details and macro-semantics is constructed; the gated residual network combined with depthwise separable convolution achieves dynamic weight distribution of cross-level features while maintaining parameter efficiency.

[0084] In order to preserve feature similarity, difference, and interaction pattern information in its long text matching model, in an embodiment of the present application, a matching sub-model is used to match each long text based on its word-level features, sentence-level features, and paragraph-level features to obtain a long text matching result. Specifically, a cascaded fully connected layer is used to perform feature integration and nonlinear mapping on the word-level features, sentence-level features, and paragraph-level features of each long text to obtain a long text matching result.

[0085] Where, Indicates the matching results. Representation layer normalization, represents a multilayer perceptron, represents the first intensifier level feature, and represents the first enhanced sentence-level feature, and Represents the semantic association features of words and sentences, and represents segment-level features, Represents the feature interaction operator.

[0086] The matching sub-model forms a 1280-dimensional joint feature vector by performing element-level Hadamard product, cosine similarity calculation and difference splicing on cross-document feature pairs.

[0087] The present application embodiment provides a long text matching device, see Figure 8As shown, a long text matching model device provided in an embodiment of the present application includes: The data acquisition module 810 is used to acquire a long text combination to be matched; wherein the long text combination includes at least two long texts; The text matching module 820 is used to input each long text into a long text matching model to obtain long text matching results between each long text; wherein, the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model and a paragraph-level feature sub-model connected in sequence, and a matching sub-model connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model respectively; the word-level feature sub-model of the long text matching model uses a first sliding window attention to perform local grammatical analysis on each long text to obtain the word-level features of each long text, uses a second sliding window attention to perform cross-sentence semantic association on the word-level features of each long text through the sentence-level feature sub-model to obtain the sentence-level features of each text, uses a third sliding window attention to perform global intention analysis on the sentence-level features of each long text through the paragraph-level feature sub-model to obtain the paragraph-level features of each long text, and uses the matching sub-model to match each long text based on the word-level features, sentence-level features and paragraph-level features of each long text to obtain the long text matching results; the semantic perception scopes of the first sliding window attention, the second sliding window attention and the third sliding window attention are in a progressive relationship.

[0088] In an optional embodiment, the long text matching model further includes a word-level feature enhancement sub-model and a sentence-level feature enhancement sub-model, the word-level feature sub-model is connected to the matching sub-model via the word-level feature enhancement sub-model, and the sentence-level feature sub-model is connected to the matching sub-model via the sentence-level feature enhancement sub-model; the long text matching method further includes: The word-level features of each long text are filtered for relational features through the word-level feature enhancer model to obtain the first enhanced word-level features of each long text; The sentence-level features of each long text are semantically reconstructed through the sentence-level feature enhancement sub-model to obtain the first enhanced sentence-level features; The matching sub-model is used to match each long text based on the first enhanced word-level features, the first enhanced sentence-level features and the paragraph-level features of each long text to obtain a long text matching result.

[0089] In an optional embodiment, the word-level feature enhancement sub-model includes a skip-connected denoising unit and a word-level feature fusion unit; the word-level feature enhancement sub-model is used to extract key features from the word-level features of each long text to obtain a first enhanced word-level feature of each long text, including: After calculating the attention scores of the word-level features of each long text through the denoising unit, based on the attention scores of the word-level features of each long text, the word-level features of each long text are screened for strong correlation features to obtain the intermediate enhanced word-level features of each long text; The word-level features and intermediate enhanced word-level features of each long text are subjected to multimodal feature splicing by a word-level feature fusion unit to obtain a spliced ​​word-level feature of each long text, and the intermediate enhanced word-level features of each long text are subjected to element-level feature fusion to obtain a fused word-level feature of each long text. Then, a first enhanced word-level feature of each long text is determined based on the spliced ​​word-level features and the fused word-level features of each long text; The sentence-level feature enhancer model includes a deep separable unit with skip connections and a sentence-level feature fusion unit. The sentence-level feature enhancer model reconstructs the semantic association of the sentence-level features of each long text to obtain the first enhanced sentence-level feature, including: The local temporal relationship features of the sentence-level features of each long text are extracted in different feature channels through the deep separable unit to obtain the local temporal relationship features of the different feature channels corresponding to each long text. Then, the local temporal relationship features of the different feature channels corresponding to each long text are fused across channels to obtain the fused sentence-level features of each long text. The sentence-level feature fusion unit is used to fuse the sentence-level features and fused sentence-level features of each long text to obtain the first enhanced sentence-level features of each long text.

[0090] In an optional embodiment, the attention score of the word-level features of each long text is calculated by the noise reduction unit, including: Calculate the initial attention score of the word-level features of each long text based on the attention score matrix; Based on the differentiable Gumbel distribution threshold gating, each initial attention score is filtered to obtain the attention score of the word-level features of each long text.

[0091] In an optional embodiment, a sentence-level feature fusion unit performs feature fusion on the sentence-level features and fused sentence-level features of each long text to obtain a first enhanced sentence-level feature of each long text, including: After the sentence-level features and fused sentence-level features of each long text are spliced ​​together to obtain the spliced ​​sentence-level features of each long text, the spliced ​​sentence-level features are normalized and regularized to obtain the first enhanced sentence-level features of each long text.

[0092] In an optional embodiment, the long text matching model further includes a word-sentence cross-fusion sub-model, the noise reduction unit and the sentence-level feature fusion unit are respectively connected to the word-sentence cross-fusion sub-model; the word-sentence cross-fusion sub-model is connected to the matching sub-model; and the long text matching method further includes: The word-sentence cross-fusion sub-model uses bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced word-level features of each long text to obtain the word-sentence semantic association features of each long text; The matching sub-model is used to match each long text based on the first enhanced word-level features, first enhanced sentence-level features, word-sentence semantic association features and paragraph-level features of each long text to obtain a long text matching result.

[0093] In an optional embodiment, the word-sentence cross-fusion sub-model includes a skip-connected interactive correction unit and a feature fusion unit; the word-sentence cross-fusion sub-model uses bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced sentence-level features of each long text to obtain word-level semantic association features of each long text, including: Mapping the intermediate enhanced word-level features of each long text to the query space and mapping the first enhanced sentence-level features of each long text to the key-value space through the interactive correction unit to obtain the word-to-sentence mapping features of each long text, and using the first enhanced sentence-level features of each long text to correct the intermediate enhanced word-level features to obtain the sentence-to-word mapping features of each long text, and then performing nonlinear transformation on the word-to-sentence mapping features and the sentence-to-word mapping features of each long text to obtain the word-to-sentence mapping features of each long text; The first enhanced sentence-level features and word-sentence mapping features of each long text are fused through a feature fusion unit to obtain the word-sentence semantic association features of each long text.

[0094] In an optional embodiment, the word-level feature sub-model uses the first sliding window attention to perform local grammatical parsing on each long text to obtain the word-level features of each long text, including: By performing word segmentation on each long text, the word element features of each long text are obtained; Through the first sliding window and local-to-global collaborative attention, each long text is locally parsed to obtain the word-level features of each long text.

[0095] In an optional embodiment, sentence-level features of each long text are obtained by performing cross-sentence semantic association on the word-level features of each long text using the second sliding window attention through the sentence-level feature sub-model, including: The word-level features of each long text are aggregated through mean pooling to obtain the sentence-level vector of each long text; The sentence-level features of each long text are obtained by performing cross-sentence semantic association on the sentence-level vectors of each long text through the second sliding window and cross-document attention.

[0096] In an optional embodiment, the segment-level feature sub-model uses the third sliding window attention to perform global intent analysis on the sentence-level features of each long text to obtain the segment-level features of each long text, including: Extract the sentence-level features of each long text through maximum pooling to obtain the initial sentence-level features of each long text; The initial sentence-level features of each long text are analyzed globally through the third sliding window and sparse global attention to obtain the paragraph-level features of each long text.

[0097] In an optional embodiment, each long text is matched based on the word-level features, sentence-level features, and paragraph-level features of each long text by a matching sub-model to obtain a long text matching result, including: The first enhanced word-level features, first enhanced sentence-level features, word-sentence semantic association features, and paragraph-level features of each long text are integrated and nonlinearly mapped through the cascaded fully connected layer to obtain the long text matching results.

[0098] The device provided in the embodiment of the present application has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.

[0099] like Figure 9 As shown, an electronic device 900 provided in an embodiment of the present application includes: a processor 910, a memory 920 and a bus, the memory 920 stores machine-readable instructions executable by the processor 910, and when the electronic device is running, the processor 910 communicates with the memory 920 through the bus, and the processor 910 executes the machine-readable instructions to perform the steps of the long text matching method as described above.

[0100] Specifically, the memory 920 and the processor 910 can be general-purpose memories and processors, which are not specifically limited here. When the processor 910 runs the computer program stored in the memory 920, the long text matching method can be executed.

[0101] The processor 910 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 910 or by instructions in the form of software. The above-mentioned processor 910 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 920, and processor 910 reads information in memory 920 and, in conjunction with its hardware, completes the steps of the above method.

[0102] Corresponding to the above-mentioned long text matching method, an embodiment of the present application also provides a computer-readable storage medium, which stores machine-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to execute the steps of the above-mentioned long text matching method.

[0103] The long text matching device provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in the embodiment of the present application are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.

[0104] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0105] For another example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of code, and a part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0106] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0107] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0108] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0109] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0110] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer" and the like indicate positions or locations based on the positions shown in the accompanying drawings, or the positions or locations in which the inventive product is typically placed when in use. These terms are intended solely to facilitate the description of the present invention and to simplify the description, and are not intended to indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third," etc., are used solely to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0111] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0112] Finally, it should be noted that the above embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. However, such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. They should all be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A long text matching method, characterized in that: include: Obtaining a long text combination to be matched; wherein the long text combination includes at least two long texts; Input each of the long texts into a long text matching model to obtain long text matching results between each of the long texts; wherein, the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model and a paragraph-level feature sub-model connected in sequence, and a matching sub-model connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model respectively; the long text matching model uses the first sliding window attention to perform local grammatical analysis on each of the long texts through the word-level feature sub-model to obtain the word-level features of each of the long texts, and uses the second sliding window attention to perform local grammatical analysis on each of the long texts through the sentence-level feature sub-model The word-level features of each of the long texts are semantically associated across sentences to obtain sentence-level features of each of the texts. The sentence-level features of each of the long texts are globally analyzed by the third sliding window attention through the paragraph-level feature sub-model to obtain paragraph-level features of each of the long texts. The long texts are matched based on the word-level features, the sentence-level features and the paragraph-level features of each of the long texts through the matching sub-model to obtain the long text matching results. The semantic perception scopes of the first sliding window attention, the second sliding window attention and the third sliding window attention are in a progressive relationship.

2. The long text matching method according to claim 1, characterized in that: The long text matching model further includes a word-level feature enhancement sub-model and a sentence-level feature enhancement sub-model, wherein the word-level feature sub-model is connected to the matching sub-model via the word-level feature enhancement sub-model, and the sentence-level feature sub-model is connected to the matching sub-model via the sentence-level feature enhancement sub-model; The long text matching method further includes: Performing relational feature screening on the word-level features of each of the long texts using the word-level feature enhancer model to obtain first enhanced word-level features of each of the long texts; Reconstructing the semantic association of sentence-level features of each of the long texts through the sentence-level feature enhancement sub-model to obtain a first enhanced sentence-level feature; The long text matching result is obtained by matching each of the long texts based on the first enhanced word-level feature, the first enhanced sentence-level feature, and the paragraph-level feature of each of the long texts through the matching sub-model.

3. The long text matching method according to claim 2, characterized in that: The word-level feature enhancement sub-model includes a skip-connected denoising unit and a word-level feature fusion unit; the word-level feature enhancement sub-model is used to extract key features from the word-level features of each of the long texts to obtain a first enhanced word-level feature of each of the long texts, including: After calculating the attention scores of the word-level features of each of the long texts by the noise reduction unit, based on the attention scores of the word-level features of each of the long texts, the word-level features of each of the long texts are subjected to strong correlation feature screening to obtain intermediate enhanced word-level features of each of the long texts; The word-level feature fusion unit performs multimodal feature splicing on the word-level features of each of the long texts and the intermediate enhancement word-level features to obtain a spliced ​​word-level feature of each of the long texts, and performs element-level feature fusion on the intermediate enhancement word-level features of each of the long texts to obtain a fused word-level feature of each of the long texts, and then determines a first enhancement word-level feature of each of the long texts based on the spliced ​​word-level features and the fused word-level features of each of the long texts; The sentence-level feature enhancement sub-model includes a jump-connected deep separable unit and a sentence-level feature fusion unit; the sentence-level features of each of the long texts are semantically reconstructed by the sentence-level feature enhancement sub-model to obtain a first enhanced sentence-level feature, including: After extracting local temporal relationship features in different feature channels from the sentence-level features of each of the long texts through the depthwise separable unit to obtain local temporal relationship features of different feature channels corresponding to each of the long texts, cross-channel feature fusion is performed on the local temporal relationship features of different feature channels corresponding to each of the long texts to obtain fused sentence-level features of each of the long texts; The sentence-level feature fusion unit performs feature fusion on the sentence-level features of each of the long texts and the fused sentence-level features to obtain first enhanced sentence-level features of each of the long texts.

4. The long text matching method according to claim 3, characterized in that: The sentence-level feature fusion unit performs feature fusion on the sentence-level features and the fused sentence-level features of each of the long texts to obtain a first enhanced sentence-level feature of each of the long texts, including: After splicing the sentence-level features and the fused sentence-level features of each of the long texts to obtain spliced ​​sentence-level features of each of the long texts, the spliced ​​sentence-level features are normalized and regularized to obtain first enhanced sentence-level features of each of the long texts.

5. The long text matching method according to claim 3, characterized in that: The long text matching model further includes a word-sentence cross-fusion sub-model, the noise reduction unit and the sentence-level feature fusion unit are respectively connected to the word-sentence cross-fusion sub-model; the word-sentence cross-fusion sub-model is connected to the matching sub-model; The long text matching method further includes: The word-sentence cross-fusion sub-model adopts bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced word-level features of each of the long texts to obtain the word-sentence semantic association features of each of the long texts; The long text matching result is obtained by matching each of the long texts through the matching sub-model based on the first enhanced word-level feature, the first enhanced sentence-level feature, the word and sentence semantic association feature, and the paragraph-level feature of each of the long texts.

6. The long text matching method according to claim 5, characterized in that: The word-sentence cross-fusion sub-model includes a skip-connected interactive correction unit and a feature fusion unit; the word-sentence cross-fusion sub-model uses bidirectional heterogeneous attention to perform word-sentence cross-correction on the intermediate enhanced word-level features and the first enhanced sentence-level features of each of the long texts to obtain word-level semantic association features of each of the long texts, including: Mapping the intermediate enhanced word-level features of each of the long texts to a query space and mapping the first enhanced sentence-level features of each of the long texts to a key-value space by the interactive correction unit to obtain a word-to-sentence mapping feature of each of the long texts, and using the first enhanced sentence-level features of each of the long texts to correct the intermediate enhanced word-level features to obtain a sentence-to-word mapping feature of each of the long texts, then performing a nonlinear transformation on the word-to-sentence mapping feature and the sentence-to-word mapping feature of each of the long texts to obtain a word-to-sentence mapping feature of each of the long texts; The feature fusion unit fuses the first enhanced sentence-level features of each of the long texts with the word-sentence mapping features to obtain word-sentence semantic association features of each of the long texts.

7. The long text matching method according to any one of claims 1 to 6, characterized in that: The word-level feature sub-model uses the first sliding window attention to perform local grammatical parsing on each of the long texts to obtain word-level features of each of the long texts, including: By performing word segmentation processing on each of the long texts, word element features of each of the long texts are obtained; Each of the long texts is locally parsed using the first sliding window and local-to-global collaborative attention to obtain word-level features of each of the long texts.

8. The long text matching method according to any one of claims 1 to 6, characterized in that: The sentence-level features of each of the long texts are obtained by performing cross-sentence semantic association on the word-level features of each of the long texts using the second sliding window attention through the sentence-level feature sub-model, including: Aggregating the word-level features of each of the long texts through mean pooling to obtain a sentence-level vector for each of the long texts; The sentence-level features of each of the long texts are obtained by performing cross-sentence semantic association on the sentence-level vectors of each of the long texts through the second sliding window and cross-document attention.

9. The long text matching method according to any one of claims 1 to 6, characterized in that: The segment-level feature sub-model uses the third sliding window attention to perform global intent analysis on the sentence-level features of each of the long texts to obtain the segment-level features of each of the long texts, including: Extracting sentence-level features of each of the long texts through maximum pooling to obtain initial sentence-level features of each of the long texts; The initial sentence-level features of each of the long texts are subjected to global intent analysis through the third sliding window and sparse global attention to obtain the paragraph-level features of each of the long texts.

10. A long text matching device, characterized in that: include: A data acquisition module is used to acquire a long text combination to be matched; wherein the long text combination includes at least two long texts; The text matching module is used to input each of the long texts into a long text matching model to obtain long text matching results between each of the long texts; wherein, the long text matching model includes a word-level feature sub-model, a sentence-level feature sub-model and a paragraph-level feature sub-model connected in sequence, and a matching sub-model respectively connected to the word-level feature sub-model, the sentence-level feature sub-model and the paragraph-level feature sub-model; the word-level feature sub-model of the long text matching model uses a first sliding window attention to perform local grammatical analysis on each of the long texts to obtain word-level features of each of the long texts, and uses a second sliding window attention to perform local grammatical analysis on each of the long texts through the sentence-level feature sub-model Attention performs cross-sentence semantic association on the word-level features of each of the long texts to obtain sentence-level features of each of the texts, and uses the third sliding window attention through the paragraph-level feature sub-model to perform global intent analysis on the sentence-level features of each of the long texts to obtain paragraph-level features of each of the long texts, and uses the matching sub-model to match each of the long texts based on the word-level features, the sentence-level features and the paragraph-level features of each of the long texts to obtain the long text matching results; the semantic perception scopes of the first sliding window attention, the second sliding window attention and the third sliding window attention are in a progressive relationship.