Text processing method and device, computer equipment, storage medium and program product

By mapping the position encoding bias information of the target text to the position encoding range of the text processing model, the problem of insufficient generalization performance of the traditional position encoding method in long text processing is solved, and better text processing effect is achieved.

CN120045646APending Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311536806.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When traditional position coding processes long text that has not been seen, the generalization performance is insufficient, resulting in poor inference effect of the model.

Method used

By obtaining the position encoding bias information of the target text, the relative position encoding is mapped to the position encoding range of the text processing model, and then semantic understanding is carried out to generate semantic associated text.

Benefits of technology

It improves the generalization performance of the text processing model, especially when processing long text, can maintain good text processing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045646A_ABST
    Figure CN120045646A_ABST
Patent Text Reader

Abstract

The invention provides a text processing method and device, computer equipment, a storage medium and a program product, and relates to the field of natural language processing of an artificial intelligence technology. The text processing method comprises the steps of obtaining position coding bias information corresponding to a target text, wherein the position coding bias information comprises relative position codes between every two text units in the target text; calling a text processing model to map a first relative position code in the position code bias information into a position code range corresponding to the text processing model to obtain a target relative position code corresponding to the first relative position code; the first relative position code is a relative position code of which the corresponding relative position exceeds a position threshold value in the position code bias information; and calling a text processing model to perform semantic understanding on the target text according to the mapped position coding bias information, and generating a semantic association text of the target text. By adopting the method, the generalization performance of the text processing model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, particularly to the field of artificial intelligence technologies, and specifically to a text processing method, a text processing device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] The Transformer model structure (a language model structure) based on the Attention mechanism has been widely applied in the field of natural language processing of artificial intelligence and has become a relatively popular language model structure at present. However, a pure Attention module cannot capture the order of each text unit (token) in the input text, which makes the role of positional encoding in the Transformer model particularly important. Positional encoding is a way of positional representation and can represent the position of each text unit in the input text through encoding.

[0003] Currently, when the model processes text using traditional positional encoding methods, for texts with some seen text lengths (the seen text length refers to the text lengths that have appeared during the model training process), the text lengths of these texts are usually short and can achieve good text processing effects; but for texts with some unseen text lengths (the unseen text length refers to those that have not appeared during the model training stage and exceed the text lengths that have appeared during the training process), the text lengths of these texts are usually long and the text processing effects are poor. That is to say, the generalization performance of traditional positional encoding methods is insufficient, which means that when the model processes inputs with unseen text lengths exceeding those during the training process, the extrapolation ability may be insufficient, resulting in poor inference effects of the model. Summary of the Invention

[0004] Embodiments of this application provide a text processing method, device, computer device, storage medium, and program product, which can improve the generalization performance of the text processing model.

[0005] On the one hand, embodiments of this application provide a text processing method, which includes:

[0006] Obtain a target text to be processed, where the target text includes multiple text units;

[0007] Obtain position encoding bias information corresponding to the target text, where the position encoding bias information includes relative position encodings between every two text units in the target text, and the relative position encodings between every two text units are determined according to the relative positions of the every two text units in the target text;

[0008] Call the text processing model to map the first relative position encoding in the position encoding bias information to the corresponding position encoding range of the text processing model, and obtain the target relative position encoding corresponding to the first relative position encoding; the first relative position encoding is the relative position encoding in the position encoding bias information where the corresponding relative position exceeds the position threshold;

[0009] Call the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information, and generate the semantic associated text of the target text.

[0010] Correspondingly, an embodiment of the present application provides a text processing device, and the text processing device includes:

[0011] An acquisition unit, configured to acquire a target text to be processed, where the target text includes a plurality of text units;

[0012] The acquisition unit is further configured to acquire the position encoding bias information corresponding to the target text, where the position encoding bias information includes the relative position encoding between two text units in the target text, and the relative position encoding between two text units is determined according to the relative position of the two text units in the target text;

[0013] A processing unit, configured to call the text processing model to map the first relative position encoding in the position encoding bias information to the corresponding position encoding range of the text processing model, and obtain the target relative position encoding corresponding to the first relative position encoding; the first relative position encoding is the relative position encoding in the position encoding bias information where the corresponding relative position exceeds the position threshold;

[0014] The processing unit is further configured to call the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information, and generate the semantic associated text of the target text.

[0015] In an implementation manner, the processing unit is further configured to perform the following steps:

[0016] Call the text processing model to determine the second relative position encoding in the position encoding bias information as the target relative position encoding corresponding to the second relative position encoding; the second relative position encoding is the relative position encoding in the position encoding bias information where the corresponding relative position does not exceed the position threshold;

[0017] Wherein, the mapped position encoding bias information includes the target relative position encoding corresponding to the first relative position encoding and the target relative position encoding corresponding to the second relative position encoding.

[0018] In one implementation, the target text includes N text units, and any one of the N text units is represented as the i-th text unit. The i-th text unit is arranged at the i-th position in the target text. N is an integer greater than 1, and i is an integer less than or equal to N. The positional encoding bias information includes the relative positional encoding between the i-th text unit and each of its previous text units. Any previous text unit refers to any text unit arranged at the 1st position - the i-th position in the target text;

[0019] The processing unit is used to call the text processing model to map the first relative positional encoding in the positional encoding bias information to the corresponding positional encoding range of the text processing model to obtain the target relative positional encoding corresponding to the first relative positional encoding. Specifically, it is used to perform the following steps:

[0020] Call the text processing model to map the first relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units to the corresponding positional encoding range of the text processing model to obtain the target relative positional encoding corresponding to the first relative positional encoding;

[0021] Wherein, the first relative positional encoding is the relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units, corresponding to the relative position exceeding the position threshold.

[0022] In one implementation, the processing unit is used to call the text processing model to determine the second relative positional encoding in the positional encoding bias information as the target relative positional encoding corresponding to the second relative positional encoding. Specifically, it is used to perform the following steps:

[0023] Call the text processing model to determine the second relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units as the target relative positional encoding corresponding to the second relative positional encoding;

[0024] Wherein, the second relative positional encoding is the relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units, corresponding to the relative position not exceeding the position threshold.

[0025] In one implementation, the relative position between the i-th text unit and the j-th previous text unit exceeds the position threshold, where j is a positive integer less than or equal to i. The processing unit is used to call the text processing model to map the first relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units to the corresponding positional encoding range of the text processing model to obtain the target relative positional encoding corresponding to the first relative positional encoding. Specifically, it is used to perform the following steps:

[0026] Obtain an interpolation function; the interpolation function is generated based on the first text length, the second text length, and a position threshold. The first text length refers to the maximum text length processed by the text processing model during training; the second text length refers to the text length of the target text, and the text length of the target text refers to the number of text units included in the target text; the value range of the interpolation function is the position encoding range corresponding to the text processing model.

[0027] According to the interpolation function, perform interpolation processing on the first relative position encoding between the i-th text unit and the j-th previous text unit to obtain the target relative position encoding corresponding to the first relative position encoding between the i-th text unit and the j-th previous text unit.

[0028] In one implementation, the mapped position encoding bias information includes the target relative position encoding between the i-th text unit and each previous text unit; the processing unit is used to call the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information and generate a semantic associated text of the target text. Specifically, it is used to perform the following steps:

[0029] Determine the first attention weight between the i-th text unit and each previous text unit according to the similarity between the i-th text unit and each previous text unit.

[0030] Update the first attention weight between the i-th text unit and each previous text unit according to the target relative position encoding between the i-th text unit and each previous text unit to obtain the second attention weight between the i-th text unit and each previous text unit.

[0031] Perform weighted summation processing on each previous text unit according to the second attention weight between the i-th text unit and each previous text unit to obtain the semantic feature corresponding to the i-th text unit.

[0032] Perform semantic understanding on the target text according to the semantic features of each text unit in the target text to generate a semantic associated text of the target text.

[0033] In one implementation, the way for the processing unit to obtain the position threshold includes any one of the following:

[0034] Set the position threshold according to the first text length, where the first text length is the maximum text length processed by the text processing model during training.

[0035] Use the initial threshold as a model parameter of the text processing model and adjust it during the fine-tuning process of the text processing model to obtain the position threshold.

[0036] In one implementation, the processing unit is configured to use the initial threshold as a model parameter of the text processing model and adjust it during the fine-tuning process of the text processing model. When obtaining the position threshold, it is specifically configured to perform the following steps:

[0037] Obtain a sample text set, where the sample text set includes multiple sample texts, and the text length of each sample text is greater than the maximum text length processed by the text processing model during training;

[0038] Iteratively fine-tune the text processing model according to the sample text set and adjust the model parameters of the text processing model;

[0039] When the iterative fine-tuning termination condition is reached, determine the model parameters in the text processing model when the iterative fine-tuning termination condition is reached as the position threshold.

[0040] In one implementation, the processing unit is configured to call the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information and generate a semantic associated text of the target text. It is specifically configured to perform the following steps:

[0041] Call the text processing model to perform the first semantic understanding on the target text according to the mapped position encoding bias information and generate a semantic associated text unit corresponding to the first semantic understanding;

[0042] Use the text composed of the semantic associated text unit corresponding to the first semantic understanding and the target text as the new target text, and continue to call the text processing model for recursive semantic understanding until the recursive termination condition is reached;

[0043] After reaching the recursive termination condition, determine the text composed of the semantic associated text unit corresponding to the first semantic understanding and each semantic associated text unit obtained by recursive semantic understanding as the semantic associated text of the target text;

[0044] Among them, reaching the recursive termination condition includes any one of the following: the predicted semantic associated text unit is a specified text unit, and the number of predicted semantic associated text units reaches a specified number.

[0045] In one implementation, the obtaining unit is configured to obtain the position encoding bias information corresponding to the target text. It is specifically configured to perform the following steps:

[0046] Obtain the relative positions of pairwise text units in the target text;

[0047] Calculate the relative positions of pairwise text units in the target text according to the relative position coefficients to obtain the relative position encoding between pairwise text units.

[0048] In one implementation, the text processing model includes multiple attention fusion modules, and the relative position encoding between every two text units in the target text is applied to the multiple attention fusion modules; the coefficients corresponding to the multiple attention fusion modules are the same or different;

[0049] If each attention fusion module includes multiple attention fusion processes, the coefficients corresponding to each attention fusion process are different.

[0050] Correspondingly, an embodiment of the present application provides a computer device, which includes:

[0051] A processor, adapted to implement a computer program;

[0052] A computer-readable storage medium storing a computer program, the computer program being adapted to be loaded and executed by the processor to perform the above-mentioned text processing method.

[0053] Correspondingly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when read and executed by a processor of a computer device, causes the computer device to perform the above-mentioned text processing method.

[0054] Correspondingly, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the above-mentioned text processing method.

[0055] In the embodiment of the present application, when the text length of the target text exceeds the text length that appears during the training process of the text processing model, the position encoding range to which the relative position encoding between every two text units in the target text belongs will exceed the position encoding range corresponding to the text processing model, which will result in poor text processing effects of the text processing model. In this case, the embodiment of the present application maps the relative position encoding with a relative position exceeding the position threshold to the position encoding range corresponding to the text processing model, so that when the text processing model performs text processing on the target text based on the mapped relative position encoding, it can achieve better text processing effects; that is to say, in the embodiment of the present application, when performing text processing on a target text with a text length exceeding the training text length (i.e., the text length that appears during the training process of the text processing model), by mapping the relative position encoding between every two text units in the target text to the position encoding range corresponding to the text processing model, the text processing model can achieve better text processing effects for the target text with a text length exceeding the training text length, thereby improving the generalization performance of the text processing model. Description of the Drawings

[0056] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a schematic diagram for calculating attention weights provided by an embodiment of the present application;

[0058] Figure 2 It is a schematic diagram of a long text understanding scenario provided by an embodiment of the present application;

[0059] Figure 3 It is a schematic diagram of a long text generation scenario provided by an embodiment of the present application;

[0060] Figure 4 It is a schematic diagram of a multi-round dialogue scenario provided by an embodiment of the present application;

[0061] Figure 5 It is a schematic flowchart of a text processing method provided by an embodiment of the present application;

[0062] Figure 6 It is a schematic structural diagram of a text processing model provided by an embodiment of the present application;

[0063] Figure 7 It is a schematic flowchart of a recursive semantic understanding provided by an embodiment of the present application;

[0064] Figure 8 It is a schematic flowchart of another text processing method provided by an embodiment of the present application;

[0065] Figure 9 It is a comparative schematic diagram of a relative position encoding method provided by an embodiment of the present application;

[0066] Figure 10 It is a schematic diagram of the role of relative position encoding in semantic understanding provided by an embodiment of the present application;

[0067] Figure 11 It is a schematic structural diagram of a text processing device provided by an embodiment of the present application;

[0068] Figure 12 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0070] In order to be able to more clearly understand the technical solutions provided by the embodiments of the present application, some key terms involved in the embodiments of the present application will be introduced here:

[0071] (1) Artificial Intelligence:

[0072] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0073] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0074] (2) Natural Language Processing:

[0075] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing involves natural language, which is the language people use in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, evolved from the large language model in the field of NLP. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0076] (3) Text processing:

[0077] Text processing, which can also be understood as text generation, refers to the process of calling a text processing model to perform semantic understanding on the text and generating semantically related text based on the understood semantics in the embodiments of the present application. The text processing process in the embodiments of the present application can generally include: First, a text processing model can be called to perform word segmentation on the text to obtain multiple text units (tokens) included in the text. A text unit refers to a language character (a language character refers to a character in natural language) or a language word (a language word refers to a word in natural language) in the text. Second, a text processing model can be called to use the attention mechanism to determine the attention weights between pairwise text units in the text based on the similarity between pairwise text units in the text; a text processing model can be called to perform weighted summation processing on the corresponding text units in the text according to the attention weights between pairwise text units in the text to obtain the attention features of each text unit in the text. The attention features of text units in the text can be understood as the semantic features of text units and can be used to represent the semantics of text units. Then, a text processing model can be called to generate semantically related text based on the semantic features of each text unit in the text. The text processing model can be a large language model in the field of natural language processing of artificial intelligence technology. A large language model refers to a deep learning model trained with a large amount of text data and can generate natural language text or understand the meaning of natural language text. For example, the text processing model can be a large prediction model based on the Transformer model structure.

[0078] It can be seen that the attention mechanism plays an important role in the text processing process. The attention mechanism is a network structure for sequence learning and has been commonly used in recent years to model language sequences in text processing for tasks such as text understanding and text generation. A common implementation of the attention mechanism can be seen in the following description:

[0079] Suppose text X is a text with a length of N, where the text length refers to the number of text units contained in the text. That is to say, text X contains N text units, and text X is a sequence input of text units (i.e., language words or tokens) with a length of N. Performing a linear mapping on text X can obtain a query matrix (Q (query) matrix), a matching matrix (K (key) matrix), and a content matrix (V (value) matrix). The attention weight matrix can be expressed as the following formula 1:

[0080]

[0081] In the above formula 1, AttSim(Q, K) represents the attention weight matrix; Q and K are two projection matrix operations, Q represents the query matrix, and K represents the matching matrix; d is a scaling factor used to maintain the stability of the model training of the text processing model. This attention weight matrix can be used to perform a weighted sum on the content matrix V to output the attention feature Z (i.e., the semantic feature Z). Specifically, please refer to the following formula 2:

[0082] Z(X) = Attsim(Q(X), K(X)) * V(X) T Formula 2

[0083] In the above formula 2, V is a projection matrix operation, and V represents the content matrix.

[0084] Specifically, during the text processing process, natural language words or tokens are generated word by word from left to right. Therefore, during the calculation of the attention feature, each text unit should only calculate the similarity between this text unit and the previous text unit and perform a weighted sum; the previous text unit of the current text unit can include the current text unit itself and the text units arranged before the current text unit in the text; generally speaking, this process can be introduced through an attention mask. In addition, the attention mechanism cannot capture the order of each text unit (token) in the text, and the order of each text unit in the text helps to understand the context semantics of the text unit. Therefore, it is necessary to introduce a relative position encoding between two text units to represent the relative position between two text units. The attention mask and relative position encoding in the attention mechanism are introduced separately below:

[0085] ① Attention mask:

[0086] For a text with a length of N, the attention mask of the text can be set as Mask ∈ R N*N, its shape is N*N, which is aligned with the shape of the attention weight matrix in the above formula 1. The attention mask defines the values of its elements as the following formula 3:

[0087]

[0088] In the above formula 3, i represents the i-th text unit in the text, and the i-th text unit is arranged at the i-th position in the text; j represents the j-th text unit in the text, and the j-th text unit is arranged at the j-th position in the text; mask i,j represents the value of the element between the i-th text unit and the j-th text unit in the attention mask. This attention mask acts on the attention weight matrix and can achieve the function of "only calculating the similarity between each text unit and its previous text unit and weighted summing".

[0089] Therefore, the above formula 1 can be rewritten as the following formula 4:

[0090]

[0091] In the above formula 4, AttSim_withMask(Q,K) represents the attention weight matrix updated based on the attention mask, and Mask represents the attention mask.

[0092] ② Relative position encoding:

[0093] Relative position encoding (taking the Alibi position encoding method as an example) can map the relative position relationship between two text units in the text into a bias term in the shape of an attention map (attention weight matrix), which acts on the attention map. Alibi relative position encoding introduces a bias term to represent the relative position relationship between two text units in the text. The original Alibi relative position encoding and its action method in the attention weight are as follows formula 5:

[0094]

[0095] In the above formula 5, represents the attention weight between the i-th text unit and each previous text unit of the i-th text unit; Mask iDenote the value of the corresponding element of the $i$-th text unit in the attention mask; $m*[-(i - 1),\ldots,-2,-1,0]$ represents the relative position encoding between the $i$-th text unit and each of its previous text units. For example, $m*[-(i - 1)]$ represents the relative position encoding between the $i$-th text unit and the 1st text unit; $m$ is the relative position coefficient, which is a fixed constant coefficient.

[0096] In summary, the attention mask and relative position encoding can be introduced into the original attention weights to update the attention weights, realizing the function of calculating the similarity between each text unit in the text and its previous text unit, and considering the order of each text unit in the text during the calculation process of the attention mechanism to improve the semantic understanding ability of the attention mechanism. Taking a text with a length of 6 as an example, the influence of the attention mask and relative position encoding on the original attention weights can be seen in Figure 1 , in Figure 1 's left matrix, $q$ 6 $k$ 1 represents the attention weight between the 6th text unit and the 1st text unit in the text, $q$ 4 $k$ 3 represents the attention weight between the 4th text unit and the 3rd text unit in the text, and so on. $q$ 1 $k$ 1 represents the attention weight between the 1st text unit and the 1st text unit in the text; in Figure 1 's left matrix, the elements in the matrix are not complete, which is the influence of the attention mask. The attention mask ensures that the attention mechanism focuses on the attention weights between each text unit and its previous text unit, and does not focus on the attention weights between each text unit and its subsequent text unit. The subsequent text unit of any text unit refers to the text unit arranged after this text unit in the text. Figure 1 The right matrix of

[0097] Based on the above introduction of key terms such as text processing, attention mechanism, and relative position encoding, it can be found that through relative position encoding, the order between each text unit in the text can be introduced into the attention mechanism, enabling the attention mechanism to capture the order between each text unit in the text. During the text processing process, improving the generalization performance of the text processing model and enhancing the length extrapolation ability of the text processing model are urgent requirements for text processing; the length extrapolation ability can be understood as the specific manifestation of the generalization performance in terms of text length. Improving the generalization performance of the text processing model and enhancing the length extrapolation ability of the text processing model are similar concepts, which refer to improving the text processing effect of the text processing model on long sequences (long sequences refer to texts with longer text lengths) for a text processing model trained on short sequences (short sequences refer to texts with shorter text lengths), so that the text processing model can also achieve good performance on long sequences.

[0098] Length extrapolation means that the text length processed by the text processing model during the inference process exceeds the text length processed by the text processing model during the training process. After length extrapolation, the relative position encoding between each pair of text units in the text will also be extrapolated accordingly. In the Alibi relative position encoding method, the relative position encoding between each pair of text units in the text can be expressed as the following formula 6:

[0099] AlibiBias(i,j) = -m*(i - j) Formula 6

[0100] In the above formula 6, AlibiBias(i,j) represents the relative position encoding between the i-th text unit and the j-th text unit in the text. i represents the arrangement position of the i-th text unit in the text as the i-th position, j represents the arrangement position of the j-th text unit in the text as the j-th position, and i ≥ j.

[0101] Assume that the longest sequence length (the longest sequence length refers to the longest text length) processed by the text processing model during the training process is l 1 , and the longest sequence length required by the actual business needs needs to be extrapolated to l 2 (l 2 > l 1 ). The embodiments of the present application provide two extrapolation methods for relative position encoding:

[0102] The first extrapolation method for relative position encoding is the direct extrapolation method. The direct extrapolation method means directly calculating the relative position encoding of the text after length extrapolation based on the above formula 6, that is, directly calculating the relative position encoding of the text after length extrapolation based on the Alibi relative position encoding method. In this method, the value range of formula 6 is [-m*(l 1 - 1), 0]. If we hope to extrapolate the longest sequence length to l during inference2 , then based on the above formula 6, we can directly calculate the relative position encoding of the text after length extrapolation. The value range of the bias term of the position encoding after extrapolation is [-m*(l 2 -1), 0].

[0103] The second extrapolation method of the relative position encoding is the interpolation extrapolation method. The interpolation extrapolation method means to retain the value range of [-m*(l 1 -1), 0], and correspondingly change the interval between each relative position encoding, and interpolate each relative position encoding geometrically into the value range [-m*(l 1 -1), 0]. The changes in the second extrapolation method of the relative position encoding can be specifically reflected in the relative position coefficient m. Please refer to the following formula 7:

[0104]

[0105] That is to say, the second extrapolation method of the relative position encoding modifies the original relative position coefficient m into a relative position coefficient with geometric interpolation function

[0106] Neither of the above two extrapolation methods of the relative position encoding has achieved good results in improving the generalization performance of the text processing model and enhancing the extrapolation ability of the text processing model. The reasons are as follows: The direct extrapolation method will cause a large change in the value range of the relative position encoding, and the generalization ability of the model to the text units in the extrapolated part is insufficient; the interpolation extrapolation method will cause the interval between the relative position encodings to be shortened, and the distribution of the attention weights between the text units with adjacent relative positions will be greatly affected; these two extrapolation methods of the relative position encoding only consider the sequence length during the training process and directly extrapolate during inference, without considering the situation where the text processing model can be further fine-tuned.

[0107] Based on this, the embodiments of the present application propose a text processing method. Based on the above two extrapolation methods of relative position encoding, an innovative extrapolation method of relative position encoding is proposed. This innovative extrapolation method of relative position encoding is a new type of functional position encoding method, which is an Alibi relative position encoding method based on piecewise interpolation and can be named LeakyAlibi relative position encoding method. The relative position encoding under this method has good extrapolation performance. Specifically, different from the previous relative position encoding, the text processing method proposed in the embodiments of the present application comprehensively considers the performance of relative position encoding at unseen positions and its impact on the local attention mechanism; the new functional position encoding method is a new piecewise position encoding bias term function. For two text units with close relative positions (i.e., similar tokens), the relative position encoding interval between the two text units remains unchanged, retaining the distribution characteristics of the text processing model in local attention; for two text units with far relative positions (i.e., distal tokens), the relative position encoding implements interpolation to control the value range of the position bias term, effectively improving the generalization ability of the text processing model when the relative position encoding is extrapolated. In addition, the text processing method proposed in the embodiments of the present application also introduces a parameter (which can be called the position threshold) to distinguish whether the relative positions between pairwise text units are close. Through this parameter, it can also be simply and effectively extended to the fine-tuning scenario. During fine-tuning, this parameter can be set as a learnable model parameter, which can improve the scalability of the text processing model.

[0108] It should be noted that the text processing method proposed in the embodiments of the present application can be integrated into a text processing model. The text processing method proposed in the embodiments of the present application can be executed by a computer device, and the computer device can be a terminal device or a server deployed with a text processing model. Among them, the terminal devices mentioned in the embodiments of the present application may include, but are not limited to, any one of the following: smart phones, tablets, laptop computers, desktop computers, smart watches, smart home appliances, smart voice interaction devices, vehicle-mounted terminals, and aircraft, but are not limited thereto; the servers mentioned in the embodiments of the present application can be a single physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of the present application do not limit this.

[0109] The embodiments of the present application do not limit the application scenarios of the text processing method. The text processing method can be applied to any application scenario that requires text semantic understanding of long texts. For example, the text processing method proposed in the embodiments of the present application can be applied to text processing scenarios such as long text understanding, long text generation, and multi-turn conversations. Specifically:

[0110] The long text understanding scenario refers to a text processing scenario in which semantic understanding is performed on the text to be processed to generate a semantic understanding result of the text to be processed. That is to say, the generated semantic understanding result is a semantically related text. For example, the semantic understanding result can be a text summary generated by performing semantic understanding on the text to be processed. Another example is that the semantic understanding result can be keywords extracted from the text to be processed after performing semantic understanding on the text to be processed. Still another example is that the semantic understanding result can be an information retrieval result of the semantics of the text to be processed after performing semantic understanding on the text to be processed. Figure 2 A scenario schematic diagram of a long text understanding scenario is shown. Taking the long text understanding scenario as a text summary generation scenario as an example, a text processing model is deployed in the terminal device used by the business object. The business object can submit a text summary generation task for the text to be processed in the terminal device. The text processing model deployed in the terminal device can, after performing semantic understanding on the text to be processed, output the generated text summary to the business object.

[0111] The long text generation scenario refers to a text processing scenario in which semantic understanding is performed on the text to be processed to generate a semantically related text that meets the text generation requirements. For example, if the text to be processed specifies an outline generation requirement, the semantically related text can be an outline text that meets the outline generation requirement. Another example is that if the text to be processed specifies a paper writing requirement, the semantically related text can be a paper text that meets the paper writing requirement. Still another example is that if the text to be processed specifies a text translation requirement, the semantically related text can be a translation text that meets the text translation requirement. Figure 3 A scenario schematic diagram of a long text generation scenario is shown. Taking the long text generation scenario as a text translation scenario as an example, the business object can submit a text translation task to the server through the used terminal device. The text translation requirement is to translate from the first language type to the second language type. A text processing model is deployed in the server. The text processing model deployed in the server can, after performing semantic understanding on the text to be processed, generate a translation text that meets the text translation requirement and send the translation text to the terminal device. The terminal device can output the translation text that meets the text translation requirement to the business object.

[0112] A multi-turn dialogue scenario refers to a text processing scenario where new dialogue text semantically related to the historical dialogue text is generated through semantic understanding of the historical dialogue text. That is to say, the semantically related text is new dialogue text that is semantically related to the historical multi-turn dialogue text. For example, the multi-turn dialogue scenario specifically includes multi-turn human-machine interaction scenarios and dialogue assistant scenarios, etc.; the multi-turn human-machine interaction scenario is similar to the dialogue assistant scenario, both referring to the scenario of having a conversation with a virtual robot. Figure 4 A schematic diagram of a multi-turn dialogue scenario is shown. A text processing model is deployed in the terminal device used by the business object. The business object can have a conversation with a virtual robot (for example, the dialogue name of the virtual robot is artificial customer service) in the terminal device. After the text processing model deployed in the terminal device performs semantic understanding on the conversation content of the business object (for example, Figure 4 "dialogue text 1" in Figure 4 ), it outputs the dialogue content of the virtual robot to the business object in the tone of the virtual robot (for example,

[0113] "dialogue text 2" in ); when the business object and the virtual robot generate multi-turn dialogue content, the text processing model can perform comprehensive semantic understanding on the multi-turn historical dialogue content between the business object and the virtual robot and then output the dialogue content of the virtual robot to the business object in the tone of the virtual robot.

[0113] It is not difficult to find that the above text processing scenarios all require the text length of the input and / or output of the text processing model to be significantly longer than that of ordinary Q&A content. This requires the text processing model to be able to better process input-output sequences of ultra-long lengths, thus posing higher requirements for the extrapolation of relative position encoding. The text processing method proposed in the embodiments of this application aims to enhance the text unit (token) length extrapolation ability of large language models (i.e., text processing models), so that the text processing model can achieve good performance without a large amount of fine-tuning when dealing with longer input-output text tasks. The text processing method proposed in the embodiments of this application can achieve better text processing effects in the above text processing scenarios; in the text understanding scenario, it can help the text processing model improve the understanding ability of long text categories when the training text length is limited; in the text generation scenario, it can help the text processing model improve the generation ability of long text categories when the training text length is limited; in the multi-turn dialogue scenario, it can help the text processing model improve the answering ability of long text categories when the training text length is limited.

[0114] Next, in conjunction with the accompanying drawings, the text processing method proposed in the embodiments of this application will be introduced in more detail.

[0115] An embodiment of the present application proposes a text processing method, which mainly introduces the model structure of a text processing model and the text processing process of the text processing model. This text processing method can be executed by a computer device deploying the text processing model, and the computer device can be, for example, a terminal device or a server. As Figure 5 shown, this text processing method may include, but is not limited to, steps S501 - S504:

[0116] S501, obtain a target text to be processed, where the target text includes multiple text units.

[0117] A target text unit refers to any text to be processed, and the target text may include multiple text units. A text unit refers to a character or a word of natural language in the target text, and the text length of the target text is the second text length, where the second text length refers to the number of text units included in the target text.

[0118] Before introducing the specific text processing process in combination with steps S502 - S504 of the embodiment of the present application, the model structure of the text processing model is introduced here first, so as to introduce the text processing process of the text processing model in combination with the model structure of the text processing model later. As Figure 6 shown, the text processing model can be a model composed of a tokenize layer, an embedding layer, multiple semantic understanding modules (transformer blocks), and a prediction layer. The connection relationship between the various components of the text processing model is as follows: The target text is input into the tokenize layer, and the output end of the tokenize layer is connected to the input end of the embedding layer; the output end of the embedding layer is connected to the input end of the first semantic understanding module; the output end of the first semantic understanding module is connected to the input end of the second semantic understanding module, the output end of the second semantic understanding module is connected to the input end of the third semantic understanding module, and so on; the output end of the last semantic understanding module is connected to the input end of the prediction layer, and the output of the prediction layer can be used as the text processing result output by the text processing model.

[0119] Among them:

[0120] ① Tokenize layer: The tokenize layer can be used to tokenize the input in the form of natural language into a sequence of text units (tokens). Tokenization refers to the tokenization process. That is to say, the tokenize layer can be used to perform tokenization processing on the input target text to obtain multiple text units included in the target text.

[0121] ② Representation layer: The sequence of text units (tokens) can be encoded into the form of feature vectors (embeddings) after passing through the representation layer. Specifically, the representation layer can be used to perform representation analysis on each text unit contained in the target text to obtain the feature vectors of each text unit, and the feature vector of each text unit can be used to uniquely represent the corresponding text unit.

[0122] ③ Semantic understanding module: The semantic understanding module can also be called the attention fusion module. The function of the semantic understanding module is mainly realized by the attention mechanism. The sequence of feature vectors (token embeddings) of text units enters multiple semantic understanding modules for attention fusion, which can enrich the feature semantics. That is to say, multiple semantic understanding modules included in the text processing model can be used to perform semantic understanding (or can be called attention fusion) on each text unit in the target text to obtain the semantic features of each text unit in the target text.

[0123] By setting multiple semantic understanding modules in the text processing model, the semantics of each text unit in the target text can be gradually deepened. Specifically, after the feature vectors of each text unit in the target text are input into the first semantic understanding module, the first semantic understanding module can perform semantic understanding on each text unit to obtain the semantic understanding results of each text unit in the target text under the first semantic understanding module; after the semantic understanding results of each text unit in the target text under the first semantic understanding module are input into the second semantic understanding module, the second semantic understanding module can perform semantic understanding on each text unit to obtain the semantic understanding results of each text unit in the target text under the second semantic understanding module; and so on, the semantic understanding results of each text unit in the target text under the last semantic understanding module can be obtained, and the semantic understanding results of each text unit in the target text under the last semantic understanding module can be used as the final semantic features of each text unit in the target text. In the semantic understanding process of the above multiple semantic understanding modules, the semantic understanding module that performs semantic understanding first outputs the shallow semantic features of the text unit, and the semantic understanding module that performs semantic understanding later outputs the deep semantic features of the text unit. The deep semantic features are richer than the shallow semantic features and have stronger text word discrimination ability; moreover, the semantic understanding processes of each semantic understanding are the same, both based on the attention mechanism for semantic understanding. The specific content of the semantic understanding process based on the attention mechanism can refer to the relevant content of Formulas 1 - 5 above.

[0124] The text processing method proposed in the embodiments of this application involves improving the relative position encoding between two text units in the target text, and the relative position encoding is applied to the attention weights of the attention mechanism. Therefore, the improvement in the embodiments of this application is mainly applied in the semantic understanding module.

[0125] ④ Prediction layer: The semantic features of each text unit in the target text can ultimately be used to predict the feature vector of the next text unit (token) in the sequence. After decoding, the feature vector of the next text unit (token) is restored to the form of natural language. That is to say, the prediction layer can be used to predict the semantically related text of the target text based on the semantic features of each text unit in the target text. Moreover, during the process of predicting the semantically related text of the target text, the text processing model can only predict one semantically related text unit at a time. The predicted semantically related text unit needs to be concatenated to the input target text as the new target text, and then the new semantically related text unit is predicted again. Subsequently, the new semantically related text units are predicted recursively until the recursive termination condition is reached. Then, the predicted semantically related text units can be concatenated as the semantically related text for the entire paragraph output.

[0126] Based on the model structure of the text processing model, the overall text processing flow of the text processing model can include: The tokenization layer of the text processing model can be called to tokenize the target text to obtain multiple text units included in the target text; the representation layer of the text processing model can be called to perform representation analysis on each text unit in the target text to obtain the feature vector of each text unit in the target text; multiple semantic understanding modules in the text processing model can be called to perform semantic understanding on each text unit in the target text based on the feature vector of each text unit in the target text to obtain the semantic feature of each text unit in the target text; the prediction layer in the text processing model can be called to perform prediction processing on the target text according to the semantic features of each text unit in the target text to obtain the predicted semantically related text unit. Then, the predicted semantically related text unit can be concatenated to the target text as the new target text, and the above text processing flow is repeated to recursively predict the new semantically related text unit until the recursive termination condition is reached. Finally, the text obtained by concatenating all the predicted semantically related text units is output as the semantically related text of the target text.

[0127] Based on the model structure of the text processing model and the overall text processing flow of the text processing model, the technical details in the text processing flow will be introduced below in combination with steps S502 - S504.

[0128] S502. Obtain the position encoding bias information corresponding to the target text. The position encoding bias information includes the relative position encoding between every two text units in the target text, and the relative position encoding between every two text units is determined according to the relative positions of the two text units in the target text.

[0129] The position encoding information corresponding to the target text may include the relative position encoding between every two text units in the target text. The relative position encoding between every two text units may be determined according to the relative positions of every two text units in the target text; the relative position encoding between every two text units may be calculated based on the Alibi relative position encoding method described in the above formula 6 for the relative positions of every two text units in the target text. Specifically, the relative position of every two text units in the target text specifically refers to the distance between every two text units in the target text, which is the difference between the position serial numbers of the arrangement positions of every two text units in the target text. For example, in the above formula 6, the i-th text unit is arranged at the i-th position in the target text, and the position serial number is i, and the j-th text unit is arranged at the j-th position in the target text, and the position serial number is j, then the relative position between the i-th text unit and the j-th text unit is (i - j); the relative position of every two text units in the target text can be calculated according to the relative position coefficient m to obtain the relative position encoding between every two text units. For example, the relative position encoding between the i-th text unit and the j-th text unit is -m*(i - j).

[0130] Further, as described above, in the process of calculating the attention features (i.e., semantic features), the attention mechanism focuses on the similarity between each text unit and the previous text unit of this text unit. Correspondingly, the relative position encoding also considers the relative position encoding between each text unit and the previous text unit of this text unit. That is to say, the position encoding bias information corresponding to the target text may include the relative position encoding between each text unit in the target text and the previous text unit of this text unit. The previous text unit of each text unit includes each text unit itself and the text units arranged before each text unit in the target text.

[0131] In addition, the relative position encoding between every two text units in the target text can be applied to multiple attention fusion modules (i.e., semantic understanding modules) to participate in the calculation of attention weights. The relative position coefficients corresponding to multiple attention fusion modules may be the same or different; if the relative position coefficients corresponding to each attention fusion module are the same, then the position encoding bias information introduced in each attention fusion module is the same, so there is no need to calculate the position encoding bias information separately for each attention fusion module, which can improve the text processing efficiency to a certain extent; if the relative position coefficients corresponding to each attention fusion module are different, then the position encoding bias information introduced in each attention fusion module is different, so different relative position coefficients can be configured according to the semantic understanding requirements of each attention fusion module to improve the semantic understanding accuracy of each attention fusion module, thereby improving the text processing accuracy to a certain extent.

[0132] Optionally, each attention fusion module may include multiple attention fusion processes, and the results of the multiple attention fusion processes are concatenated to obtain the output result of each attention fusion module. It can be further understood that the attention mechanism adopted in each attention fusion module may be a multi-head attention mechanism; the multi-head attention mechanism means that in the attention fusion process of each attention fusion module, the input is mapped to different dimensions, and an attention fusion process is performed separately in each dimension. Finally, the attention fusion processing results of each dimension are concatenated as the output result of each attention fusion module. In the multi-head attention mechanism, the relative position coefficients corresponding to each attention fusion process may be different, that is, the relative position coefficient m may take different values in different heads. In this way, different relative position coefficients can be configured for different dimension mapping requirements to improve the semantic understanding accuracy of each dimension, thereby improving the text processing accuracy to a certain extent.

[0133] S503. Call the text processing model to map the first relative position encoding in the position encoding bias information to the corresponding position encoding range of the text processing model, and obtain the target relative position encoding corresponding to the first relative position encoding; the first relative position encoding is the relative position encoding in the position encoding bias information whose corresponding relative position exceeds the position threshold.

[0134] After obtaining the position encoding bias information corresponding to the target text (including the relative position encoding between two text units in the target text), the text processing model can be called to map the first relative position encoding in the position encoding bias information to the corresponding position encoding range of the text processing model, and obtain the target relative position encoding corresponding to the first relative position encoding; the text processing model can be called to determine the second relative position encoding in the position encoding bias information as the target relative position encoding corresponding to the second relative position encoding; the first relative position encoding is the relative position encoding in the position encoding bias information whose corresponding relative position exceeds (exceeds means greater than) the position threshold, and the second relative position encoding is the relative position encoding in the position encoding bias information whose corresponding relative position does not exceed (does not exceed means less than or equal to) the position threshold. The mapped position encoding bias information may include the target relative position encoding corresponding to the first relative position encoding and the target relative position encoding corresponding to the second relative position encoding.

[0135] Among them, the position encoding range corresponding to the text processing model refers to the position encoding range that enables the text processing model to achieve a better text processing effect. After introducing the attention mechanism, the relative position encoding within the position encoding range enables the text processing model to achieve a better text processing effect. The position encoding range corresponding to the text processing model can be determined according to the first text length, where the first text length refers to the longest text length processed by the text processing model during training. The position encoding range corresponding to the text processing model refers to the value range of the relative position encoding between any two text units in the text under the longest text length; for example, when the first text length is l 1 At this time, the position encoding range corresponding to the text processing model is [-m*(l 1 -1), 0]. When the text length of the target text to be processed (i.e., the second text length) is greater than the first text length, the position encoding range corresponding to the target text (i.e., the value range of the relative position encoding between any two text units in the target text) exceeds the position encoding range corresponding to the text processing model; for example, when the second text length is l 2 (l 2 >l 1 ) At this time, the position encoding range corresponding to the text processing model is [-m*(l 2 -1), 0], [-m*(l 2 -1), 0] exceeds [-m*(l 1 -1), 0], which will affect the text processing effect of the text processing model. Therefore, the embodiments of the present application map the relative position encoding between any two text units in the target text to the position encoding range corresponding to the text processing model, improving the text processing effect of the text processing model on long texts (a long text refers to a text whose length exceeds the maximum text length that appears during the training of the text processing model).

[0136] During the mapping process, the embodiments of the present application provide a custom parameter, which is a position threshold with a value range of [0, l 1 . The physical meaning of this parameter is the inflection point of the relative position encoding. When the relative position between any two text units in the target text (i.e., the distance between any two text units in the target text) is less than or equal to the position threshold, the relative position encoding between the two text units remains unchanged, that is, it is consistent with the relative position encoding obtained by encoding in the Alibi relative position encoding method. In this way, the characteristics of the local attention mechanism can be kept unchanged within the window m, reducing the impact on the local attention weight (attention score) and maintaining the dependence relationship between adjacent text units; when the relative position between any two text units in the target text (i.e., the distance between any two text units in the target text) is greater than the position threshold and less than the extended second text length l 2When it is possible to keep the position encoding range corresponding to the text processing model (i.e., the value range of the relative position encoding corresponding to the text processing model at the first text length l 1 unchanged, map the relative position encoding between every two text units to the position encoding range corresponding to the text processing model, so that the value range of the extrapolated relative position encoding is the same as that before extrapolation, reducing the impact of unseen value ranges on the generalization of relative position encoding. 1 -1), 0], and map the relative position encoding between every two text units to the position encoding range corresponding to the text processing model, so that the value range of the extrapolated relative position encoding is the same as that before extrapolation, reducing the impact of unseen value ranges on the generalization of relative position encoding.

[0137] S504. Invoke the text processing model to perform semantic understanding on the target text according to the mapped position encoding offset information, and generate a semantic associated text of the target text.

[0138] As described above, in one semantic understanding of the text processing model, only one semantic associated text unit can be generated. Therefore, the text processing model needs to perform recursive semantic understanding to generate multiple semantic associated text units and splice them to obtain the semantic associated text. Recursive semantic understanding means that on the basis of the previous semantic understanding, new semantic understanding is continuously performed until the recursive termination condition is reached. The text corresponding to the new semantic understanding is obtained by splicing the text corresponding to the previous semantic understanding and the semantic associated text unit generated by the previous semantic understanding.

[0139] Specifically, in combination with step S502 - step S503, after obtaining the mapped position encoding offset information, the text processing model can be invoked to perform the first semantic understanding on the target text according to the mapped position encoding offset information, and generate a semantic associated text unit corresponding to the first semantic understanding; use the text composed of the semantic associated text unit corresponding to the first semantic understanding and the target text as the new target text, and continue to invoke the text processing model to perform recursive semantic understanding until the recursive termination condition is reached; after reaching the recursive termination condition, determine the text composed of the semantic associated text unit corresponding to the first semantic understanding and each semantic associated text unit obtained by recursive semantic understanding as the semantic associated text of the target text.

[0140] Among them, reaching the recursive termination condition includes any one of the following: the predicted semantic associated text unit is a specified text unit. For example, the specified text unit refers to a special end symbol; the number of predicted semantic associated text units reaches a specified number. For example, when it is required to generate a text summary of 500 units, when the number of predicted semantic associated text units reaches 500, the recursive termination condition is reached.

[0141] For example, Figure 7As shown, after the text processing model is called for semantic understanding three times, the recursive termination condition is reached. The text processing model is called to perform the first semantic understanding on the target text to generate the first semantic associated text unit; after the target text and the first semantic associated text are concatenated into a new target text, the text processing model is called to perform the second semantic understanding on the new target text to generate the second semantic associated text unit; after the new target text and the second semantic associated text are concatenated into an updated target text, the text processing model is called to perform the third semantic understanding on the updated target text to generate the third semantic associated text unit; finally, the text obtained by concatenating the first semantic associated text unit, the second semantic associated text unit, and the third semantic associated text unit can be output as the semantic associated text unit of the target text.

[0142] In the embodiments of the present application, when the text length of the target text to be processed exceeds the text length processed during the training of the text processing model, for the relative position encoding between two text units in the target text, when the relative position corresponding to the relative position encoding is greater than the position threshold, the relative position encoding can be mapped to the corresponding position encoding range of the text processing model, and the target text is processed based on the relative position encoding within the position encoding range, which can enable the text processing model to maintain a good text processing effect; when the relative position corresponding to the relative position encoding is less than or equal to the position threshold, the original relative position encoding can be kept unchanged to avoid affecting the attention weight distribution between adjacent text units after mapping. That is to say, the embodiments of the present application can extend the text length processed by the text processing model on the premise of ensuring that the text processing effect of the text processing model is not affected, so that the text processing model can process texts with text lengths exceeding the training text length. Experiments prove that compared with the Alibi relative position encoding method, the relative position encoding method proposed in the embodiments of the present application can effectively extend the context window of the text length from 2k (2k refers to 2000 text units) to 8k (8k refers to 8000 text units), and the accuracy rate on the long text validation set is increased by more than 10%.

[0143] The embodiments of the present application propose a text processing method, which mainly introduces the specific mapping process of relative position encoding and the setting method of the position threshold. This text processing method can be executed by a computer device deploying the text processing model, and the computer device can be, for example, a terminal device or a server.

[0144] As Figure 8 shown, this text processing method may include but is not limited to steps S801 - S805:

[0145] S801, obtain the target text to be processed, where the target text includes multiple text units.

[0146] In the embodiment of the present application, the execution process of step S801 is the same as that of step S501 in the above-mentioned Figure 5 illustrated embodiment, and for the specific description, reference can be made to the relevant description of step S501 in the above-mentioned Figure 5 illustrated embodiment, which will not be elaborated here.

[0147] S802. Obtain the position encoding offset information corresponding to the target text. The position encoding offset information includes the relative position encoding between every two text units in the target text, and the relative position encoding between every two text units is determined according to the relative positions of every two text units in the target text.

[0148] In the embodiment of the present application, the execution process of step S802 is the same as that of step S502 in the above-mentioned Figure 5 illustrated embodiment, and for the specific description, reference can be made to the relevant description of step S502 in the above-mentioned Figure 5 illustrated embodiment, which will not be elaborated here.

[0149] S803. Call the text processing model to map the first relative position encoding in the position encoding offset information to the corresponding position encoding range of the text processing model, and obtain the target relative position encoding corresponding to the first relative position encoding; the first relative position encoding is the relative position encoding in the position encoding offset information whose corresponding relative position exceeds the position threshold.

[0150] For the convenience of understanding, here, taking any text unit in the target text (any text unit can be represented as the i-th text unit) as an example, the mapping process of the relative position encoding of the i-th text unit is introduced:

[0151] The target text may include N text units. Any text unit among the N text units can be represented as the i-th text unit, and the i-th text unit is arranged at the i-th position in the target text. N is an integer greater than 1, and i is an integer less than or equal to N; the position encoding offset information may include the relative position encoding between the i-th text unit and each of its previous text units. Any previous text unit refers to any text unit arranged at the 1st position - the i-th position in the target text. That is to say, the previous text units of the i-th text unit are the 1st text unit, the 2nd text unit,..., the i-th text unit in the target text units.

[0152] Based on this, the text processing model can be called to map the first relative position encoding in the relative position encodings between the i-th text unit and each previous text unit to the corresponding position encoding range of the text processing model, so as to obtain the target relative position encoding corresponding to the first relative position encoding; wherein, the first relative position encoding is the relative position encoding in the relative position encodings between the i-th text unit and each previous text unit, corresponding to the relative positions exceeding the position threshold.

[0153] Further, the essence of the mapping is interpolation processing. For the purpose of facilitating understanding, taking the relative position encoding between the i-th text unit and any previous text unit of the i-th text unit (any previous text unit of the i-th text unit can be expressed as the j-th previous text unit, where j is a positive integer less than or equal to i) as an example, the mapping process of the relative position encoding between the i-th text unit and the j-th previous text unit is introduced. Specifically, if the relative position between the i-th text unit and the j-th previous text unit exceeds the position threshold, an interpolation function can be obtained; the interpolation function is generated according to the second text length, the first text length, and the position threshold. The second text length refers to the text length of the target text, and the text length of the target text refers to the number of text units included in the target text. The first text length is the maximum text length processed by the text processing model during training. The value range of the interpolation function is the corresponding position encoding range of the text processing model; then, according to the interpolation function, interpolation processing can be performed on the first relative position encoding between the i-th text unit and the j-th previous text unit to obtain the target relative position encoding corresponding to the first relative position encoding between the i-th text unit and the j-th previous text unit.

[0154] S804, call the text processing model to determine the second relative position encoding in the position encoding bias information as the target relative position encoding corresponding to the second relative position encoding; the second relative position encoding is the relative position encoding in the position encoding bias information corresponding to the relative positions not exceeding the position threshold.

[0155] For the purpose of facilitating understanding, taking any text unit in the target text (any text unit can be expressed as the i-th text unit) as an example, the mapping process of the relative position encoding regarding the i-th text unit is introduced:

[0156] The text processing model can be called to determine the second relative position encoding in the relative position encodings between the i-th text unit and each previous text unit as the target relative position encoding corresponding to the second relative position encoding; wherein, the second relative position encoding is the relative position encoding in the relative position encodings between the i-th text unit and each previous text unit, corresponding to the relative positions not exceeding the position threshold.

[0157] In summary, for the mapping process of the relative position encoding of the i-th text unit in steps S803 - S804 above, it can be summarized into the following formula 8 (formula 8 can also be expressed as a piecewise function in formula 9 below):

[0158]

[0159]

[0160] In the above formulas 9 and 10, represents the target relative position encoding between the i-th text unit and the j-th previous text unit, and l 1 represents the first text length, and l 2 represents the second text length. represents the position threshold. It can be seen from the above formulas 8 and 9 that when the relative position between the i-th text unit and the j-th previous text unit (i.e., the distance between the i-th text unit and the j-th previous text unit) is greater than the position threshold , the first relative position encoding [-m*(i - j)] between the i-th text unit and the j-th previous text unit can be mapped to the corresponding position encoding range of the text processing model (i.e., the value range [-m*(l 1 -1), 0]) of the text processing model, and is actually mapped to the target relative position encoding When the relative position between the i-th text unit and the j-th previous text unit (i.e., the distance between the i-th text unit and the j-th previous text unit) is less than or equal to the position threshold , the second relative position encoding [-m*(i - j)] between the i-th text unit and the j-th previous text unit can be kept unchanged.

[0161] Figure 9Shows the comparison diagram between the relative position encoding in the Alibi relative position encoding method (which can be understood as before text length extrapolation) and the relative position encoding in the LeakyAlibi relative position encoding method proposed in the embodiment of the present application (which can be understood as after text length extrapolation). It can be seen that: First, the relationship between the relative position encoding and the relative position is the same in both the Alibi relative position encoding method and the LeakyAlibi relative position encoding method. The larger the relative position between text units (i.e., the farther the distance), the larger the absolute value of the relative position encoding. Conversely, the smaller the relative position between text units (i.e., the closer the distance), the smaller the absolute value of the relative position encoding. Second, in the Alibi relative position encoding method and the LeakyAlibi relative position encoding method, the value range of the relative position encoding remains consistent. That is to say, before and after text length extrapolation, the value range of the relative position encoding remains consistent, reducing the impact of unseen value ranges on the generalization of relative position encoding and improving the text processing effect of the text processing model after text length extrapolation. Third, before the inflection point (the inflection point refers to the position threshold ), the relative position encoding in the Alibi relative position encoding method and the LeakyAlibi relative position encoding method is the same, with little impact on the local attention weight and maintaining the dependency relationship between adjacent text units; after the inflection point (the position threshold ), the relative position encoding in the LeakyAlibi relative position encoding method is interpolated into the value range of the relative position encoding before text length extrapolation, expanding the text length supported by the relative position encoding while ensuring the text processing effect of the text processing model.

[0162] S805, call the text processing model to perform semantic understanding on the target text according to the mapped position encoding offset information, and generate the semantic related text of the target text.

[0163] For ease of understanding, here, taking any text unit in the target text (any text unit can be represented as the i-th text unit) as an example, introduce the process in which the relative position encoding of the i-th text unit participates in semantic understanding:

[0164] The mapped position encoding bias information may include the target relative position encoding between the $i$-th text unit and each previous text unit; the first attention weight between the $i$-th text unit (specifically, the element value of the $i$-th text unit in the Q matrix) and each previous text unit (specifically, the element value of the $j$-th previous text unit in the K matrix) can be determined according to their similarity; the first attention weight between the $i$-th text unit and each previous text unit can be updated according to the target relative position encoding between the $i$-th text unit and each previous text unit to obtain the second attention weight between the $i$-th text unit and each previous text unit; the semantic feature corresponding to the $i$-th text unit can be obtained by performing weighted summation processing on each previous text unit (specifically, the element value of the previous text unit in the V matrix) according to the second attention weight between the $i$-th text unit and each previous text unit.

[0165] For example, as Figure 10 shown, the target text includes 4 text units, namely the 1st text unit, the 2nd text unit, the 3rd text unit, and the 4th text unit. Taking the 3rd text unit as an example, the process of determining the semantic feature of the 3rd text unit according to its relative position encoding is introduced. First, the first attention weight $q$ 3 $k$ 1 between the 3rd text unit and the 1st text unit can be determined according to their similarity, and so on. The first attention weight $q$ 3 $k$ 2 between the 3rd text unit and the 2nd text unit can be obtained, and the first attention weight $q$ 3 $k$ 3 between the 3rd text unit and the 3rd text unit can be obtained. Second, the first attention weight $q$ between the 3rd text unit and the 1st text unit can be updated according to the target relative position encoding 3 $k$ 1 between the 3rd text unit and the 1st text unit to obtain the second attention weight $q$ 3 $k$ ′ 1 between the 3rd text unit and the 1st text unit, and so on. The second attention weight $q$ 3 $k$ ′ 2 between the 3rd text unit and the 2nd text unit can be obtained, and the second attention weight $q$ 3 $k$ ′ 3Then, based on the obtained second attention weights, weighted summation processing can be performed on the first text unit, the second text unit, and the third text unit (q 3 k ′ 1 v 1 +q 3 k ′ 2 v 2 +q 3 k ′ 3 v 3 ), to obtain the semantic feature of the third text unit, where v 1 represents the element value of the first text unit in the V matrix, v 2 represents the element value of the second text unit in the V matrix, v 3 represents the element value of the third text unit in the V matrix.

[0166] According to the determination method of the semantic feature of the above-mentioned i-th text unit, the semantic features of other text units in the target text except the i-th text unit can be determined; for the determination method of the semantic features of other text units, specifically, reference can be made to the determination method of the semantic feature of the i-th text unit, which will not be elaborated here. After determining the semantic features of each text unit in the target text, based on the semantic features of each text unit in the target text, semantic understanding of the target text can be performed to generate a semantic association text of the target text.

[0167] Among them, semantic understanding can specifically refer to the recursive semantic understanding described in step S504 of the above-mentioned Figure 5 illustrated embodiment. Specifically, the text processing model can be called to determine the semantic features of each text unit in the target text according to the mapped position encoding bias information, and semantic understanding of the target text can be performed based on the semantic features of each text unit in the target text to generate a semantic association text unit corresponding to the first semantic understanding; the text composed of the semantic association text unit corresponding to the first semantic understanding and the target text is used as a new target text, and the text processing model is continuously called for recursive semantic understanding until the recursive termination condition is reached; after reaching the recursive termination condition, the semantic association text unit corresponding to the first semantic understanding, and the text composed of each semantic association text unit obtained by recursive semantic understanding are determined as the semantic association text of the target text.

[0168] It should be noted that the text processing model is a pre-trained large language model, and the text processing model has text processing capabilities (the text processing capabilities here refer to the ability to semantically understand the input text and generate semantically related text of the input text). Therefore, the text processing model can be directly used for reasoning. Directly using for reasoning means directly using the text processing model for text processing; or, when resources are sufficient to support the secondary fine-tuning of the text processing model, the text processing model can be fine-tuned according to the task requirements of the text processing task.

[0169] For the above two usage scenarios of the text processing model (i.e., direct reasoning or secondary fine-tuning), different settings can be made for the position threshold Specifically, the acquisition method of the position threshold can include any of the following:

[0170] First, for the case of direct reasoning of the text processing model, the position threshold can be set according to the first text length. The first text length is the maximum text length processed by the text processing model during training. For example, if the first text length is l 1 , the position threshold can be set to l 1 / 2. In this case, the position threshold is set to l 1 / 2 according to the empirical value, so that when the text processing model processes long texts, it can achieve good performance without a large amount of fine-tuning and reach a better text processing effect.

[0171] Second, for the case of secondary fine-tuning of the text processing model, the position threshold can be set as a learnable model parameter and adjusted to a relatively ideal value during the fine-tuning process of the text processing model. This relatively ideal value can enable the text processing model to achieve a better text processing effect when processing long texts; and, the initial value of this learnable model parameter can be set to l 1 , and it is restricted to the interval [0, l 1 during the fine-tuning process. That is to say, the initial threshold (the initial threshold refers to the initial value l 1 ) can be used as the model parameter of the text processing model and adjusted during the fine-tuning process of the text processing model to obtain the position threshold.

[0172] Specifically, taking the initial threshold as a model parameter of the text processing model and adjusting it during the fine-tuning process of the text processing model to obtain the position threshold may include: First, a sample text set can be obtained. The sample text set may include multiple sample texts, and the text length of each sample text is greater than the maximum text length (i.e., the first text length) processed by the text processing model during training. Moreover, if the text processing task specifies a text length, the text length of each sample text in the sample text set can be the specified text length, which can enable the fine-tuned text processing model and the position threshold to better adapt to the task requirements of the text processing task. Second, the text processing model can be iteratively fine-tuned according to the sample text set to adjust the model parameters of the text processing model. When adjusting the model parameters of the text processing model, the model parameters serving as the position threshold can be adjusted, and other model parameters can be not adjusted, or the model parameters serving as the position threshold and other model parameters can be adjusted together. Then, when the iterative fine-tuning termination condition is reached, the model parameters in the text processing model when the iterative fine-tuning termination condition is reached can be determined as the position threshold.

[0173] Furthermore, iterative fine-tuning refers to performing multiple fine-tuning operations on the text processing model, and the subsequent fine-tuning is based on the previous fine-tuning. The sample text used in any fine-tuning process during the iterative fine-tuning process can be represented as a reference text, and the sample text set can also include the marked semantic associated text of the reference text. Any fine-tuning process during the iterative fine-tuning process may include: First, the reference text can be obtained. The reference text may include multiple sample text units, and the position encoding bias information corresponding to the reference text can be obtained. The position encoding bias information corresponding to the reference text may include the relative position encoding between two sample text units in the reference text, and the relative position encoding between two sample text units is determined according to the relative positions of the two sample text units in the reference text. Second, the text processing model can be called to map the first relative position encoding in the position encoding bias information corresponding to the reference text to the corresponding position encoding range of the text processing model to obtain the target relative position encoding corresponding to the first relative position encoding. The first relative position encoding is the relative position encoding in the position encoding bias information corresponding to the reference text where the corresponding relative position exceeds the position threshold. Then, the text processing model can be called to perform semantic understanding on the reference text according to the mapped position encoding bias information to generate the actual semantic associated text of the reference text. The loss information of the text processing model can be determined according to the difference between the actual semantic associated text of the reference text and the marked semantic associated text of the reference text, and the model parameters of the text processing model can be adjusted according to the loss information of the text processing model.

[0174] It should be noted that meeting the termination condition of iterative fine-tuning can include any of the following: the number of fine-tuning times included in iterative fine-tuning reaches the number threshold, and the loss information of the text processing model is less than the loss threshold. In addition, the fine-tuning process of the text processing model is similar to the inference process of the text processing model. The similar steps in the fine-tuning process of the text processing model and in the inference process of the text processing model will not be elaborated here. Specifically, reference can be made to the inference process of the text processing model. By fine-tuning the position threshold and adjusting it to a relatively ideal value during the fine-tuning process of the text processing model, on the one hand, this relatively ideal value can enable the text processing model to achieve better text processing effects when processing long texts, and on the other hand, it can enable the text processing model to better meet the task requirements of the text processing task.

[0175] In the embodiments of the present application, when the text length of the target text to be processed exceeds the text length processed during the training of the text processing model, for the relative position encoding between two text units in the target text, when the relative position corresponding to the relative position encoding is greater than the position threshold, the relative position encoding can be mapped to the corresponding position encoding range of the text processing model, and the target text can be processed based on the relative position encoding within the position encoding range, which can enable the text processing model to maintain good text processing effects; when the relative position corresponding to the relative position encoding is less than or equal to the position threshold, the original relative position encoding can be kept unchanged to avoid affecting the attention weight distribution between adjacent text units after mapping. That is to say, the embodiments of the present application can expand the text length processed by the text processing model on the premise of ensuring that the text processing effects of the text processing model are not affected, so that the text processing model can process texts with text lengths exceeding the training text length. In addition, by adopting different setting methods for the position threshold in different usage scenarios (direct inference or secondary fine-tuning) of the text processing model, for the usage scenario of direct inference of the text processing model, the position threshold is set according to the empirical value, so that the text processing model can achieve good performance and better text processing effects when processing long texts without a large amount of fine-tuning; for the usage scenario of secondary fine-tuning of the text processing model, during the fine-tuning process of the text processing model, the position threshold is adjusted as a learnable model parameter, so that the adjusted position threshold can enable the text processing model to achieve better text processing effects when processing long texts.

[0176] The above has elaborated in detail the method of the embodiments of the present application. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below.

[0177] Please refer to Figure 11 , Figure 11It is a schematic structural diagram of a text processing device provided by an embodiment of the present application. The text processing device can be set in the computer device provided by the embodiment of the present application, and the computer device can be a terminal device or a server. Figure 11 The text processing device shown can be a computer program running in a computer device, and this text processing device can be used to execute Figure 5 or Figure 8 part or all of the steps in the method embodiment shown. Please refer to Figure 12 , and this text processing device can include the following units:

[0178] An acquisition unit 1101, configured to acquire a target text to be processed, where the target text includes a plurality of text units;

[0179] The acquisition unit 1101 is further configured to acquire position encoding offset information corresponding to the target text, where the position encoding offset information includes relative position encodings between two text units in the target text, and the relative position encodings between two text units are determined according to the relative positions of the two text units in the target text;

[0180] A processing unit 1102, configured to call a text processing model to map the first relative position encoding in the position encoding offset information to the corresponding position encoding range of the text processing model, so as to obtain a target relative position encoding corresponding to the first relative position encoding; the first relative position encoding is the relative position encoding in the position encoding offset information where the corresponding relative position exceeds a position threshold;

[0181] The processing unit 1102 is further configured to call the text processing model to perform semantic understanding on the target text according to the mapped position encoding offset information, and generate a semantic association text of the target text.

[0182] In one implementation, the processing unit 1102 is further configured to perform the following steps:

[0183] Call the text processing model to determine the second relative position encoding in the position encoding offset information as the target relative position encoding corresponding to the second relative position encoding; the second relative position encoding is the relative position encoding in the position encoding offset information where the corresponding relative position does not exceed the position threshold;

[0184] Wherein, the mapped position encoding offset information includes the target relative position encoding corresponding to the first relative position encoding and the target relative position encoding corresponding to the second relative position encoding.

[0185] In one implementation, the target text includes N text units. Any one of the N text units is denoted as the i-th text unit, and the i-th text unit is arranged at the i-th position in the target text. N is an integer greater than 1, and i is an integer less than or equal to N. The positional encoding bias information includes the relative positional encoding between the i-th text unit and each of its previous text units. Any previous text unit refers to any text unit arranged at the 1st position - the i-th position in the target text;

[0186] The processing unit 1102, when used to call the text processing model to map the first relative positional encoding in the positional encoding bias information to the corresponding positional encoding range of the text processing model to obtain the target relative positional encoding corresponding to the first relative positional encoding, is specifically used to perform the following steps:

[0187] Call the text processing model to map the first relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units to the corresponding positional encoding range of the text processing model to obtain the target relative positional encoding corresponding to the first relative positional encoding;

[0188] Wherein, the first relative positional encoding is the relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units, corresponding to the relative position exceeding the position threshold.

[0189] In one implementation, the processing unit 1102, when used to call the text processing model to determine the second relative positional encoding in the positional encoding bias information as the target relative positional encoding corresponding to the second relative positional encoding, is specifically used to perform the following steps:

[0190] Call the text processing model to determine the second relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units as the target relative positional encoding corresponding to the second relative positional encoding;

[0191] Wherein, the second relative positional encoding is the relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units, corresponding to the relative position not exceeding the position threshold.

[0192] In one implementation, the relative position between the i-th text unit and the j-th previous text unit exceeds the position threshold, where j is a positive integer less than or equal to i. The processing unit 1102, when used to call the text processing model to map the first relative positional encoding in the relative positional encoding between the i-th text unit and each of its previous text units to the corresponding positional encoding range of the text processing model to obtain the target relative positional encoding corresponding to the first relative positional encoding, is specifically used to perform the following steps:

[0193] Obtain an interpolation function; the interpolation function is generated based on the first text length, the second text length, and a position threshold. The first text length refers to the maximum text length processed by the text processing model during training. The second text length refers to the text length of the target text, and the text length of the target text refers to the number of text units included in the target text. The value range of the interpolation function is the position encoding range corresponding to the text processing model.

[0194] According to the interpolation function, perform interpolation processing on the first relative position encoding between the i-th text unit and the j-th previous text unit to obtain the target relative position encoding corresponding to the first relative position encoding between the i-th text unit and the j-th previous text unit.

[0195] In one implementation, the mapped position encoding bias information includes the target relative position encoding between the i-th text unit and each previous text unit. The processing unit 1102, when used to call the text processing model to perform semantic understanding on the target text and generate a semantic associated text of the target text, is specifically used to perform the following steps:

[0196] Determine the first attention weight between the i-th text unit and each previous text unit according to the similarity between the i-th text unit and each previous text unit.

[0197] Update the first attention weight between the i-th text unit and each previous text unit according to the target relative position encoding between the i-th text unit and each previous text unit to obtain the second attention weight between the i-th text unit and each previous text unit.

[0198] Perform weighted summation processing on each previous text unit according to the second attention weight between the i-th text unit and each previous text unit to obtain the semantic feature corresponding to the i-th text unit.

[0199] Perform semantic understanding on the target text according to the semantic features of each text unit in the target text to generate a semantic associated text of the target text.

[0200] In one implementation, the way for the processing unit 1102 to obtain the position threshold includes any one of the following:

[0201] Set the position threshold according to the first text length, where the first text length is the maximum text length processed by the text processing model during training.

[0202] Use the initial threshold as a model parameter of the text processing model and adjust it during the fine-tuning process of the text processing model to obtain the position threshold.

[0203] In one implementation, the processing unit 1102 is configured to use the initial threshold as a model parameter of the text processing model and adjust it during the fine-tuning process of the text processing model. When obtaining the position threshold, it is specifically configured to perform the following steps:

[0204] Obtain a sample text set, where the sample text set includes multiple sample texts, and the text length of each sample text is greater than the maximum text length processed by the text processing model during training;

[0205] Iteratively fine-tune the text processing model according to the sample text set and adjust the model parameters of the text processing model;

[0206] When the iterative fine-tuning termination condition is reached, determine the model parameters in the text processing model when the iterative fine-tuning termination condition is reached as the position threshold.

[0207] In one implementation, the processing unit 1102 is configured to call the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information and generate a semantic association text of the target text. It is specifically configured to perform the following steps:

[0208] Call the text processing model to perform a first semantic understanding on the target text according to the mapped position encoding bias information and generate a semantic association text unit corresponding to the first semantic understanding;

[0209] Use the text composed of the semantic association text unit corresponding to the first semantic understanding and the target text as the new target text, and continue to call the text processing model for recursive semantic understanding until the recursive termination condition is reached;

[0210] After the recursive termination condition is reached, determine the text composed of the semantic association text unit corresponding to the first semantic understanding and the semantic association text units obtained by recursive semantic understanding as the semantic association text of the target text;

[0211] Among them, reaching the recursive termination condition includes any one of the following: the predicted semantic association text unit is a specified text unit, and the number of predicted semantic association text units reaches a specified number.

[0212] In one implementation, the obtaining unit 1101 is configured to obtain the position encoding bias information corresponding to the target text. It is specifically configured to perform the following steps:

[0213] Obtain the relative positions of pairwise text units in the target text;

[0214] Calculate the relative positions of pairwise text units in the target text according to the relative position coefficients to obtain the relative position encoding between pairwise text units.

[0215] In one implementation, the text processing model includes multiple attention fusion modules, and the relative position encoding between every two text units in the target text is applied to the multiple attention fusion modules; the coefficients corresponding to the multiple attention fusion modules are the same or different;

[0216] If each attention fusion module includes multiple attention fusion processes, the coefficients corresponding to each attention fusion process are different.

[0217] According to another embodiment of the present application, Figure 11 Each unit in the text processing device shown can be separately or entirely combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with functional divisions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the data processing device based on the blockchain may also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0218] According to another embodiment of the present application, it can be achieved by running a computer program that can execute the steps involved in some or all of the methods shown in Figure 5 or Figure 8 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct the text processing device shown in Figure 11 and to implement the text processing method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the above computing device through the computer-readable storage medium, and run therein.

[0219] In the embodiments of the present application, when the text length of the target text exceeds the text length that appears during the training of the text processing model, the position encoding range to which the relative position encoding between two text units in the target text belongs will exceed the position encoding range corresponding to the text processing model, which will result in poor text processing effects of the text processing model. In this case, the embodiments of the present application map the relative position encoding with a relative position exceeding the position threshold to the position encoding range corresponding to the text processing model, so that when the text processing model performs text processing on the target text based on the mapped relative position encoding, it can achieve better text processing effects; that is to say, in the embodiments of the present application, when performing text processing on a target text whose text length exceeds the training text length (i.e., the text length that appears during the training of the text processing model), by mapping the relative position encoding between two text units in the target text to the position encoding range corresponding to the text processing model, the text processing model can achieve better text processing effects for the target text whose text length exceeds the training text length, thereby improving the generalization performance of the text processing model.

[0220] Based on the above method and apparatus embodiments, the embodiments of the present application provide a computer device. Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a computer device provided by the embodiments of the present application. Figure 12 The computer device shown at least includes a processor 1201, an input interface 1202, an output interface 1203, and a computer-readable storage medium 1204. Among them, the processor 1201, the input interface 1202, the output interface 1203, and the computer-readable storage medium 1204 can be connected through a bus or other means.

[0221] The computer-readable storage medium 1204 can be stored in the memory of the computer device. The computer-readable storage medium 1204 is used to store a computer program, and the computer program includes computer instructions. The processor 1201 is used to execute the computer program stored in the computer-readable storage medium 1204. The processor 1201 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, and is suitable for implementing the computer program, specifically suitable for loading and executing the computer program to implement the corresponding method flow or corresponding function.

[0222] An embodiment of the present application also provides a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in a computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the computer device. And, a computer program suitable for being loaded and executed by a processor is also stored in this storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0223] The computer device can be a terminal device or a server. In a specific implementation, the processor 1201 can load and execute the computer program stored in the computer-readable storage medium 1204 to implement the corresponding steps in the above-mentioned Figure 5 or Figure 8 text processing method shown. In a specific implementation, the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to perform the following steps:

[0224] Obtain a target text to be processed, where the target text includes multiple text units;

[0225] Obtain position encoding offset information corresponding to the target text, where the position encoding offset information includes relative position encodings between pairwise text units in the target text, and the relative position encodings between pairwise text units are determined according to the relative positions of the pairwise text units in the target text;

[0226] Call a text processing model to map the first relative position encoding in the position encoding offset information to the corresponding position encoding range of the text processing model to obtain a target relative position encoding corresponding to the first relative position encoding; the first relative position encoding is the relative position encoding in the position encoding offset information where the corresponding relative position exceeds a position threshold;

[0227] Call a text processing model to perform semantic understanding on the target text according to the mapped position encoding offset information to generate a semantic association text of the target text.

[0228] In one implementation, the computer program in the computer-readable storage medium 1204 is loaded and further used by the processor 1201 to perform the following steps:

[0229] Call the text processing model to determine the second relative position encoding in the position encoding bias information as the target relative position encoding corresponding to the second relative position encoding; the second relative position encoding is the relative position encoding in the position encoding bias information where the corresponding relative position does not exceed the position threshold.

[0230] Among them, the mapped position encoding bias information includes the target relative position encoding corresponding to the first relative position encoding and the target relative position encoding corresponding to the second relative position encoding.

[0231] In one implementation, the target text includes N text units, any one of the N text units is represented as the i-th text unit, the i-th text unit is arranged at the i-th position in the target text, N is an integer greater than 1, and i is an integer less than or equal to N; the position encoding bias information includes the relative position encoding between the i-th text unit and each of its previous text units, and any previous text unit refers to any text unit arranged at the 1st position - the i-th position in the target text.

[0232] When the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to call the text processing model to map the first relative position encoding in the position encoding bias information to the corresponding position encoding range of the text processing model to obtain the target relative position encoding corresponding to the first relative position encoding, it is specifically used to execute the following steps:

[0233] Call the text processing model to map the first relative position encoding in the relative position encoding between the i-th text unit and each of its previous text units to the corresponding position encoding range of the text processing model to obtain the target relative position encoding corresponding to the first relative position encoding.

[0234] Among them, the first relative position encoding is the relative position encoding in the relative position encoding between the i-th text unit and each of its previous text units where the corresponding relative position exceeds the position threshold.

[0235] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to call the text processing model to determine the second relative position encoding in the position encoding bias information as the target relative position encoding corresponding to the second relative position encoding, it is specifically used to execute the following steps:

[0236] Call the text processing model to determine the second relative position encoding in the relative position encoding between the i-th text unit and each of its previous text units as the target relative position encoding corresponding to the second relative position encoding.

[0237] Among them, the second relative position encoding is the relative position encoding in which the relative position between the i-th text unit and each previous text unit does not exceed the position threshold among the relative position encodings between the i-th text unit and each previous text unit.

[0238] In one implementation, the relative position between the i-th text unit and the j-th previous text unit exceeds the position threshold, where j is a positive integer less than or equal to i; when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to call the text processing model to map the first relative position encoding in the relative position encodings between the i-th text unit and each previous text unit to the corresponding position encoding range of the text processing model to obtain the target relative position encoding corresponding to the first relative position encoding, it is specifically used to perform the following steps:

[0239] Obtain an interpolation function; the interpolation function is generated according to the first text length, the second text length, and the position threshold. The first text length refers to the maximum text length processed by the text processing model during training; the second text length refers to the text length of the target text, and the text length of the target text refers to the number of text units included in the target text; the value range of the interpolation function is the corresponding position encoding range of the text processing model;

[0240] According to the interpolation function, perform interpolation processing on the first relative position encoding between the i-th text unit and the j-th previous text unit to obtain the target relative position encoding corresponding to the first relative position encoding between the i-th text unit and the j-th previous text unit.

[0241] In one implementation, the mapped position encoding bias information includes the target relative position encoding between the i-th text unit and each previous text unit; when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to call the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information and generate the semantic associated text of the target text, it is specifically used to perform the following steps:

[0242] Determine the first attention weight between the i-th text unit and each previous text unit according to the similarity between the i-th text unit and each previous text unit;

[0243] Update the first attention weight between the i-th text unit and each previous text unit according to the target relative position encoding between the i-th text unit and each previous text unit to obtain the second attention weight between the i-th text unit and each previous text unit;

[0244] Perform weighted summation processing on each previous text unit according to the second attention weight between the i-th text unit and each previous text unit to obtain the semantic feature corresponding to the i-th text unit;

[0245] Semantically understand the target text according to the semantic features of each text unit in the target text, and generate a semantically related text of the target text.

[0246] In one implementation, the method for obtaining the position threshold includes any of the following:

[0247] Set the position threshold according to the first text length, where the first text length is the maximum text length processed by the text processing model during training;

[0248] Use the initial threshold as a model parameter of the text processing model, and adjust it during the fine-tuning process of the text processing model to obtain the position threshold.

[0249] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to use the initial threshold as a model parameter of the text processing model and adjust it during the fine-tuning process of the text processing model to obtain the position threshold, it is specifically used to execute the following steps:

[0250] Obtain a sample text set, where the sample text set includes multiple sample texts, and the text length of each sample text is greater than the maximum text length processed by the text processing model during training;

[0251] Iteratively fine-tune the text processing model according to the sample text set, and adjust the model parameters of the text processing model;

[0252] When the iterative fine-tuning termination condition is reached, determine the model parameters in the text processing model when the iterative fine-tuning termination condition is reached as the position threshold.

[0253] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to call the text processing model to semantically understand the target text according to the mapped position encoding bias information and generate a semantically related text of the target text, it is specifically used to execute the following steps:

[0254] Call the text processing model to perform a first semantic understanding on the target text according to the mapped position encoding bias information, and generate a semantic related text unit corresponding to the first semantic understanding;

[0255] Use the text composed of the semantic related text unit corresponding to the first semantic understanding and the target text as the new target text, and continue to call the text processing model for recursive semantic understanding until the recursive termination condition is reached;

[0256] After reaching the recursive termination condition, the semantic association text unit corresponding to the first semantic understanding and the text composed of each semantic association text unit obtained by recursive semantic understanding are determined as the semantic association text of the target text;

[0257] Among them, reaching the recursive termination condition includes any of the following: the predicted semantic association text unit is a specified text unit, and the number of predicted semantic association text units reaches a specified number.

[0258] In one implementation, when the computer program in the computer-readable storage medium 1204 is loaded and executed by the processor 1201 to obtain the position encoding bias information corresponding to the target text, it is specifically used to execute the following steps:

[0259] Obtain the relative positions of pairwise text units in the target text;

[0260] According to the relative position coefficient, calculate the relative positions of pairwise text units in the target text to obtain the relative position encoding between pairwise text units.

[0261] In one implementation, the text processing model includes multiple attention fusion modules, and the relative position encoding between pairwise text units in the target text is applied to multiple attention fusion modules; the coefficients corresponding to the multiple attention fusion modules are the same or different;

[0262] If each attention fusion module includes multiple attention fusion processes, the coefficients corresponding to each attention fusion process are different.

[0263] In the embodiments of the present application, when the text length of the target text exceeds the text length that appears during the training of the text processing model, the position encoding range to which the relative position encoding between pairwise text units in the target text belongs will exceed the position encoding range corresponding to the text processing model, which will result in poor text processing effects of the text processing model. In this case, the embodiments of the present application map the relative position encoding with a relative position exceeding the position threshold to the position encoding range corresponding to the text processing model, so that when the text processing model performs text processing based on the mapped relative position encoding for the target text, it can achieve better text processing effects; that is to say, in the embodiments of the present application, when performing text processing on a target text with a text length exceeding the training text length (i.e., the text length that appears during the training of the text processing model), by mapping the relative position encoding between pairwise text units in the target text to the position encoding range corresponding to the text processing model, the text processing model can achieve better text processing effects for target texts with text lengths exceeding the training text length, thereby improving the generalization performance of the text processing model.

[0264] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned text processing method.

[0265] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0266] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0267] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (for example, infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0268] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.

Claims

1. A text processing method, It is characterized in that include: Acquire a target text to be processed, wherein the target text includes a plurality of text units; Acquire position coding bias information corresponding to the target text, wherein the position coding bias information includes relative position coding between two text units in the target text, and the relative position coding between two text units is determined according to the relative positions of the two text units in the target text; Calling the text processing model to map the first relative position code in the position code bias information to the position code range corresponding to the text processing model, and obtaining a target relative position code corresponding to the first relative position code; The first relative position code is a relative position code in the position code bias information whose corresponding relative position exceeds a position threshold; The text processing model is called to perform semantic understanding on the target text according to the mapped position encoding bias information, and generate semantically associated text of the target text.

2. The method according to claim 1, It is characterized in that The method further comprises: Calling the text processing model to determine the second relative position code in the position code bias information as the target relative position code corresponding to the second relative position code; the second relative position code is a relative position code whose corresponding relative position in the position code bias information does not exceed the position threshold; The mapped position code offset information includes a target relative position code corresponding to the first relative position code and a target relative position code corresponding to the second relative position code.

3. The method according to claim 2, It is characterized in that The target text includes N text units, any one of the N text units is represented as the i-th text unit, the i-th text unit is arranged at the i-th position in the target text, N is an integer greater than 1, and i is an integer less than or equal to N; the position coding bias information includes the relative position coding between the i-th text unit and each preceding text unit of the i-th text unit, and any preceding text unit refers to any text unit arranged at the 1st position to the i-th position in the target text; The calling of the text processing model to map the first relative position code in the position code bias information to a position code range corresponding to the text processing model to obtain a target relative position code corresponding to the first relative position code includes: Calling the text processing model to map a first relative position code in the relative position codes between the i-th text unit and each of the preceding text units into a position code range corresponding to the text processing model, and obtaining a target relative position code corresponding to the first relative position code; Among them, the first relative position code is the relative position code between the i-th text unit and each of the preceding text units, and the corresponding relative position exceeds the position threshold.

4. The method according to claim 3, It is characterized in that The calling of the text processing model to determine the second relative position code in the position code offset information as the target relative position code corresponding to the second relative position code includes: Calling the text processing model to determine the second relative position code in the relative position code between the i-th text unit and each of the preceding text units as the target relative position code corresponding to the second relative position code; Among them, the second relative position code is a relative position code in which the corresponding relative position does not exceed the position threshold in the relative position code between the i-th text unit and each of the preceding text units.

5. The method according to claim 3, It is characterized in that The relative position between the i-th text unit and the j-th preceding text unit exceeds the position threshold, where j is a positive integer less than or equal to i; the calling of the text processing model to map the first relative position code in the relative position code between the i-th text unit and each of the preceding text units to the position code range corresponding to the text processing model, and obtaining the target relative position code corresponding to the first relative position code, includes: Obtain an interpolation function; the interpolation function is generated according to a first text length, a second text length and the position threshold, the first text length refers to the maximum text length processed by the text processing model during training; the second text length refers to the text length of the target text, and the text length of the target text refers to the number of text units included in the target text; the value range of the interpolation function is the position coding range corresponding to the text processing model; According to the interpolation function, the first relative position code between the i-th text unit and the j-th preceding text unit is interpolated to obtain a target relative position code corresponding to the first relative position code between the i-th text unit and the j-th preceding text unit.

6. The method according to claim 3, It is characterized in that The mapped position coding bias information includes the target relative position coding between the i-th text unit and each of the preceding text units; the calling of the text processing model to perform semantic understanding on the target text according to the mapped position coding bias information to generate semantically associated text of the target text includes: Determine a first attention weight between the i-th text unit and each of the preceding text units according to the similarity between the i-th text unit and each of the preceding text units; According to the target relative position encoding between the i-th text unit and each of the preceding text units, a first attention weight between the i-th text unit and each of the preceding text units is updated to obtain a second attention weight between the i-th text unit and each of the preceding text units; According to the second attention weight between the i-th text unit and each of the preceding text units, weighted sum processing is performed on each of the preceding text units to obtain a semantic feature corresponding to the i-th text unit; According to the semantic features of each text unit in the target text, the target text is semantically understood to generate a semantically associated text of the target text.

7. The method according to claim 1, It is characterized in that The method for obtaining the position threshold includes any of the following: Setting the position threshold according to a first text length, where the first text length is a maximum text length processed by the text processing model during training; The initial threshold is used as a model parameter of the text processing model and is adjusted during the fine-tuning process of the text processing model to obtain the position threshold.

8. The method according to claim 7, It is characterized in that The initial threshold is used as a model parameter of the text processing model, and is adjusted during the fine-tuning process of the text processing model to obtain the position threshold, including: Acquire a sample text set, the sample text set comprising a plurality of sample texts, the text length of each sample text being greater than the maximum text length processed by the text processing model during training; Iteratively fine-tune the text processing model according to the sample text set to adjust model parameters of the text processing model; When an iterative fine-tuning termination condition is reached, the model parameter in the text processing model when the iterative fine-tuning termination condition is reached is determined as the position threshold.

9. The method according to claim 1, It is characterized in that The calling of the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information to generate semantically associated text of the target text includes: Calling the text processing model to perform a first semantic understanding on the target text according to the mapped position encoding bias information, and generating a semantically associated text unit corresponding to the first semantic understanding; The text composed of the semantically associated text unit corresponding to the first semantic understanding and the target text is used as a new target text, and the text processing model is continuously called to perform recursive semantic understanding until a recursive termination condition is reached; After the recursive termination condition is reached, the semantically associated text unit corresponding to the first semantic understanding and the text composed of the semantically associated text units obtained by the recursive semantic understanding are determined as the semantically associated text of the target text; The recursive termination condition is reached, including any one of the following: the predicted semantically associated text unit is a specified text unit, and the number of the predicted semantically associated text units reaches a specified number.

10. The method according to claim 1, It is characterized in that The obtaining of the position encoding bias information corresponding to the target text includes: Obtaining the relative positions of the two text units in the target text; The relative positions of the pairwise text units in the target text are calculated according to the relative position coefficients to obtain the relative position codes between the pairwise text units.

11. The method according to claim 10, It is characterized in that The text processing model includes a plurality of attention fusion modules, and the relative position encoding between the two text units in the target text is applied to the plurality of attention fusion modules; the coefficients corresponding to the plurality of attention fusion modules are the same or different; If each attention fusion module includes multiple attention fusion processes, the coefficients corresponding to each attention fusion process are different.

12. A text processing device, It is characterized in that include: An acquisition unit, used for acquiring a target text to be processed, wherein the target text includes a plurality of text units; The acquisition unit is further used to acquire position coding bias information corresponding to the target text, wherein the position coding bias information includes relative position coding between two text units in the target text, and the relative position coding between two text units is determined according to the relative positions of the two text units in the target text; A processing unit, used for calling a text processing model to map the first relative position code in the position code bias information to a position code range corresponding to the text processing model, so as to obtain a target relative position code corresponding to the first relative position code; The first relative position code is a relative position code in the position code bias information whose corresponding relative position exceeds a position threshold; The processing unit is also used to call the text processing model to perform semantic understanding on the target text according to the mapped position encoding bias information, and generate semantically associated text of the target text.

13. A computer device, It is characterized in that The computer device comprises: a processor suitable for implementing a computer program; A computer-readable storage medium storing a computer program, wherein the computer-readable storage medium stores a computer program, wherein the computer program is suitable for being loaded by the processor and executing the text processing method according to any one of claims 1 to 11.

14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the text processing method according to any one of claims 1 to 11.

15. A computer program product, It is characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the text processing method according to any one of claims 1 to 11 is implemented.