Text processing method and apparatus, and computer device, storage medium and program product

By mapping relative position coding in the text processing model, the problem of insufficient generalization performance when processing long text is solved by traditional position coding methods, and better generalization performance and length extrapolation capabilities are achieved.

WO2025102915A1PCT designated stage expired Publication Date: 2025-05-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2024/115952
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-08-30
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

When traditional position coding processes long text that has not been seen, the generalization performance is insufficient, resulting in poor inference effect of the model.

Method used

A text processing method is proposed, which improves the processing capability of the model for long text by obtaining relative position encoding information in the source text text and mapping these encodings into the position encoding range of the text processing model according to the distance threshold.

Benefits of technology

This method effectively improves the generalization performance and length extrapolation capabilities of the text processing model, so that the model can maintain a good text processing effect when processing long text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115952_22052025_PF_FP_ABST
    Figure CN2024115952_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A text processing method, which is executed by means of a computer device. The method comprises: acquiring source text to be processed, wherein the source text comprises a plurality of tokens (S501); on the basis of a relative distance between paired tokens in the source text, determining a relative position code corresponding to the relative distance between the paired tokens, and obtaining position code information comprising the relative position code (S502); acquiring a distance threshold value, determining from the position code information a first relative position code that corresponds to a relative distance exceeding the distance threshold value, and calling a text processing model to map the first relative position code in the position code information to a position code range of the text processing model, so as to obtain mapped position code information (S503); and calling the text processing model to perform semantic understanding on the source text on the basis of the mapped position code information, so as to generate semantically associated text for the source text (S504).
Need to check novelty before this filing date? Find Prior Art

Description

Text processing method, device, computer equipment, storage medium, and program product

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 2023115368068, filed on November 17, 2023, entitled “Text processing method, device, computer equipment, storage medium, program product,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of computer technology, in particular to the field of artificial intelligence technology, and specifically to a text processing method, a text processing apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0004] The Transformer model architecture (a language model structure) based on the attention mechanism has been widely used in the field of natural language processing in artificial intelligence and has become a popular language model architecture. However, the pure attention module cannot capture the order of individual text units (tokens) in the input text. This makes positional encoding particularly important in the Transformer model. Positional encoding is a positional representation method that uses encoding to represent the position of each text unit in the input text.

[0005] Currently, when the model uses traditional positional encoding methods for text processing, for some texts of seen text lengths (seen text lengths refer to text lengths that have appeared during model training), the text lengths are usually shorter, and good text processing effects can be achieved. However, for some texts of unseen text lengths (unseen text lengths refer to text lengths that have not appeared during the model training phase and exceed the lengths of texts that have appeared during training), the text lengths are usually longer, and the text processing effects are poor. In other words, the traditional positional encoding method has insufficient generalization performance, which means that when the model processes unseen inputs with text lengths that exceed the text lengths during training, it may have insufficient extrapolation capabilities, resulting in poor model reasoning effects.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide a text processing method, apparatus, computer equipment, storage medium, and program product.

[0008] In one aspect, an embodiment of the present application provides a text processing method, which is executed by a computer device and includes:

[0009] Acquire a source text to be processed, where the source text includes a plurality of text units;

[0010] Determining, based on relative distances between paired text units in the source text, relative position codes corresponding to the relative distances between the paired text units, and obtaining position code information including the relative position codes;

[0011] Obtaining a distance threshold, and determining, from the position code information, a first relative position code corresponding to a relative distance exceeding the distance threshold;

[0012] calling a text processing model to map the first relative position code in the position code information to a position code range of the text processing model to obtain mapped position code information; and

[0013] The text processing model is called to perform semantic understanding on the source text according to the mapped position coding information, and generate a semantically associated text of the source text.

[0014] On the other hand, an embodiment of the present application provides a text processing device, including:

[0015] An acquisition unit, configured to acquire a source text to be processed, wherein the source text includes a plurality of text units;

[0016] The acquisition unit is further configured to determine, based on the relative distances between the paired text units in the source text, relative position codes corresponding to the relative distances between the paired text units, and obtain position code information including the relative position codes;

[0017] a processing unit configured to obtain a distance threshold, determine a first relative position code corresponding to a relative distance exceeding the distance threshold from the position code information, invoke a text processing model, map the first relative position code in the position code information to a position code range of the text processing model, and obtain mapped position code information; and

[0018] The processing unit is further configured to call the text processing model, perform semantic understanding on the source text according to the mapped position coding information, and generate a semantically associated text of the source text.

[0019] In another aspect, an embodiment of the present application provides a computer device, comprising:

[0020] a processor suitable for implementing a computer program;

[0021] A computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the above-mentioned text processing method.

[0022] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is read and executed by a processor of a computer device, the computer device executes the above-mentioned text processing method.

[0023] In another aspect, embodiments of the present application provide a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the above-described text processing method.

[0024] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0026] FIG1 is a schematic diagram of calculating an attention weight provided in an embodiment of the present application;

[0027] FIG2 is a schematic diagram of a long text comprehension scenario provided by an embodiment of the present application;

[0028] FIG3 is a schematic diagram of a long text generation scenario provided in an embodiment of the present application;

[0029] FIG4 is a schematic diagram of a multi-round dialogue scenario provided by an embodiment of the present application;

[0030] FIG5 is a flow chart of a text processing method provided in an embodiment of the present application;

[0031] FIG6 is a schematic diagram of the structure of a text processing model provided in an embodiment of the present application;

[0032] FIG7 is a schematic diagram of a process of recursive semantic understanding provided by an embodiment of the present application;

[0033] FIG8 is a flow chart of another text processing method provided in an embodiment of the present application;

[0034] FIG9 is a comparative schematic diagram of a relative position encoding method provided in an embodiment of the present application;

[0035] FIG10 is a schematic diagram of the role of relative position coding in semantic understanding provided by an embodiment of the present application;

[0036] FIG11 is a schematic structural diagram of a text processing device provided in an embodiment of the present application;

[0037] FIG12 is a schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0039] In order to more clearly understand the technical solutions provided by the embodiments of the present application, some key terms involved in the embodiments of the present application are first introduced here:

[0040] Text processing can also be understood as text generation. In the embodiment of the present application, it refers to the process of calling a text processing model to perform semantic understanding on the text and generating semantically associated text based on the understood semantics. The text processing flow in the embodiment of the present application can generally include: first, the text processing model can be called to perform word segmentation processing on the text to obtain multiple text units (tokens) included in the text. The text unit refers to the language word (language word refers to the word of natural language) or language word (language word refers to the word of natural language) in the text. Secondly, the text processing model can be called to adopt an attention mechanism to determine the attention weights between paired text units in the text based on the similarity between paired text units in the text; the text processing model can be called to perform weighted summation processing on the corresponding text units in the text according to the attention weights between paired text units in the text to obtain the attention features of each text unit in the text. The attention features of the text units in the text can be understood as the semantic features of the text units and can be used to characterize the semantics of the text units. Then, the text processing model can be called to generate semantically associated text based on the semantic features of each text unit in the text. The text processing model can be a large language model in the field of natural language processing of artificial intelligence technology. A large language model refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of natural language text. For example, the text processing model can be a large prediction model based on the Transformer model structure.

[0041] This shows that the attention mechanism plays an important role in text processing. The attention mechanism is a network structure for sequence learning. In recent years, it has been commonly used in text processing to model language sequences for tasks such as text understanding and text generation. A common implementation of the attention mechanism can be seen in the following description:

[0042] Assume that text X is a text of length N. The text length refers to the number of text units contained in the text. In other words, text X contains N text units and is a sequence of text units (i.e., language words or tokens) of length N. Linearly mapping text X yields a query matrix (Q (query) matrix), a matching matrix (K (key) matrix), and a content matrix (V (value) matrix). The attention weight matrix can be expressed as the following formula 1:

[0043] In the above formula 1, AttSim(Q,K) represents the attention weight matrix; Q and K are two projection matrix operations, Q represents the query matrix, and K represents the matching matrix; d is a scaling factor used to maintain the stability of model training for text processing models. This attention weight matrix can be used to weight the content matrix V and output the attention feature Z (i.e., semantic feature Z). For details, please refer to the following formula 2: Z(X) = Attsim(Q(X),K(X)) * V(X) T Formula 2

[0044] In the above formula 2, V is the projection matrix operation and V represents the content matrix.

[0045] In particular, in the text processing process, natural language characters or words are generated word by word from left to right. Therefore, in the process of calculating the attention feature, each text unit should only calculate the similarity between the text unit and the previous text unit, and perform a weighted sum; the previous text unit of the current text unit may include the current text unit itself, as well as the text unit that is arranged before the current text unit in the text; generally speaking, this process can be introduced through the attention mask. In addition, the attention mechanism cannot capture the order of each text unit (token) in the text, and the order of each text unit in the text helps to understand the contextual semantics of the text unit. Therefore, it is necessary to introduce the relative position encoding between pairs of text units to represent the relative distance between pairs of text units. The following introduces the attention mask and relative position encoding in the attention mechanism respectively:

[0046] ①Attention mask:

[0047] For a text with a length of N, the attention mask of the text can be set to Mask∈R N*N , the matrix form is N*N, which is aligned with the matrix form of the attention weight matrix in the above formula 1. The attention mask defines the values ​​of its elements as follows:

[0048] In the above formula 3, i represents the i-th text unit in the text, and the i-th text unit is arranged at the i-th position in the text; j represents the j-th text unit in the text, and the j-th text unit is arranged at the j-th position in the text; mask i,j Represents the value of the element between the i-th text unit and the j-th text unit in the attention mask. This attention mask acts on the attention weight matrix, which can achieve the function of "each text unit only calculates the similarity with its previous text unit and takes the weighted sum". Therefore, the above formula 1 can be rewritten as the following formula 4:

[0049] In the above formula 4, AttSim_withMask(Q,K) represents the attention weight matrix updated based on the attention mask, and Mask represents the attention mask.

[0050] ②Relative position encoding:

[0051] Relative position encoding (using Alibi's position encoding as an example) can map the relative position relationship between paired text units in a text into a bias term in the form of an attention map (attention weight matrix), which acts on the attention map. Alibi relative position encoding introduces a bias term to represent the relative position relationship between paired text units in a text. The original Alibi relative position encoding and its role in attention weight are shown in Formula 5:

[0052] In the above formula 5, Represents the attention weight between the i-th text unit and each preceding text unit of the i-th text unit; Mask i Represents the value of the element corresponding to the i-th text unit in the attention mask; m*[-(i-1),…,-2,-1,0] represents the relative position encoding between the i-th text unit and each preceding text unit of the i-th text unit. For example, m*[-(i-1)] represents the relative position encoding between the i-th text unit and the first text unit; m is the relative distance coefficient, which is a fixed constant coefficient.

[0053] In summary, the attention mask and relative position encoding can be introduced into the original attention weights to update the attention weights, realize the function of calculating the similarity between each text unit in the text and its previous text unit, and consider the order of each text unit in the text in the calculation process of the attention mechanism to improve the semantic understanding ability of the attention mechanism. Taking a text with a length of 6 as an example, the influence of the attention mask and relative position encoding on the original attention weight can be seen in Figure 1. In the matrix on the left of Figure 1, q6k1 represents the attention weight between the 6th text unit and the 1st text unit in the text, q4k3 represents the attention weight between the 4th text unit and the 3rd text unit in the text, and so on, q1k1 represents the attention weight between the 1st text unit and the 1st text unit in the text; in the matrix on the left of Figure 1, the elements in the matrix are not complete, which is the influence of the attention mask. The attention mask ensures that the attention mechanism focuses on the attention weight between each text unit and its predecessor text unit, and does not focus on the attention weight between each text unit and its subsequent text unit. The subsequent text unit of any text unit refers to the text unit arranged after the text unit in the text. The matrix on the right side of Figure 1 represents the bias term, including the relative position encoding between pairs of text units. For example, the relative position encoding between the 6th text unit and the 1st text unit is m*[-(6-1)]=m*(-5); consistent with the attention weight matrix, the relative position encoding focuses on the relative position encoding between each text unit and its previous text unit, and does not focus on the relative position encoding between each text unit and its subsequent text unit.

[0054] Based on the above introduction to key terms such as text processing, attention mechanism, and relative position encoding, we can find that relative position encoding can be used to introduce the order between text units in the text into the attention mechanism, allowing the attention mechanism to capture the order between text units in the text. In the text processing process, improving the generalization performance of text processing models and enhancing their length extrapolation capabilities are urgent needs. Length extrapolation capabilities can be understood as the specific manifestation of generalization performance in terms of text length. Improving the generalization performance of text processing models and enhancing their length extrapolation capabilities are similar concepts. They refer to improving the text processing performance of text processing models trained on short sequences (short sequences refer to texts with shorter lengths) when processing long sequences (long sequences refer to texts with longer lengths), so that text processing models can also achieve good performance on long sequences.

[0055] Length extrapolation means that the length of the text processed by the text processing model during inference exceeds the length of the text processed by the text processing model during training. After the length extrapolation, the relative position encoding between paired text units in the text will also be extrapolated. Under the Alibi relative position encoding method, the relative position encoding between paired units in the text can be expressed as the following formula 6: AlibiBias(i,j)=-m*(ij) Formula 6

[0056] In the above formula 6, AlibiBias(i,j) represents the relative position encoding between the i-th text unit and the j-th text unit in the text, i represents the arrangement position of the i-th text unit in the text is the i-th position, j represents the arrangement position of the j-th text unit in the text is the j-th position, and i≥j.

[0057] Assume that the longest sequence length processed by the text processing model during training (the longest sequence length refers to the longest text length) is l1, and the longest sequence length required by actual business needs needs to be extrapolated to l2 (l2>l1). This embodiment of the application provides two extrapolation methods for relative position encoding:

[0058] The first method for extrapolating relative position encoding is direct extrapolation. Direct extrapolation involves directly calculating the relative position encoding of the text after length extrapolation based on Formula 6 above. In other words, it involves directly calculating the relative position encoding of the text after length extrapolation based on the Alibi relative position encoding method. In this method, the range of Formula 6 is [-m*(l1-1),0]. If we wish to extrapolate the longest sequence length to l2 during inference, we can directly calculate the relative position encoding of the text after length extrapolation based on Formula 6 above. The range of the extrapolated position encoding bias term is [-m*(l2-1),0].

[0059] The second relative position code extrapolation method is interpolation. This method involves retaining the range of [-m*(l1-1),0] and adjusting the interval between each relative position code to geometrically interpolate each relative position code to within the range of [-m*(l1-1),0]. The changes to the second relative position code extrapolation method are specifically reflected in the relative distance coefficient m, as shown in Formula 7 below:

[0060] That is to say, the second relative position coding extrapolation method modifies the original relative distance coefficient m into a relative distance coefficient with geometric interpolation function.

[0061] Neither of the two relative position encoding extrapolation methods mentioned above has achieved good results in improving the generalization performance and extrapolation capabilities of text processing models. The reasons are as follows: direct extrapolation causes the value range of relative position encodings to vary significantly, and the model's generalization ability for text units in the extrapolated part is insufficient; interpolation extrapolation shortens the intervals between relative position encodings, significantly affecting the distribution of attention weights between relatively close text units; and both relative position encoding extrapolation methods only consider the sequence length during training and directly extrapolate during inference, without considering the possibility of further fine-tuning the text processing model.

[0062] Based on this, the embodiment of the present application proposes a text processing method, which proposes an innovative extrapolation method of relative position coding on the basis of the extrapolation methods of the above two relative position coding methods. The extrapolation method of the innovative relative position coding is a new functional position coding method, which is an Alibi relative position coding method based on segmented interpolation and can be named LeakyAlibi relative position coding method. The relative position coding under this method has good extrapolation performance. Specifically, unlike the previous relative position coding, the text processing method proposed in the embodiment of the present application comprehensively considers the performance of relative position coding in unseen positions and the impact on the local attention mechanism; the new functional position coding method is a new segmented position coding bias term function. For two text units with close relative distances (i.e., similar tokens), the relative position coding interval between the two text units remains unchanged, retaining the distribution characteristics of the text processing model in local attention; for two text units with relatively large relative distances (i.e., remote tokens), the relative position coding is interpolated to control the value range of the position bias term, which effectively improves the generalization ability of the text processing model when the relative position coding is extrapolated. In addition, the text processing method proposed in the embodiment of the present application also introduces a parameter (this parameter can be called a distance threshold) to distinguish whether the relative distance between paired text units is close or not. Through this parameter, it can also be simply and effectively expanded to fine-tunable scenarios. During fine-tuning, this parameter can be set as a learnable model parameter, which can improve the scalability of the text processing model.

[0063] It should be noted that the text processing method proposed in the embodiment of the present application can be integrated into a text processing model, and the text processing method proposed in the embodiment of the present application can be executed by a computer device, which can be a terminal device or server deployed with a text processing model. Among them, the terminal device mentioned in the embodiment of the present application can include but is not limited to any of the following: smart phones, tablet computers, laptops, desktop computers, smart watches, smart home appliances, smart voice interaction devices, vehicle terminals and aircraft, but is not limited to these; the server mentioned in the embodiment of the present application can be a single physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The embodiment of the present application does not limit this.

[0064] The embodiments of this application do not limit the application scenarios of the text processing method. The text processing method can be applied to any application scenario that requires semantic understanding of long texts. For example, the text processing method proposed in the embodiments of this application can be applied to text processing scenarios such as long text understanding, long text generation, and multi-round dialogue. Specifically:

[0065] A long text understanding scenario refers to a text processing scenario in which semantic understanding of the text to be processed is performed to generate a semantic understanding result of the text to be processed. In other words, the generated semantic understanding result is a semantically associated text. For example, the semantic understanding result can be a text summary generated after semantic understanding of the text to be processed. For another example, the semantic understanding result can be a keyword extracted from the text to be processed after semantic understanding of the text to be processed. For another example, the semantic understanding result can be an information retrieval result of the semantics of the text to be processed after semantic understanding of the text to be processed. Figure 2 shows a schematic diagram of a long text understanding scenario. Taking the long text understanding scenario as a text summary generation scenario as an example, a text processing model is deployed in the terminal device used by the business object. The business object can submit a text summary generation task for the text to be processed in the terminal device. The text processing model deployed in the terminal device can output the generated text summary to the business object after semantic understanding of the text to be processed.

[0066] A long-text generation scenario refers to a text processing scenario in which semantically related text that meets text generation requirements is generated by performing semantic understanding on the text to be processed. For example, if the text to be processed specifies outline generation requirements, the semantically related text can be an outline text that meets the outline generation requirements. For another example, if the text to be processed specifies paper writing requirements, the semantically related text can be a paper text that meets the paper writing requirements. For another example, if the text to be processed specifies text translation requirements, the semantically related text can be a translated text that meets the text translation requirements. Figure 3 shows a scenario diagram of a long-text generation scenario. Taking the long-text generation scenario as an example of a text translation scenario, a business object can submit a text translation task to a server using a terminal device. The text translation requirement is to translate from a first language type to a second language type. A text processing model is deployed in the server. After performing semantic understanding on the text to be processed, the text processing model deployed in the server can generate a translated text that meets the text translation requirements and send the translated text to the terminal device. The terminal device can output the translated text that meets the text translation requirements to the business object.

[0067] Multi-turn dialogue scenarios refer to text processing scenarios that generate new dialogue text that is semantically related to historical dialogue text by semantically understanding historical dialogue text. In other words, semantically related text is new dialogue text that is semantically related to historical multi-turn dialogue text. For example, multi-turn dialogue scenarios specifically include multi-turn human-computer interaction scenarios and dialogue assistant scenarios. Multi-turn human-computer interaction scenarios and dialogue assistant scenarios are similar in that they both involve conversations with virtual robots. Figure 4 shows a scenario diagram of a multi-round dialogue scenario. A text processing model is deployed in the terminal device used by the business object. The business object can communicate with a virtual robot (for example, the virtual robot's dialogue name is human customer service) in the terminal device. The text processing model deployed in the terminal device semantically understands the dialogue content of the business object (for example, "Dialogue Text 1" in Figure 4), and outputs the dialogue content of the virtual robot to the business object in the tone of the virtual robot (for example, "Dialogue Text 2" in Figure 4); when the business object and the virtual robot generate multi-round dialogue content, the text processing model can perform a comprehensive semantic understanding of the multi-round historical dialogue content between the business object and the virtual robot, and output the dialogue content of the virtual robot to the business object in the tone of the virtual robot.

[0068] It is not difficult to find that the above text processing scenarios all require that the input and / or output text length of the text processing model is significantly longer than the ordinary question and answer content. This also requires the text processing model to be able to better handle ultra-long input and output sequences, and therefore puts forward higher requirements on the extrapolation of relative position encoding. The text processing method proposed in the embodiment of the present application is intended to enhance the text unit (token) length extrapolation capability of a large language model (i.e., a text processing model), so that the text processing model can achieve good performance without a lot of fine-tuning when processing longer input and output text tasks. The text processing method proposed in the embodiment of the present application can achieve better text processing effects in the above text processing scenarios; in the text understanding scenario, it can help the text processing model improve the comprehension ability of long text categories when the training text length is limited; in the text generation scenario, it can help the text processing model improve the generation ability of long text categories when the training text length is limited; in the multi-round dialogue scenario, it can help the text processing model improve the answer ability of long text categories when the training text length is limited.

[0069] The text processing method proposed in the embodiment of the present application is described in more detail below with reference to the accompanying drawings.

[0070] This embodiment of the present application proposes a text processing method. This text processing method mainly introduces the model structure of a text processing model and the text processing flow of the text processing model. This text processing method can be executed by a computer device that deploys the text processing model. The computer device can be, for example, a terminal device or a server. As shown in Figure 5, this text processing method may include but is not limited to steps S501 to S504:

[0071] S501: Obtain a source text to be processed, where the source text includes multiple text units.

[0072] The source text refers to any text to be processed. The source text may include multiple text units. A text unit refers to a natural language character or a natural language word in the source text. The text length of the source text is the second text length. The second text length refers to the number of text units contained in the source text.

[0073] Before introducing the specific text processing flow in conjunction with step S502-step S504 in the embodiment of the present application, the model structure of the text processing model is first introduced here, so as to facilitate the subsequent introduction of the text processing flow of the text processing model in conjunction with the model structure of the text processing model. As shown in Figure 6, the text processing model can be a model composed of a word segmentation layer (tokenize layer), a representation layer (embedding layer), multiple semantic understanding modules (transformer block) and a prediction layer. The connection relationship between the various components of the text processing model is as follows: the source text is input into the word segmentation layer, the output end of the word segmentation layer is connected to the input end of the representation layer, the output end of the representation layer is connected to the input end of the first semantic understanding module; the output end of the first semantic understanding module is connected to the input end of the second semantic understanding module, the output end of the second semantic understanding module is connected to the input end of the third semantic understanding module, and so on; the output end of the last semantic understanding module is connected to the input end of the prediction layer, and the output of the prediction layer can be used as the text processing result output by the text processing model. Among them:

[0074] ① Word Segmentation Layer: The word segmentation layer can be used to tokenize natural language input, converting it into a sequence of text units (tokens). Word tokenization refers to word segmentation. In other words, the word segmentation layer can be used to segment the input source text into multiple text units.

[0075] ② Representation Layer: Text unit (token) sequences are encoded into feature vectors (embeddings) through the representation layer. Specifically, the representation layer analyzes each text unit contained in the source text, obtaining a feature vector for each text unit. Each feature vector uniquely represents the corresponding text unit.

[0076] ③ Semantic Understanding Module: This module, also known as the Attention Fusion Module, is primarily implemented by the attention mechanism. The feature vectors (token embeddings) of text units are sequentially fed into multiple semantic understanding modules for attention fusion, enriching the semantics of the features. In other words, the multiple semantic understanding modules included in the text processing model can be used to perform semantic understanding (or attention fusion) on each text unit in the source text, obtaining the semantic features of each text unit in the source text.

[0077] By setting up multiple semantic understanding modules in the text processing model, the semantics of each text unit in the source text can be analyzed from shallow to deep. Specifically, after the feature vectors of each text unit in the source text are input into the first semantic understanding module, the first semantic understanding module can perform semantic understanding on each text unit and obtain the semantic understanding results of each text unit in the source text under the first semantic understanding module; after the semantic understanding results of each text unit in the source text under the first semantic understanding module are input into the second semantic understanding module, the second semantic understanding module can perform semantic understanding on each text unit and obtain the semantic understanding results of each text unit in the source text under the second semantic understanding module; and so on, the semantic understanding results of each text unit in the source text under the last semantic understanding module can be obtained, and the semantic understanding results of each text unit in the source text under the last semantic understanding module can be used as the final semantic features of each text unit in the source text. In the semantic understanding process of the above-mentioned multiple semantic understanding modules, the semantic understanding module that performs semantic understanding first outputs the shallow semantic features of the text unit, and the semantic understanding module that performs semantic understanding later outputs the deep semantic features of the text unit. Compared with the shallow semantic features, the deep semantic features are richer and have stronger text word distinction capabilities; moreover, the semantic understanding process of each semantic understanding module is the same, and they are all based on the attention mechanism for semantic understanding. For the specific semantic understanding process based on the attention mechanism, please refer to the relevant content of the above formulas 1 to 5.

[0078] The text processing method proposed in the embodiment of the present application involves the improvement of the relative position encoding between pairs of text units in the source text, and the relative position encoding is applied to the attention weight of the attention mechanism. Therefore, the improvement of the embodiment of the present application is mainly used in the semantic understanding module.

[0079] ④ Prediction layer: The semantic features of each text unit in the source text can ultimately be used to predict the feature vector of the next text unit (token) in the sequence. The feature vector of the next text unit (token) is restored to the form of natural language after decoding. In other words, the prediction layer can be used to predict the semantically associated text of the source text based on the semantic features of each text unit in the source text. Moreover, in the process of predicting the semantically associated text of the source text, the text processing model can only predict one semantically associated text unit at a time. The predicted semantically associated text unit needs to be spliced ​​into the input source text as the new source text, and a new semantically associated text unit is re-predicted. Subsequently, the new semantically associated text unit is recursively predicted until the recursive termination condition is reached. At this point, the predicted semantically associated text units can be spliced ​​together as the semantically associated text of the entire output.

[0080] Based on the model structure of the text processing model, the overall text processing process of the text processing model may include: calling the word segmentation layer of the text processing model to perform word segmentation processing on the source text to obtain multiple text units included in the source text; calling the representation layer of the text processing model to perform representation analysis on each text unit in the source text to obtain a feature vector of each text unit in the source text; calling multiple semantic understanding modules in the text processing model to perform semantic understanding on each text unit in the source text based on the feature vector of each text unit in the source text to obtain the semantic features of each text unit in the source text; calling the prediction layer in the text processing model to perform prediction processing on the source text based on the semantic features of each text unit in the source text to obtain predicted semantically associated text units. Then, the predicted semantically associated text units can be spliced ​​into the source text as the new source text, repeating the above text processing process to recursively predict new semantically associated text units until the recursive termination condition is reached, and outputting the text spliced ​​from all the predicted semantically associated text units as the semantically associated text of the source text.

[0081] Based on the model structure of the text processing model and the overall text processing process of the text processing model, the technical details of the text processing process are introduced below in conjunction with steps S502 to S504.

[0082] S502 : Determine relative position codes corresponding to the relative distances between paired text units in the source text, and obtain position code information including the relative position codes.

[0083] The position coding information corresponding to the source text can be obtained. The position coding information includes relative position coding between paired text units in the source text. The relative position coding between paired text units is determined based on the relative distance between the paired text units in the source text.

[0084] The position coding information corresponding to the source text may include relative position coding between pairs of text units in the source text. The relative position coding between pairs of text units may be determined based on the relative distances between the pairs of text units in the source text. The relative position coding between pairs of text units may be calculated based on the Alibi relative position coding method described in Formula 6 above, based on the relative distances between the pairs of text units in the source text.

[0085] In some embodiments, relative position codes corresponding to the relative distances between paired text units in a source text are determined, including: obtaining the relative distances between paired text units in the source text; obtaining relative distance coefficients, and calculating relative position codes between paired text units based on the relative distance coefficients and the relative distances between paired text units in the source text.

[0086] Specifically, the relative distance between paired text units in the source text specifically refers to the distance between paired text units in the source text, which is the difference between the position numbers of the arrangement positions of paired text units in the source text. For example, in the above formula 6, the i-th text unit is arranged at the i-th position in the source text with the position number i, and the j-th text unit is arranged at the j-th position in the source text with the position number j. Then the relative distance between the i-th text unit and the j-th text unit is (ij). The relative distance between paired text units in the source text can be calculated based on the relative distance coefficient m to obtain the relative position code between paired text units. For example, the relative position code between the i-th text unit and the j-th text unit is -m*(ij).

[0087] Furthermore, as described above, in the process of calculating attention features (i.e., semantic features), the attention mechanism focuses on the similarity between each text unit and its predecessor text unit. Correspondingly, the relative position encoding also considers the relative position encoding between each text unit and its predecessor text unit. That is, the position encoding information corresponding to the source text may include the relative position encoding between each text unit in the source text and its predecessor text unit. The predecessor text unit of each text unit includes each text unit itself, as well as the text unit arranged before each text unit in the source text.

[0088] In addition, the relative position encoding between pairs of text units in the source text can be applied to multiple attention fusion modules (i.e., semantic understanding modules) to participate in the calculation of attention weights, and the relative distance coefficients corresponding to multiple attention fusion modules can be the same or different; if the relative distance coefficients corresponding to each attention fusion module are the same, then the position encoding information introduced in each attention fusion module is the same, so there is no need to calculate the position encoding information separately for each attention fusion module, which can improve the text processing efficiency to a certain extent; if the relative distance coefficients corresponding to each attention fusion module are different, then the position encoding information introduced in each attention fusion module is different, so that different relative distance coefficients can be configured according to the semantic understanding requirements of each attention fusion module, thereby improving the semantic understanding accuracy of each attention fusion module, thereby improving the text processing accuracy to a certain extent.

[0089] Optionally, each attention fusion module may include multiple attention fusion processes, and the results of the multiple attention fusion processes are spliced ​​together to obtain the output result of each attention fusion module. It can be further understood that the attention mechanism adopted in each attention fusion module may be a multi-head attention mechanism (multi-head attention); the multi-head attention mechanism refers to mapping the input to different dimensions during the attention fusion process of each attention fusion module, performing an attention fusion process in each dimension, and finally splicing the attention fusion results of each dimension as the output result of each attention fusion module. In the multi-head attention mechanism, the relative distance coefficient corresponding to each attention fusion process may be different, that is, the relative distance coefficient m may take different values ​​in different heads. In this way, different relative distance coefficients can be configured according to the mapping requirements of different dimensions, thereby improving the semantic understanding accuracy of each dimension, thereby improving the text processing accuracy to a certain extent.

[0090] S503: Obtain a distance threshold, determine a first relative position code from the position code information whose corresponding relative distance exceeds the distance threshold, invoke a text processing model, and map the first relative position code in the position code information to a position code range within the text processing model to obtain mapped position code information.

[0091] The text processing model can be called to map the first relative position code in the position code information to the position code range of the text processing model to obtain the target relative position code corresponding to the first relative position code; the first relative position code is the relative position code in the position code information whose corresponding relative distance exceeds the distance threshold.

[0092] After obtaining the position coding information corresponding to the source text (including the relative position coding between paired text units in the source text), the text processing model can be called to map the first relative position coding in the position coding information to the position coding range of the text processing model to obtain the target relative position coding corresponding to the first relative position coding; the text processing model can be called to determine the second relative position coding in the position coding information as the target relative position coding corresponding to the second relative position coding; the first relative position coding is the relative position coding in the position coding information whose corresponding relative distance exceeds (exceeds means greater than) a distance threshold, and the second relative position coding is the relative position coding in the position coding information whose corresponding relative distance does not exceed (does not exceed means less than or equal to) the distance threshold. The mapped position coding information may include the target relative position coding corresponding to the first relative position coding and the target relative position coding corresponding to the second relative position coding.

[0093] The position coding range of the text processing model refers to the position coding range that enables the text processing model to achieve better text processing results. After the introduction of the attention mechanism, the relative position coding within the position coding range enables the text processing model to achieve better text processing results. The position coding range of the text processing model can be determined based on the first text length. The first text length refers to the longest text length processed by the text processing model during training. The position coding range of the text processing model refers to the value range of the relative position coding between paired text units in the text with the longest text length. For example, when the first text length is l1, the position coding range of the text processing model is [-m*(l1-1),0]. When the text length of the source text (i.e., the second text length) is greater than the first text length, the position coding range corresponding to the source text (i.e., the value range of the relative position coding between paired text units in the source text) exceeds the position coding range of the text processing model; for example, when the second text length is l2 (l2>l1), the position coding range of the text processing model is [-m*(l2-1),0], and [-m*(l2-1),0] exceeds [-m*(l1-1),0], which will affect the text processing effect of the text processing model. Therefore, the embodiment of the present application maps the relative position coding between paired text units in the source text to the position coding range of the text processing model, thereby improving the text processing effect of the text processing model on long texts (long text refers to texts whose length exceeds the maximum text length that has appeared during the training of the text processing model).

[0094] During the mapping process, an embodiment of the present application provides a custom parameter, which is a distance threshold with a value range of [0, l1]. The physical meaning of this parameter is the inflection point of the relative position encoding. When the relative distance between paired text units in the source text (i.e., the distance between paired text units in the source text) is less than or equal to the distance threshold, the relative position encodings between the paired text units remain unchanged, that is, they remain consistent with the relative position encodings obtained using the Alibi relative position encoding method. This allows the characteristics of the local attention mechanism to remain unchanged within the window m, reducing the impact on the local attention score and maintaining the dependency between adjacent text units. When the relative distance between paired text units in the source text (i.e., the distance between paired text units in the source text) is greater than the distance threshold and less than the expanded second text length l2, the position encoding range of the text processing model (i.e., the value range of the relative position encoding corresponding to the text processing model under the first text length l1) [-m*(l1-1),0] can be maintained unchanged, and the relative position encodings between the paired text units are mapped to the position encoding range of the text processing model, so that the value range of the relative position encoding after extrapolation remains consistent with that before extrapolation, reducing the impact of the unseen value range on the generalization of the relative position encoding.

[0095] S504 , calling a text processing model, performing semantic understanding on the source text according to the mapped position coding information, and generating a semantically associated text of the source text.

[0096] As previously described, a text processing model can only generate one semantically associated text unit per semantic understanding. Therefore, the text processing model needs to perform recursive semantic understanding, generating multiple semantically associated text units and concatenating them to form semantically associated text. Recursive semantic understanding involves continuously performing new semantic understandings based on the previous semantic understanding until the recursive termination condition is met. The text corresponding to the new semantic understanding is the concatenation of the text corresponding to the previous semantic understanding and the semantically associated text units generated by the previous semantic understanding.

[0097] Specifically, in combination with step S502-step S503, after obtaining the mapped position coding information, the text processing model can be called to perform a first semantic understanding of the source text according to the mapped position coding information, and generate a semantically associated text unit corresponding to the first semantic understanding; the text composed of the semantically associated text unit corresponding to the first semantic understanding and the source text is used as the new source text, and the text processing model is continued to be called for recursive semantic understanding until the recursive termination condition is reached; after the recursive termination condition is reached, the semantically associated text unit corresponding to the first semantic understanding and the text composed of the various semantically associated text units obtained by the recursive semantic understanding are determined as the semantically associated text of the source text.

[0098] Among them, reaching the recursive termination condition includes any of the following: the predicted semantically associated text unit is a specified text unit, for example, the specified text unit refers to a special end symbol; the number of predicted semantically associated text units reaches a specified number, for example, it is required to generate a text summary of 500 units. When the number of predicted semantically associated text units reaches 500, the recursive termination condition is reached.

[0099] For example, as shown in Figure 7, the recursive termination condition is reached after the text processing model is called to perform semantic understanding three times. The text processing model is called to perform the first semantic understanding of the source text to generate a first semantically associated text unit; after the source text and the first semantically associated text are spliced ​​into a new source text, the text processing model is called to perform the second semantic understanding of the new source text to generate a second semantically associated text unit; after the new source text and the second semantically associated text are spliced ​​into an updated source text, the text processing model is called to perform the third semantic understanding of the updated source text to generate a third semantically associated text unit; finally, the text obtained by splicing the first semantically associated text unit, the second semantically associated text unit, and the third semantically associated text unit can be output as the semantically associated text unit of the source text.

[0100] In an embodiment of the present application, when the text length of the source text exceeds the text length processed during the training of the text processing model, for the relative position coding between paired text units in the source text, when the relative distance corresponding to the relative position coding is greater than the distance threshold, the relative position coding can be mapped to the position coding range of the text processing model, and the source text is processed based on the relative position coding within the position coding range, so that the text processing model can maintain a good text processing effect. When the relative distance corresponding to the relative position coding is less than or equal to the distance threshold, the original relative position coding can be kept unchanged to avoid affecting the attention weight distribution between adjacent text units after mapping.

[0101] That is to say, the embodiment of the present application can expand the length of text processed by the text processing model while ensuring that the text processing effect of the text processing model is not affected, so that the text processing model can process text whose length exceeds the training text length.

[0102] Experiments have shown that compared with Alibi's relative position encoding method, the relative position encoding method proposed in the embodiment of the present application can effectively expand the context window of text length from 2k (2k refers to 2000 text units) to 8k (8k refers to 8000 text units), and the accuracy rate on the long text verification set is improved by more than 10%.

[0103] This embodiment of the present application proposes a text processing method, which mainly introduces the specific mapping process of relative position encoding and the setting method of distance threshold. The text processing method can be executed by a computer device that deploys a text processing model, and the computer device can be, for example, a terminal device or a server. As shown in Figure 8, the text processing method may include but is not limited to steps S801-S805:

[0104] S801: Obtain a source text to be processed, where the source text includes multiple text units.

[0105] In the embodiment of the present application, the execution process of step S801 is the same as the execution process of step S501 in the embodiment shown in Figure 5 above. For details, please refer to the relevant description of step S501 in the embodiment shown in Figure 5 above, which will not be repeated here.

[0106] S802: Determine relative position codes corresponding to the relative distances between paired text units in the source text, and obtain position code information including the relative position codes.

[0107] In the embodiment of the present application, the execution process of step S802 is the same as the execution process of step S502 in the embodiment shown in Figure 5 above. For details, please refer to the relevant description of step S502 in the embodiment shown in Figure 5 above, which will not be repeated here.

[0108] S803, obtaining a distance threshold, determining the first relative position code whose corresponding relative distance exceeds the distance threshold from the position code information; calling the text processing model, mapping the first relative position code in the position code information to the position code range of the text processing model to obtain the target relative position code.

[0109] In some embodiments, the source text includes N text units, the i-th text unit among the N text units is arranged at the i-th position in the source text, N is an integer greater than 1, and i is an integer less than or equal to N; the position coding information includes the relative position coding between the i-th text unit and each preceding text unit of the i-th text unit, and the preceding text unit is any text unit arranged at the 1st position to the i-th position in the source text. Calling the text processing model to map the first relative position coding in the position coding information to the position coding range of the text processing model to obtain the mapped position coding information includes: calling the text processing model to map the first relative position coding in the relative position coding between the i-th text unit and each preceding text unit to the target relative position coding corresponding to the first relative position coding within the position coding range of the text processing model to obtain the mapped position coding information.

[0110] For ease of understanding, here we take any text unit in the source text (any text unit can be represented as the i-th text unit) as an example to introduce the mapping process of the relative position encoding of the i-th text unit:

[0111] The source text may include N text units, any one of the N text units may be represented as the i-th text unit, the i-th text unit is arranged at the i-th position in the source text, N is an integer greater than 1, and i is an integer less than or equal to N; the position coding information may include the relative position coding between the i-th text unit and each preceding text unit of the i-th text unit, any preceding text unit refers to any text unit arranged at the 1st position to the i-th position in the source text, that is, the preceding text unit of the i-th text unit is the 1st text unit, the 2nd text unit, ..., the i-th text unit in the source text.

[0112] Based on this, the text processing model can be called to map the first relative position code in the relative position code between the i-th text unit and each previous text unit to the position code range of the text processing model to obtain the target relative position code corresponding to the first relative position code; wherein, the first relative position code is the relative position code in the relative position code between the i-th text unit and each previous text unit, and the corresponding relative distance exceeds the distance threshold.

[0113] Furthermore, the essence of mapping is interpolation, which is interpolation to the position coding range of the text processing model. For ease of understanding, the relative position coding between the i-th text unit and any predecessor text unit of the i-th text unit (any predecessor text unit of the i-th text unit can be expressed as the j-th predecessor text unit, j is a positive integer less than or equal to i) is taken as an example to introduce the mapping process of the relative position coding between the i-th text unit and the j-th predecessor text unit. In detail, when the relative distance between the i-th text unit and the j-th preceding text unit exceeds a distance threshold, an interpolation function can be obtained; the interpolation function is generated based on the second text length, the first text length, and the distance threshold, the second text length refers to the text length of the source text, the text length of the source text refers to the number of text units included in the source text, the first text length is the maximum text length among the text lengths of the text processing model processed during the training process, and the value range of the interpolation function is the position coding range of the text processing model; then, according to the interpolation function, the first relative position coding between the i-th text unit and the j-th preceding text unit can be interpolated to obtain the target relative position coding corresponding to the first relative position coding between the i-th text unit and the j-th preceding text unit.

[0114] S804: Determine, from the position code information, a second relative position code corresponding to a relative distance that does not exceed a distance threshold; invoke a text processing model to determine the second relative position code in the position code information as a target relative position code corresponding to the second relative position code, thereby obtaining mapped position code information. The mapped position code information includes the target relative position code obtained by mapping the first relative position code, and the target relative position code corresponding to the second relative position code.

[0115] For ease of understanding, here we take any text unit in the source text (any text unit can be represented as the i-th text unit) as an example to introduce the mapping process of the relative position encoding of the i-th text unit:

[0116] Calling a text processing model to determine the second relative position code in the position code information as the target relative position code corresponding to the second relative position code, including: calling a text processing model to determine the second relative position code in the relative position code between the i-th text unit and each preceding text unit as the target relative position code corresponding to the second relative position code.

[0117] The text processing model can be called to determine the second relative position code in the relative position code between the i-th text unit and each previous text unit as the target relative position code corresponding to the second relative position code; wherein the second relative position code is a relative position code in the relative position code between the i-th text unit and each previous text unit, in which the relative distance does not exceed the distance threshold.

[0118] In summary, the mapping process of the relative position coding of the i-th text unit from step S803 to step S804 can be summarized as the following formula 8 (formula 8 can also be expressed as a piecewise function form of the following formula 9):

[0119] In the above formulas 9 and 10, represents the target relative position encoding between the i-th text unit and the j-th preceding text unit, l1 represents the length of the first text, l2 represents the length of the second text, It is known from the above formula 8 and formula 9 that when the relative distance between the i-th text unit and the j-th preceding text unit (i.e. the distance between the i-th text unit and the j-th preceding text unit) is greater than the distance threshold When the first relative position code [-m*(ij)] between the i-th text unit and the j-th preceding text unit is mapped to the position code range of the text processing model (i.e., the value range [-m*(l1-1),0] corresponding to the text processing model), the actual mapping is the target relative position code When the relative distance between the i-th text unit and the j-th preceding text unit (i.e., the distance between the i-th text unit and the j-th preceding text unit) is less than or equal to the distance threshold When , the second relative position code [-m*(ij)] between the i-th text unit and the j-th preceding text unit can be kept unchanged.

[0120] Figure 9 shows a schematic diagram of the comparison between the relative position coding under the Alibi relative position coding method (which can be understood as before the text length is extrapolated) and the relative position coding under the LeakyAlibi relative position coding method proposed in the embodiment of the present application (which can be understood as after the text length is extrapolated). It can be seen that: first, the relationship between the relative position coding and the relative distance is the same in the Alibi relative position coding method and the LeakyAlibi relative position coding method. The greater the relative distance between the text units (that is, the farther the distance), the greater the absolute value of the relative position coding. Conversely, the smaller the relative distance between the text units (that is, the closer the distance), the smaller the absolute value of the relative position coding. Second, under the Alibi relative position coding method and the LeakyAlibi relative position coding method, the value range of the relative position coding remains consistent. That is to say, the value range of the relative position coding remains consistent before and after the text length is extrapolated, reducing the influence of the unseen value range on the generalization of the relative position coding and improving the text processing effect of the text processing model after the text length is extrapolated. Third, at the inflection point (the inflection point refers to the distance threshold ), the relative position encoding in Alibi relative position encoding and LeakyAlibi relative position encoding is the same, which has little effect on the local attention weight and maintains the dependency between adjacent text units; at the inflection point (distance threshold ), the relative position coding under the LeakyAlibi relative position coding method is interpolated into the relative position coding value range before the text length is extrapolated, and the text length supported by the relative position coding is expanded while ensuring the text processing effect of the text processing model.

[0121] S805 , calling the text processing model, performing semantic understanding on the source text according to the mapped position coding information, and generating a semantically associated text of the source text.

[0122] For ease of understanding, we take any text unit in the source text (any text unit can be represented as the i-th text unit) as an example to introduce the process of how the relative position encoding of the i-th text unit participates in semantic understanding:

[0123] The mapped position coding information may include the target relative position coding between the i-th text unit and each preceding text unit; the first attention weight between the i-th text unit and each preceding text unit may be determined based on the similarity between the i-th text unit (specifically, the element value of the i-th text unit in the Q matrix) and each preceding text unit (specifically, the element value of the j-th preceding text unit in the K matrix); the first attention weight between the i-th text unit and each preceding text unit may be updated based on the target relative position coding between the i-th text unit and each preceding text unit to obtain the second attention weight between the i-th text unit and each preceding text unit; and based on the second attention weight between the i-th text unit and each preceding text unit, each preceding text unit (specifically, the element value of the preceding text unit in the V matrix) is weightedly summed to obtain the semantic feature corresponding to the i-th text unit.

[0124] For example, as shown in Figure 10, the source text includes four text units, namely the first text unit, the second text unit, the third text unit and the fourth text unit. Here, taking the third text unit as an example, the process of determining the semantic features of the third text unit based on the relative position encoding of the third text unit is introduced. First, based on the similarity between the third text unit and the first text unit, the first attention weight q3k1 between the third text unit and the first text unit can be determined. Similarly, the first attention weight q3k2 between the third text unit and the second text unit can be obtained, and the first attention weight q3k3 between the third text unit and the third text unit can be obtained. Secondly, based on the target relative position encoding between the third text unit and the first text unit The first attention weight q3k1 between the third text unit and the first text unit is updated to obtain the second attention weight q3k′1 between the third text unit and the first text unit. Similarly, the second attention weight q3k′2 between the third text unit and the second text unit can be obtained, and the second attention weight q3k′3 between the third text unit and the third text unit can be obtained. Then, based on the obtained second attention weights, the first text unit, the second text unit, and the third text unit can be weighted and summed (q3k′1v1+q3k′2v2+q3k′3v3) to obtain the semantic features of the third text unit, where v1 represents the element value of the first text unit in the V matrix, v2 represents the element value of the second text unit in the V matrix, and v3 represents the element value of the third text unit in the V matrix.

[0125] Following the aforementioned method for determining the semantic features of the i-th text unit, the semantic features of all text units in the source text other than the i-th text unit can be determined. The method for determining the semantic features of other text units can be found in the method for determining the semantic features of the i-th text unit and will not be further elaborated here. After determining the semantic features of each text unit in the source text, semantic understanding of the source text can be performed based on the semantic features of each text unit in the source text to generate semantically associated text for the source text.

[0126] Specifically, semantic understanding may refer to the recursive semantic understanding described in step S504 of the embodiment shown in FIG5 . Specifically, a text processing model may be called to perform a first semantic understanding of the source text based on the mapped position coding information, generating a semantically associated text unit corresponding to the first semantic understanding; the text composed of the semantically associated text unit corresponding to the first semantic understanding and the source text is used as a new source text, and the text processing model is continuously called to perform recursive semantic understanding until a recursive termination condition is met; after the recursive termination condition is met, the text composed of the semantically associated text unit corresponding to the first semantic understanding and the various semantically associated text units obtained by the recursive semantic understanding is determined as the semantically associated text of the source text.

[0127] It is worth noting that the text processing model is a pre-trained large language model with text processing capabilities (text processing capabilities here refer to the ability to semantically understand the input text and generate semantically associated text of the input text). Therefore, the text processing model can be directly used for reasoning. Direct use for reasoning means directly using the text processing model for text processing; or, if resources are sufficient to support secondary fine-tuning of the text processing model, the text processing model can be fine-tuned according to the task requirements of the text processing task.

[0128] For the usage of the above two text processing models (i.e. direct inference or secondary fine-tuning), the distance threshold can be Specifically, the distance threshold can be obtained in any of the following ways:

[0129] First, for direct inference of the text processing model, the maximum text length among the text lengths processed by the text processing model during training can be obtained; and a distance threshold can be set according to the first text length. Specifically, the distance threshold is set according to the first text length, where the first text length is the maximum text length among the text lengths processed by the text processing model during training. For example, if the first text length is l1, the distance threshold can be set as Set to l1 / 2. In this case, the distance threshold is set based on the empirical value Setting it to l1 / 2 allows the text processing model to perform well when processing long texts without a lot of fine-tuning, achieving better text processing results.

[0130] Second, for the case of secondary fine-tuning of the text processing model, the distance threshold It can be set as a learnable model parameter and adjusted to a relatively ideal value during the fine-tuning process of the text processing model. This relatively ideal value can enable the text processing model to achieve better text processing results when processing long texts. In addition, the initial value of this learnable model parameter can be set to l1 and restricted to the interval [0, l1] during the fine-tuning process. In other words, an initial threshold (the initial threshold refers to the initial value l1) can be obtained and used as a model parameter of the text processing model. During the fine-tuning process of the text processing model, it is adjusted to obtain the distance threshold.

[0131] In detail, the process of using the initial threshold as a model parameter of the text processing model and adjusting it during the fine-tuning process of the text processing model to obtain the distance threshold can include: first, obtaining a sample text set, the sample text set including multiple sample texts, the text length of each sample text being greater than the maximum text length (i.e., the first text length) among the text lengths of the texts processed by the text processing model during the training process. Moreover, if the text processing task specifies a text length, the text length of each sample text in the sample text set can be the specified text length, so that the fine-tuned text processing model and the distance threshold can better adapt to the task requirements of the text processing task. Secondly, the text processing model can be iteratively fine-tuned based on the sample text set to adjust the model parameters of the text processing model; when adjusting the model parameters of the text processing model, the model parameters serving as the distance threshold can be adjusted without adjusting other model parameters, or the model parameters serving as the distance threshold and other model parameters can be adjusted together. Then, when the iterative fine-tuning termination condition is reached, the model parameters in the text processing model when the iterative fine-tuning termination condition is reached are determined as the distance threshold.

[0132] Furthermore, iterative fine-tuning refers to fine-tuning the text processing model multiple times, with the subsequent fine-tuning being performed based on the previous fine-tuning. The sample text used in any fine-tuning process during the iterative fine-tuning process can be represented as a reference text, and the sample text set can also include marked semantically associated text of the reference text; any fine-tuning process during the iterative fine-tuning process can include: first, obtaining a reference text, which can include multiple sample text units; and obtaining position coding information corresponding to the reference text, which can include relative position coding between pairs of sample text units in the reference text, where the relative position coding between pairs of sample text units is determined based on the relative distance between the paired sample text units in the reference text. Secondly, the text processing model can be called to map the first relative position coding in the position coding information corresponding to the reference text to the position coding range of the text processing model, thereby obtaining a target relative position coding corresponding to the first relative position coding; the first relative position coding is the relative position coding in the position coding information corresponding to the reference text whose corresponding relative distance exceeds a distance threshold. Then, the text processing model can be called to perform semantic understanding of the reference text based on the mapped position encoding information to generate the actual semantically associated text of the reference text; the loss information of the text processing model can be determined based on the difference between the actual semantically associated text of the reference text and the marked semantically associated text of the reference text, and the model parameters of the text processing model can be adjusted based on the loss information of the text processing model.

[0133] It should be noted that the conditions for terminating iterative fine-tuning may include any of the following: the number of fine-tuning times included in the iterative fine-tuning reaches a number threshold, and the loss information of the text processing model is less than the loss threshold. In addition, the fine-tuning process of the text processing model is similar to the reasoning process of the text processing model. The steps in the fine-tuning process of the text processing model that are similar to the reasoning process of the text processing model are not described here. For details, please refer to the reasoning process of the text processing model. By fine-tuning the distance threshold, it is adjusted to a relatively ideal value during the fine-tuning process of the text processing model. This relatively ideal value can, on the one hand, enable the text processing model to achieve better text processing effects when processing long texts, and on the other hand, enable the text processing model to better meet the task requirements of the text processing task.

[0134] In an embodiment of the present application, when the text length of the source text exceeds the text length processed during the training of the text processing model, for the relative position coding between paired text units in the source text, when the relative distance corresponding to the relative position coding is greater than the distance threshold, the relative position coding can be mapped to the position coding range of the text processing model, and the source text is processed based on the relative position coding within the position coding range, so that the text processing model can maintain a good text processing effect; when the relative distance corresponding to the relative position coding is less than or equal to the distance threshold, the original relative position coding can be kept unchanged to avoid affecting the attention weight distribution between adjacent text units after mapping. In other words, the embodiment of the present application can expand the text length processed by the text processing model under the premise of ensuring that the text processing effect of the text processing model is not affected, so that the text processing model can process texts whose text length exceeds the training text length. In addition, by adopting different setting methods for the distance threshold under different usage scenarios of the text processing model (direct inference or secondary fine-tuning), for the usage scenario of direct inference of the text processing model, the distance threshold is set according to the empirical value, so that the text processing model can achieve good performance when processing long texts without a lot of fine-tuning, and achieve better text processing effects; for the usage scenario of secondary fine-tuning of the text processing model, during the fine-tuning process of the text processing model, the distance threshold is adjusted as a learnable model parameter, so that the adjusted distance threshold can enable the text processing model to achieve better text processing effects when processing long texts.

[0135] The above describes in detail the method of the embodiment of the present application. In order to facilitate better implementation of the above scheme of the embodiment of the present application, the device of the embodiment of the present application is provided below accordingly.

[0136] Please refer to Figure 11, which is a schematic diagram of the structure of a text processing device provided in an embodiment of the present application. The text processing device can be installed in a computer device provided in an embodiment of the present application, and the computer device can be a terminal device or a server. The text processing device shown in Figure 11 can be a computer program running on the computer device. The text processing device can be used to perform some or all of the steps in the method embodiments shown in Figures 5 or 8. Please refer to Figure 12, the text processing device can include the following units:

[0137] The acquisition unit 1101 is configured to acquire a source text to be processed, where the source text includes multiple text units.

[0138] The acquisition unit 1101 is further configured to determine relative position codes corresponding to the relative distances between paired text units in the source text, and obtain position code information including the relative position codes.

[0139] Processing unit 1102 is used to obtain a distance threshold, determine the first relative position code whose corresponding relative distance exceeds the distance threshold from the position code information; call the text processing model, map the first relative position code in the position code information to the position code range of the text processing model, and obtain the mapped position code information.

[0140] The processing unit 1102 is further configured to call a text processing model to perform semantic understanding on the source text according to the mapped position coding information, and generate a semantically associated text of the source text.

[0141] In some embodiments, the processing unit 1102 is further used to perform the following steps: determining, from the position coding information, a second relative position coding whose corresponding relative distance does not exceed a distance threshold; calling a text processing model to determine the second relative position coding in the position coding information as a target relative position coding corresponding to the second relative position coding; wherein the mapped position coding information includes the target relative position coding obtained after mapping the first relative position coding, and the target relative position coding corresponding to the second relative position coding.

[0142] In some embodiments, the source text includes N text units, the i-th text unit among the N text units is arranged at the i-th position in the source text, N is an integer greater than 1, and i is an integer less than or equal to N; the position coding information includes the relative position coding between the i-th text unit and each preceding text unit of the i-th text unit, and the preceding text unit is any text unit arranged at the 1st position to the i-th position in the source text.

[0143] Processing unit 1102 is used to call the text processing model, map the first relative position code in the relative position code between the i-th text unit and each previous text unit to the target relative position code corresponding to the first relative position code within the position code range of the text processing model, and obtain the mapped position code information.

[0144] In some embodiments, the processing unit 1102 is used to call the text processing model to determine the second relative position code in the relative position code between the i-th text unit and each previous text unit as the target relative position code corresponding to the second relative position code.

[0145] In some embodiments, the relative distance between the i-th text unit and the j-th preceding text unit exceeds a distance threshold, where j is a positive integer less than or equal to i; a processing unit 1102 is configured to obtain an interpolation function; the interpolation function is generated based on a first text length, a second text length, and a distance threshold, where the first text length refers to the maximum text length among the text lengths of the texts processed by the text processing model during training; the second text length refers to the text length of the source text, where the text length of the source text refers to the number of text units included in the source text; the value range of the interpolation function is the position coding range of the text processing model; and the interpolation function is configured to interpolate the first relative position coding between the i-th text unit and the j-th preceding text unit according to the interpolation function to obtain a target relative position coding corresponding to the first relative position coding between the i-th text unit and the j-th preceding text unit.

[0146] In some embodiments, the mapped position coding information includes the target relative position coding between the i-th text unit and each preceding text unit; the processing unit 1102 is used to determine the first attention weight between the i-th text unit and each preceding text unit based on the similarity between the i-th text unit and each preceding text unit; based on the target relative position coding between the i-th text unit and each preceding text unit, the first attention weight between the i-th text unit and each preceding text unit is updated to obtain the second attention weight between the i-th text unit and each preceding text unit; based on the second attention weight between the i-th text unit and each preceding text unit, each preceding text unit is weightedly summed to obtain the semantic feature corresponding to the i-th text unit; based on the semantic features of each text unit in the source text, the source text is semantically understood to generate a semantically associated text of the source text.

[0147] In some embodiments, the processing unit 1102 is configured to obtain a maximum text length among text lengths processed by the text processing model during training; and set a distance threshold according to the first text length.

[0148] In some embodiments, the processing unit 1102 is configured to obtain an initial threshold value, use the initial threshold value as a model parameter of the text processing model, and adjust the initial threshold value during the fine-tuning process of the text processing model to obtain a distance threshold value.

[0149] In some embodiments, the processing unit 1102 is used to obtain a sample text set, which includes multiple sample texts, and the text length of each sample text is greater than the maximum text length of the text processed by the text processing model during the training process; based on the sample text set, the text processing model is iteratively fine-tuned to adjust the model parameters of the text processing model; when the iterative fine-tuning termination condition is reached, the model parameters in the text processing model when the iterative fine-tuning termination condition is reached are determined as the distance threshold.

[0150] In some embodiments, the processing unit 1102 is used to call the text processing model, perform a first semantic understanding of the source text according to the mapped position coding information, and generate a semantically associated text unit corresponding to the first semantic understanding; use the text composed of the semantically associated text unit corresponding to the first semantic understanding and the source text as the new source text, and continue to call the text processing model for recursive semantic understanding until the recursive termination condition is reached; after the recursive termination condition is reached, the semantically associated text unit corresponding to the first semantic understanding and the text composed of the various semantically associated text units obtained by the recursive semantic understanding are determined as the semantically associated text of the source text.

[0151] In some embodiments, reaching the recursive termination condition includes any one of the following: the predicted semantically associated text unit is a specified text unit, or the number of the predicted semantically associated text units reaches a specified number.

[0152] In some embodiments, the acquisition unit 1101 is used to obtain the relative distance between paired text units in the source text; obtain the relative distance coefficient; and the processing unit 1102 is used to calculate the relative position code between the paired text units based on the relative distance coefficient and the relative distance between the paired text units in the source text.

[0153] In some embodiments, the text processing model includes multiple attention fusion modules, and the relative position encoding between pairs of text units in the source text is applied to multiple attention fusion modules; if each attention fusion module includes multiple attention fusion processes, the coefficients corresponding to each attention fusion process are different.

[0154] According to another embodiment of the present application, the various units in the text processing device shown in Figure 11 can be individually or all merged into one or several other units to form, or one (or some) of the units can be further divided into multiple functionally smaller units to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the blockchain-based data processing device can also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by collaboration of multiple units.

[0155] According to another embodiment of the present application, a text processing apparatus as shown in FIG11 can be constructed and the text processing method of the embodiment of the present application can be implemented by running a computer program capable of executing some or all of the steps involved in the method shown in FIG5 or FIG8 on a general-purpose computing device such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). The computer program can be recorded on a computer-readable storage medium, for example, and loaded into the computing device via the computer-readable storage medium and executed therein.

[0156] In an embodiment of the present application, when the text length of the source text exceeds the text length that appears during the training process of the text processing model, the position coding range to which the relative position coding between the paired text units in the source text belongs will exceed the position coding range of the text processing model, which will result in poor text processing effect of the text processing model. In this case, the embodiment of the present application maps the relative position coding whose relative distance exceeds the distance threshold to the position coding range of the text processing model, so that when the text processing model performs text processing on the source text based on the mapped relative position coding, it can achieve better text processing effect; that is, when the embodiment of the present application performs text processing on the source text whose text length exceeds the training text length (that is, the text length that appears during the training process of the text processing model), by mapping the relative position coding between the paired text units in the source text to the position coding range of the text processing model, the text processing model can achieve better text processing effect for the source text whose text length exceeds the training text length, thereby improving the generalization performance of the text processing model.

[0157] Based on the above method and device embodiments, an embodiment of the present application provides a computer device. Please refer to Figure 12, which is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device shown in Figure 12 includes at least a processor 1201, an input interface 1202, an output interface 1203, and a computer-readable storage medium 1204. The processor 1201, input interface 1202, output interface 1203, and computer-readable storage medium 1204 may be connected via a bus or other means.

[0158] Computer-readable storage medium 1204 may be stored in a memory of a computer device. Computer-readable storage medium 1204 is used to store a computer program, which includes computer instructions. Processor 1201 is used to execute the computer program stored in computer-readable storage medium 1204. Processor 1201 (or CPU (Central Processing Unit)) is the computing and control core of the computer device and is suitable for implementing computer programs, specifically loading and executing computer programs to implement corresponding method flows or corresponding functions.

[0159] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the operating system of the computer device. In addition, a computer program suitable for being loaded and executed by the processor is also stored in the storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk storage; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.

[0160] The computer device may be a terminal device or a server. In a specific implementation, the processor 1201 may load and execute a computer program stored in a computer-readable storage medium 1204 to implement the corresponding steps in the text processing method shown in FIG5 or FIG8. In a specific implementation, the processor 1201 loads the computer program in the computer-readable storage medium 1204 and executes the following steps: obtaining a source text to be processed, the source text including multiple text units; determining relative position codes corresponding to the relative distances between paired text units based on relative distances between paired text units in the source text, and obtaining position code information including the relative position codes; obtaining a distance threshold, and determining, from the position code information, a first relative position code corresponding to a relative distance exceeding the distance threshold; calling a text processing model to map the first relative position code in the position code information to a position code range of the text processing model, and obtaining mapped position code information; and calling the text processing model to perform semantic understanding on the source text based on the mapped position code information, and generating semantically associated text of the source text.

[0161] In an embodiment of the present application, when the text length of the source text exceeds the text length that appears during the training process of the text processing model, the position coding range to which the relative position coding between the paired text units in the source text belongs will exceed the position coding range of the text processing model, which will result in poor text processing effect of the text processing model. In this case, the embodiment of the present application maps the relative position coding whose relative distance exceeds the distance threshold to the position coding range of the text processing model, so that when the text processing model performs text processing on the source text based on the mapped relative position coding, it can achieve better text processing effect; that is, when the embodiment of the present application performs text processing on the source text whose text length exceeds the training text length (that is, the text length that appears during the training process of the text processing model), by mapping the relative position coding between the paired text units in the source text to the position coding range of the text processing model, the text processing model can achieve better text processing effect for the source text whose text length exceeds the training text length, thereby improving the generalization performance of the text processing model.

[0162] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described text processing method.

[0163] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0164] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0165] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0166] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0167] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A text processing method, executed by a computer device, comprising: Acquire a source text to be processed, wherein the source text includes a plurality of text units; Determine, according to the relative distances between the paired text units in the source text, relative position codes corresponding to the relative distances between the paired text units, and obtain position code information including the relative position codes; Obtaining a distance threshold, and determining, from the position code information, a first relative position code whose corresponding relative distance exceeds the distance threshold; Calling a text processing model, mapping the first relative position code in the position code information to a position code range of the text processing model, and obtaining mapped position code information; and The text processing model is called to perform semantic understanding on the source text according to the mapped position encoding information, and to generate semantically associated text of the source text.

2. The method of claim 1, further comprising: Determine, from the position code information, a second relative position code corresponding to which the relative distance does not exceed the distance threshold; Calling the text processing model to determine the second relative position code in the position code information as a target relative position code corresponding to the second relative position code; The mapped position coding information includes a target relative position coding obtained by mapping the first relative position coding and a target relative position coding corresponding to the second relative position coding.

3. The method according to claim 2, wherein the source text comprises N text units, the i-th text unit among the N text units is arranged at the i-th position in the source text, N is an integer greater than 1, and i is an integer less than or equal to N; the position encoding information comprises a relative position encoding between the i-th text unit and each preceding text unit of the i-th text unit, and the preceding text unit is any text unit arranged at the 1st position to the i-th position in the source text; The calling of the text processing model, mapping the first relative position code in the position code information to the position code range of the text processing model, and obtaining the mapped position code information, comprises: The text processing model is called to map the first relative position code in the relative position code between the i-th text unit and each of the preceding text units to a target relative position code corresponding to the first relative position code within the position code range of the text processing model to obtain the mapped position code information.

4. The method according to claim 3, wherein calling the text processing model to determine the second relative position code in the position code information as the target relative position code corresponding to the second relative position code comprises: The text processing model is called to determine the second relative position code in the relative position code between the i-th text unit and each of the preceding text units as the target relative position code corresponding to the second relative position code.

5. The method according to claim 3 or 4, wherein the relative distance between the i-th text unit and the j-th preceding text unit exceeds a distance threshold, and j is a positive integer less than or equal to i; The method further comprises: Obtain an interpolation function; the interpolation function is generated according to a first text length, a second text length and the distance threshold, the first text length refers to the maximum text length of the texts processed by the text processing model during the training process; the second text length refers to the text length of the source text, and the text length of the source text refers to the number of text units included in the source text; the value range of the interpolation function is the position coding range of the text processing model; The mapping of the first relative position code in the relative position code between the i-th text unit and each of the preceding text units to a target relative position code corresponding to the first relative position code within the position code range of the text processing model comprises: According to the interpolation function, the first relative position code between the i-th text unit and the j-th preceding text unit is interpolated to obtain the first relative position code between the i-th text unit and the j-th preceding text unit. The target relative position encoding corresponding to the position encoding.

6. The method according to any one of claims 3 to 5, wherein the mapped position coding information includes the target relative position coding between the i-th text unit and each of the preceding text units; and the calling of the text processing model to perform semantic understanding on the source text according to the mapped position coding information to generate semantically associated text of the source text comprises: Determine a first attention weight between the i-th text unit and each of the preceding text units according to the similarity between the i-th text unit and each of the preceding text units; According to the target relative position encoding between the i-th text unit and each of the preceding text units, a first attention weight between the i-th text unit and each of the preceding text units is updated to obtain a second attention weight between the i-th text unit and each of the preceding text units; According to the second attention weight between the i-th text unit and each of the preceding text units, weighted sum processing is performed on each of the preceding text units to obtain a semantic feature corresponding to the i-th text unit; According to the semantic features of each text unit in the source text, the source text is semantically understood to generate a semantically associated text of the source text.

7. The method according to any one of claims 1 to 6, further comprising: Obtaining a maximum text length among text lengths processed by the text processing model during training; The distance threshold is set according to the first text length.

8. The method according to any one of claims 1 to 6, further comprising: An initial threshold is obtained, and the initial threshold is used as a model parameter of the text processing model. The initial threshold is adjusted during the fine-tuning process of the text processing model to obtain the distance threshold.

9. The method according to claim 8, wherein the initial threshold is used as a model parameter of the text processing model and is adjusted during the fine-tuning process of the text processing model to obtain the distance threshold, comprising: Acquire a sample text set, the sample text set comprising a plurality of sample texts, the text length of each sample text being greater than the maximum text length of texts processed by the text processing model during training; Iteratively fine-tune the text processing model according to the sample text set to adjust model parameters of the text processing model; When the iterative fine-tuning termination condition is reached, the model parameter in the text processing model when the iterative fine-tuning termination condition is reached is determined as the distance threshold.

10. The method according to any one of claims 1 to 9, wherein calling the text processing model to perform semantic understanding on the source text according to the mapped position encoding information to generate semantically associated text of the source text comprises: Calling the text processing model to perform a first semantic understanding on the source text according to the mapped position encoding information, and generating a semantically associated text unit corresponding to the first semantic understanding; The text composed of the semantically associated text unit corresponding to the first semantic understanding and the source text is used as a new source text, and the text processing model is continuously called to perform recursive semantic understanding until a recursive termination condition is reached; After the recursive termination condition is reached, the semantically associated text unit corresponding to the first semantic understanding and the text composed of the semantically associated text units obtained by the recursive semantic understanding are determined as the semantically associated text of the source text.

11. According to the method of claim 10, the recursive termination condition is reached, including any one of the following: the predicted semantically associated text unit is a specified text unit, or the number of predicted semantically associated text units reaches a specified number.

12. The method according to any one of claims 1 to 11, wherein determining the relative position code corresponding to the relative distance between the paired text units in the source text comprises: Obtaining relative distances between pairs of text units in the source text; A relative distance coefficient is obtained, and a relative position code between the paired text units is calculated based on the relative distance coefficient and the relative distance between the paired text units in the source text.

13. According to the method described in any one of claims 1 to 12, the text processing model includes multiple attention fusion modules, and the relative position encoding between the pairs of text units in the source text is applied to the multiple attention fusion modules; if each attention fusion module includes multiple attention fusion processes, the coefficients corresponding to each attention fusion process are different.

14. A text processing device, comprising: An acquisition unit, used for acquiring a source text to be processed, wherein the source text includes a plurality of text units; The acquisition unit is further configured to determine, based on the relative distances between the paired text units in the source text, relative position codes corresponding to the relative distances between the paired text units, and obtain position code information including the relative position codes; A processing unit, configured to obtain a distance threshold, and determine, from the position code information, a first relative position code corresponding to a relative distance exceeding the distance threshold; calling a text processing model to map the first relative position code in the position code information to a position code range of the text processing model to obtain mapped position code information; and The processing unit is further used to call the text processing model, perform semantic understanding on the source text according to the mapped position encoding information, and generate semantically associated text of the source text.

15. A computer device, comprising: a processor suitable for implementing a computer program; A computer-readable storage medium storing a computer program, wherein the computer program is suitable for being loaded by the processor and executing the text processing method according to any one of claims 1 to 13.

16. A computer-readable storage medium storing a computer program, wherein the computer program is suitable for being loaded by a processor and executing the text processing method according to any one of claims 1 to 13.

17. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the text processing method according to any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Position coding method and device, electronic equipment and storage medium

    CN115796127A

  • Punctuation prediction method and device, equipment and storage medium

    CN117034863A

  • Large scale retrieval for sequence generation

    US20230177334A1

Cited By

  • Music understanding coding model construction method and device

    CN116129916A

  • A method and device for constructing a music understanding coding model

    CN116129916B