Method and device for carrying out position coding on extensible text, and electronic equipment

By flexibly encoding the extensible text in the form of flexible functions, the problem of insufficient universality and robustness of the position encoding method in the prior art in long text processing is solved, and efficient performance improvements on texts of different lengths are achieved.

CN120337994APending Publication Date: 2025-07-18启元实验室
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510403697.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing position coding methods have poor versatility and robustness when dealing with long text, especially when the model is degraded when it is extrapolated from a short text window to a long text window.

Method used

A flexible function form is used to encode the extensible text position. By determining the dimensions and interpolation multiples of the text vector, the sigmoid function is used to adjust the interpolation multiples of each dimension to realize the position encoding of the extensible text.

Benefits of technology

It improves adaptability to texts of different lengths, can naturally capture relative positional relationships, and improves the performance of the model in long text scenarios, especially when the context length is greatly expanded, without significantly increasing computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337994A_ABST
    Figure CN120337994A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for carrying out position coding on an extensible text and electronic equipment, and the method comprises the steps: determining the dimension of a text vector corresponding to the extensible text according to the extensible text; determining an interpolation multiple corresponding to each dimension in the dimensions according to a preset minimum dimension difference multiple and a preset maximum dimension interpolation multiple; and performing position coding on the extensible text by using the interpolation multiple. According to the embodiment of the invention, the position information of the extensible text is coded in a flexible function form, compared with a traditional position coding mode, the relative position relation in the extensible text can be more naturally captured, the method can adapt to application scenes of texts with different lengths, the method can be applied to language models with different scales, and the user experience is improved. The problem that a current position coding method is poor in universality and robustness is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of e-commerce. Specifically, it relates to a method and device for position encoding of extensible text, an electronic device, and a non-transitory computer-readable storage medium. Background Art

[0002] Since the Transformer architecture was proposed in 2017, it has quickly become the mainstream technology in the field of natural language processing and is widely used in multiple fields such as machine translation, text generation, and speech recognition. This model architecture, with its unique self-attention mechanism and parallelized computing method, overcomes the limitations of traditional recurrent neural networks in dealing with long-distance dependence problems, thus significantly improving the performance of the model in handling complex language tasks.

[0003] Currently, many mainstream open-source large model architectures are based on the Transformer architecture. These models demonstrate strong generalization ability in various natural language processing tasks through large-scale pre-training. However, a major challenge of the Transformer is that its computational complexity grows quadratically with the increase in the length of the input sequence, which makes the computational resource requirements extremely high when dealing with long texts and the training efficiency is low.

[0004] To address this problem, researchers have proposed a commonly used training strategy, that is, first pre-train on a shorter text length (such as 4k) to obtain an LLM (Large Language Model, abbreviated as large language) model with a context window length of 4k. Subsequently, through continued pre-training, fine-tune the LLM model on a longer text length (for example, 32k) to obtain an LLM with a context window length of 32k. This strategy is called length extrapolation. Its feasibility lies in that the number of training texts required for continued pre-training is much lower than that required in the initial pre-training stage. Therefore, by performing a small number of steps of continued pre-training on the short-text model, the ability of the model on the short-text window can be effectively generalized to the long-text window. This method significantly reduces the consumption of computational resources while retaining the model performance.

[0005] Position encoding is a core component in the Transformer architecture, which is used to provide the position information of each element in the embedded sequence, enabling the model to perceive the order of the input sequence. Traditional Transformer uses a combination of sine and cosine functions to generate position encoding, which is simple but effective. However, as the model is applied to long texts, the standard position encoding method gradually exposes its limitations. For this reason, researchers have introduced Rotary Position Encoding (RoPE), which captures the relative position information in the text more naturally by rotating vectors.

[0006] Since the proposed RoPE has been widely adopted in many large language models due to its excellent performance in processing long texts. The usual pre-training setting is to set the base frequency of RoPE to 10,000 and train the model on texts of 4k length. However, theoretical studies have shown that when RoPE with a base frequency of 10,000 is extended to a longer context window, its corresponding wavelength cannot effectively cover the extended context window, which leads to a decline in the performance of the model when performing length extrapolation.

[0007] To solve this problem, many variants of RoPE have emerged, such as position interpolation (PI), adaptive base frequency (ABF), NTK, and YaRN and other position encoding methods. These variants introduce scaling and adjustment of position encoding in their respective designs in order to maintain the performance of the model in long text scenarios. These methods show certain advantages in specific scenarios, but there is still room for improvement in terms of generality and robustness. Summary of the Invention

[0008] The present application aims to propose a method and device for position encoding of scalable texts, an electronic device, and a non-transitory computer-readable storage medium to solve the problem of poor generality and robustness of current position encoding methods.

[0009] According to one aspect of the present application, a method for position encoding of scalable texts is proposed, including:

[0010] Determine the dimension of the corresponding text vector according to the scalable text;

[0011] Determine the interpolation multiple corresponding to each dimension in the dimension according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple; and

[0012] Perform position encoding on the scalable text using the interpolation multiple.

[0013] According to some embodiments, before determining the dimension of the corresponding text vector according to the scalable text, it further includes:

[0014] Perform vector quantization encoding on the scalable text to obtain the text vector.

[0015] According to some embodiments, determining the interpolation multiple corresponding to each dimension in the dimension according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple includes:

[0016] Determine the interpolation multiple corresponding to each dimension in the dimension through a preset function according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple.

[0017] According to some embodiments, the preset function includes a sigmoid function.

[0018] According to some embodiments, in the step of performing positional encoding on the scalable text using the interpolation multiple, positional encoding is performed by the following formula:

[0019]

[0020] where x is the text vector, t is the t-th vector among the N text vectors corresponding to the scalable text, j is the dimension, d is the dimension length, b is a preset fundamental frequency, α j is the scaling factor corresponding to the j-th dimension position, m is the attention scaling factor in the self-attention model, and i is the imaginary number.

[0021] According to some embodiments, the method further includes:

[0022] Calculating the evaluation index scores of the scalable text within different ranges of the interpolation multiple for each dimension;

[0023] Using the evaluation index scores within different ranges of the interpolation multiple for each dimension to calculate the average value of the evaluation index scores.

[0024] According to some embodiments, calculating the evaluation index scores of the scalable text within different ranges of the interpolation multiple for each dimension includes:

[0025] Using the RuLES framework to calculate the evaluation index scores of the scalable text within different ranges of the interpolation multiple for each dimension.

[0026] According to one aspect of the present application, there is provided an apparatus for performing positional encoding on a scalable text, including:

[0027] A dimension determination unit, configured to determine the dimension of the text vector corresponding to the scalable text according to the scalable text;

[0028] An interpolation multiple determination unit, configured to determine the interpolation multiple corresponding to each dimension in the dimension according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple; and

[0029] A positional encoding unit, configured to perform positional encoding on the scalable text using the interpolation multiple.

[0030] According to one aspect of the present application, there is provided an electronic device, including: a processor; a memory for storing a computer program; when the computer program is executed by the processor, the processor is caused to implement the method as described in any of the previous embodiments.

[0031] According to one aspect of the present application, a non-transitory computer-readable storage medium is provided, on which computer-readable instructions are stored. When the instructions are executed by a processor, the processor is caused to execute the method described in any of the previous embodiments.

[0032] According to the embodiments of the present application, by encoding the position information of extensible text in a flexible function form, compared with traditional position encoding methods, it can capture the relative position relationship in extensible text more naturally, can adapt to application scenarios of texts with different lengths, can be applied to language models of different scales, and solves the problem that the current position encoding methods have poor generality and robustness.

[0033] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application. Brief Description of the Drawings

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. By referring to the drawings and describing its exemplary embodiments in detail, the above and other objectives, features, and advantages of the present application will become more obvious.

[0035] Figure 1 The flowchart of a method for position encoding of extensible text according to an exemplary embodiment of the present application is shown.

[0036] Figure 2a The schematic diagram showing the relationship between dimensions and interpolation multiples in a position encoding method PI is shown.

[0037] Figure 2b The schematic diagram showing the relationship between dimensions and interpolation multiples in a position encoding method ABF is shown.

[0038] Figure 2c The schematic diagram showing the relationship between dimensions and interpolation multiples in a position encoding method NTK is shown.

[0039] Figure 2d The schematic diagram showing the relationship between dimensions and interpolation multiples in a position encoding method YaRN is shown.

[0040] Figure 2e The schematic diagram showing the relationship between dimensions and interpolation multiples in a position encoding method LongRoPE is shown.

[0041] Figure 2f The position encoding S 3 PE showing the relationship between dimensions and interpolation multiples according to an exemplary embodiment of the present application is shown.

[0042] Figure 3aShows a schematic diagram of the relationship between the dimension and the interpolation multiple in another position encoding method PI.

[0043] Figure 3b Shows a schematic diagram of the relationship between the dimension and the interpolation multiple in another position encoding method ABF.

[0044] Figure 3c Shows a schematic diagram of the relationship between the dimension and the interpolation multiple in another position encoding method NTK.

[0045] Figure 3d Shows a schematic diagram of the relationship between the dimension and the interpolation multiple in another position encoding method YaRN.

[0046] Figure 3e Shows a schematic diagram of the relationship between the dimension and the interpolation multiple in another position encoding method LongRoPE.

[0047] Figure 3f Shows another position encoding S according to the exemplary embodiments of the present application 3 PE, a schematic diagram of the relationship between the dimension and the interpolation multiple.

[0048] Figure 4 Shows a flowchart of another method for position encoding scalable text according to the exemplary embodiments of the present application.

[0049] Figure 5 Shows a schematic diagram of the training parameters of the training model according to the exemplary embodiments of the present application.

[0050] Figure 6 Shows a schematic diagram of the average index results of the NIAH index in the calculation RuLES framework under different position encodings.

[0051] Figure 7 Shows a schematic diagram of the score comparison results of the NIAH index in the calculation RuLES framework under different position encodings and when the text vector length is 16k - 32k.

[0052] Figure 8 Shows a block diagram of a device for position encoding scalable text according to the exemplary embodiments of the present application.

[0053] Figure 9 Shows an electronic device according to the exemplary embodiments of the present application. Detailed implementation manners

[0054] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.

[0055] The features, structures, or characteristics described may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will recognize that the technical solutions of the present disclosure may be practiced without one or more of these specific details, or may be practiced using other methods, components, materials, devices, or operations, etc. In such cases, well-known structures, methods, devices, implementations, materials, or operations will not be shown or described in detail.

[0056] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily include all of the content and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0057] The terms "first", "second", etc. in the description, claims, and above-mentioned drawings of this application are used to distinguish different objects and not to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.

[0058] Before describing the embodiments of this application, the terms that appear in this application will be explained first.

[0059] Position Interpolation (PI): A method of mapping between different text lengths by scaling position encoding.

[0060] Rotary Position Embedding (RoPE), by rotating vectors, to capture relative position information in text.

[0061] Adaptive Base Frequency (ABF for short): By dynamically adjusting the base frequency in RoPE, the positional encoding can better adapt to texts of different lengths.

[0062] Neural Tangent Kernel (NTK for short): Based on the theoretical positional encoding method, it designs a more accurate encoding form by simulating the learning behavior of neural networks.

[0063] YaRN: An improved method based on RoPE positional encoding. It further adjusts RoPE to address challenges in long text processing.

[0064] LongRoPE: A specific implementation in YaRN, aiming to perform interpolation on texts of different lengths to enhance the generalization ability of the model.

[0065] S3PE: A positional interpolation method proposed in the embodiments of this application.

[0066] RULER: A test framework for evaluating the performance of large language models in handling complex tasks, especially for evaluating the performance of models in complex tasks such as long text processing, reasoning, and precise information retrieval.

[0067] Needle in a Haystack (NIAH for short): One of the core tasks of RULER, aiming to evaluate the ability of the model to accurately locate key information in long text scenarios through specific subtasks.

[0068] The following will describe in detail the specific embodiments according to this application with reference to the accompanying drawings.

[0069] Figure 1 A flowchart of a method for positional encoding of extensible text according to an exemplary embodiment of this application is shown. As Figure 1 shown, the method includes steps S101, S103, and S105. The following will refer to Figure 1 to describe in detail a method for positional encoding of extensible text according to an exemplary embodiment of this application.

[0070] As Figure 1 shown, in step S101, according to the extensible text, determine the dimension of its corresponding text vector.

[0071] According to the embodiments of this application, before step S101, it is necessary to perform vectorized encoding on the extensible text to obtain a text vector.

[0072] Vectorized encoding represents text information as vectors that can express the semantics of the text, using numerical vectors to represent the semantics of the text.

[0073] It should be noted here that this application does not limit the method of vectorized encoding for extensible text. In specific embodiments, the methods for vectorized encoding of input text include, but are not limited to, one-hot encoding vectorization, bag-of-words model vectorization, term frequency-inverse document frequency model vectorization, n-gram model vectorization, word-vector model vectorization, and / or document-vector model.

[0074] In step S103, according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple, determine the interpolation multiple corresponding to each dimension in the dimensions.

[0075] According to an embodiment of the present application, each dimension and the interpolation multiple can be defined by a preset function, so that according to the preset minimum dimension difference multiple and the maximum dimension interpolation multiple, the interpolation multiple corresponding to each dimension in the dimensions can be determined through the preset function.

[0076] In a specific embodiment, the preset function includes the sigmoid function, so that according to the minimum dimension difference multiple corresponding to the preset minimum dimension and the maximum dimension interpolation multiple corresponding to the maximum dimension, the corresponding sigmoid function can be determined.

[0077] In step S105, use the interpolation multiple to perform position encoding on the extensible text.

[0078] According to an embodiment of the present application, in step S105, position encoding is performed through formula (1).

[0079]

[0080] Among them, x is the text vector, t is the t-th vector among the N text vectors corresponding to the extensible text, j is the dimension, d is the dimension length, b is the preset base frequency, α j is the scaling factor corresponding to the j-th dimension position, m is the attention scaling factor in the self-attention model, i is the imaginary number, and the interpolation multiple = 1 / α j .

[0081] According Figure 1 to the shown embodiment, by encoding the position information of the extensible text in a flexible function form, compared with the traditional position encoding method, it can capture the relative position relationship in the extensible text more naturally, can adapt to the application scenarios of texts of different lengths, can be applied to language models of different scales, and solves the problem that the current position encoding method has poor generality and robustness.

[0082] Figures 2a - 2f andFigures 3a - 3f It shows a schematic diagram of the relationship between the dimension (the abscissa in the figure) and the interpolation multiple (the ordinate in the figure) according to the prior art and the exemplary embodiments of the present application. Among them, Figures 2a - 2f It is used to compare the relationship schematic diagram of the dimension and the interpolation multiple in the position encoding methods (including PI, ABF, NTK, YaRN, LongRoPE, and S 3 PE) when the maximum difference multiple is 8. Figures 3a - 3f It is used to compare the relationship schematic diagram of the dimension and the interpolation multiple in the position encoding methods (including PI, ABF, NTK, YaRN, LongRoPE, and S3PE) when the maximum difference multiple is 50. Among them, S 3 PE is the relationship schematic diagram of the dimension and the interpolation multiple when the preset function according to the exemplary embodiments of the present application is the sigmoid function.

[0083] According to Figures 2a - 2f and Figures 3a - 3f the results shown, for the position encoding methods NTK, YaRN, and LongRoPE, the smoothness corresponding to different position encodings in the figure is different. For example, NTK, YaRN, and LongRoPE all make a trade-off between the original RoPE and PI. That is, the low dimension is close to RoPE, and the high dimension is close to PI, only the interpolation strategies are different. However, the smoothness of different position encodings is different. For example, in the image of YaRN, there are two turning points, namely low and high. Before low, the original RoPE is adopted, and after high, PI is adopted. And for the S 3 PE position encoding method of the exemplary embodiments of the present application, when the dimension j is small, α j tends to 1 to retain as much position information as possible; when j is large, α j correspondingly decreases to make the text vector at position t interpolate as much as possible within the range seen during training.

[0084] Figure 4 It shows a flowchart of another method for position encoding of extensible text according to the exemplary embodiments of the present application. And Figure 1 compared with Figure 4 the method shown, in addition to including steps S101, S103, and S105, also includes steps S107 and S109. For simplicity, only the differences from Figure 1 are described here, and the same parts will not be elaborated again.

[0085] As Figure 4 shown, in step S107, calculate the evaluation index scores of the extensible text within different dimension interpolation multiple ranges.

[0086] According to an embodiment of the present application, the RuLES framework is used to calculate the evaluation index scores of the scalable text within different ranges of interpolation multiples in different dimensions.

[0087] In step S109, the average value of the evaluation index scores is calculated by using the evaluation index scores within different ranges of interpolation multiples in different dimensions.

[0088] To comprehensively evaluate the performance of different positional encoding schemes in long text processing, in this embodiment, five models of different sizes are used, and each model is preliminarily pre-trained on a text of length 4k to lay a foundation for subsequent continued pre-training. In this way, it is ensured that the basic language modeling capabilities of each model on shorter texts are relatively consistent. On this basis, we continue to pre-train these models on a dataset of longer texts (such as text vectors of length 32k) to observe the performance of different positional encoding schemes when dealing with longer texts. The goal of the experiment is to systematically compare the extrapolation capabilities of different positional encodings when the model is extended to handle long text tasks.

[0089] As Figure 5 shown are five models of different sizes for training, and the scales of these models gradually increase from small to large. Specifically, each model is first pre-trained on a text of length 4k to help the model learn basic language representation capabilities, thus laying a good foundation for subsequent long text processing. In the continued pre-training stage, a dataset containing 32k text vectors is used and mixed according to the characteristics of the data. For example, the data is mixed by the Per-Source Upsampling technique to balance data from different sources, thereby avoiding biases in the model training results caused by uneven data distribution.

[0090] During the training process, the first two stages of the WSD (Warmup, Static, Decay) learning rate strategy are adopted. In this strategy, the learning rate is first linearly increased from 0 to the set peak learning rate and then remains unchanged.

[0091] In terms of model evaluation, to more accurately evaluate the performance of the model in long text scenarios, the RuLES framework is used to calculate the evaluation metric scores of the scalable text within different ranges of interpolation multiples in different dimensions. Among them, NIAH, as the core task of the RuLES framework, includes sub-tasks such as single-shot retrieval and multi-shot retrieval, and can effectively simulate the difficulty of accurately locating key information in long texts. Specifically, the NIAH task requires the model to find one or more target information given a long text. Such tasks have relatively high requirements for the model's long text understanding ability, so it is an ideal tool for evaluating the performance of different position encoding schemes in long text tasks. For each text interval of different lengths, in this implementation, the scores of all sub-tasks of NIAH are averaged, and this is used as a comprehensive evaluation metric for the model performance.

[0092] Figure 6 Shows a schematic diagram of the average metric results of the NIAH metric in the RuLES framework under different position encodings. Figure 7 Shows a schematic diagram of the score comparison results of the NIAH metric in the RuLES framework under different position encodings and when the text vector length is 16k - 32k.

[0093] In this embodiment, different position encoding schemes (including ABF, PI, NTK, YaRN, LongRoPE, and S 3 PE) are compared in terms of their effects when processing ultra-long texts at maximum interpolation multiples of 8 times and 50 times.

[0094] It should be particularly noted that except for ABF, most of the existing position encoding schemes (such as PI, NTK, YaRN, LongRoPE) are closely related to the context expansion multiple. Therefore, we marked these existing expansion multiples with "[]" in the table, and also provided a control group to observe the changes in the position encoding scheme when expanding by different multiples.

[0095] As Figure 6 and Figure 7 shown, among all position encoding schemes, when the model scale is not less than model s2, the model with an expansion multiple of 50 times generally performs better than the model with an expansion multiple of 8 times in long text tasks with a length of 16k - 32k. That is to say, within a reasonable range, the larger the expansion multiple, the better the model's performance in long text processing. This also indicates that increasing the expansion multiple can help the model better understand and process ultra-long context information. As the model scale increases, the performance difference between different position encoding schemes gradually shrinks. Especially in larger models, although different position encodings perform differently when processing long texts, the overall difference is not as significant as in smaller models. However, according to the S of the embodiments of the present application 3PE maintains high performance in all position encoding schemes. Especially at an expansion multiple of 50, it significantly leads other encoding schemes by 2 to 3 points. On the premise of the same expansion multiple, different position encoding schemes essentially correspond to different expansion shapes. Although different encoding schemes perform differently in their respective tasks, according to the S of the embodiments of this application 3 PE performs stably in long text tasks. Whether in a relatively long sequence interval or the overall average score, it reaches or approaches the optimal state.

[0096] Thus, it can be seen that position encoding is crucial for the long text processing ability of the model, especially when the context window length is greatly expanded. As the model scale increases, the relative importance of position encoding decreases.

[0097] According to the S of the embodiments of this application 3 PE encoding can capture relative position information in the text more naturally, thus improving the model's understanding ability, especially when the context length increases significantly. And it has obvious advantages in terms of scalability and robustness. Experiments show that under different model scales, S 3 PE encoding always maintains the optimal or near-optimal performance in long text tasks. Whether it is short text or ultra-long text processing, S 3 PE can handle it well. Therefore, S 3 PE encoding has strong adaptability and can be widely applied to language models of different scales.

[0098] In addition, according to the S of the embodiments of this application 3 PE encoding does not require a significant increase in model parameters or computational complexity, but can significantly improve performance in long text scenarios. Therefore, it provides an efficient solution for the application of large-scale language models. Especially when computational resources are limited, excellent results can be achieved through a small amount of pre-training.

[0099] The above mainly introduced the embodiments of this application from the perspective of methods. Those skilled in the art should easily realize that, combined with the operations or steps of each example described in the embodiments disclosed herein, this application can be implemented in the form of hardware or a combination of hardware and computer software. Those skilled in the art can use different methods to implement the described functions for each specific operation or method, and such implementation should not be considered to exceed the scope of this application.

[0100] Next, the device embodiments of this application are described. For details not described in the device embodiments of this application, reference can be made to the method embodiments of this application.

[0101] Figure 8The block diagram of a device for performing position encoding on extensible text according to an exemplary embodiment of the present application is shown. As Figure 8 The device shown includes a dimension determination unit 801, an interpolation multiple determination unit 803, and a position encoding unit 805. Among them, the dimension determination unit 801 is configured to determine the dimension of the corresponding text vector according to the extensible text; the interpolation multiple determination unit 803 determines the interpolation multiple corresponding to each dimension in the dimension according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple; and the position encoding unit 805 is configured to perform position encoding on the extensible text by using the interpolation multiple.

[0102] Figure 9 An electronic device according to an exemplary embodiment of the present application is shown. Referring below to Figure 9 describe the electronic device 200 according to this embodiment of the present application. Figure 9 The electronic device 200 shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0103] As Figure 9 shown, the electronic device 200 is presented in the form of a general-purpose computing device. The components of the electronic device 200 may include, but are not limited to: at least one processing unit 210, at least one storage unit 220, a bus 230 connecting different system components (including the storage unit 220 and the processing unit 210), a display unit 240, etc.

[0104] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 210, so that the processing unit 210 executes the methods according to various exemplary embodiments of the present application described in this specification. For example, the processing unit 210 can execute the method as Figure 1 shown in.

[0105] The storage unit 220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 2201 and / or a cache storage unit 2202, and may further include a read-only storage unit (ROM) 2203.

[0106] The storage unit 220 may further include a program / utility 2204 having a set (at least one) of program modules 2205. Such program modules 2205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0107] The bus 230 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of the various bus structures.

[0108] The electronic device 200 may also communicate with one or more external devices 300 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 200, and / or may communicate with any device that enables the electronic device 200 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be performed through the input / output (I / O) interface 250. Moreover, the electronic device 200 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 260. The network adapter 260 may communicate with other modules of the electronic device 200 through the bus 230. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0109] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. The technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiments of the present application.

[0110] The software product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0111] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0112] The program code for performing the operations of this application may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0113] The above computer-readable medium carries one or more programs, which when executed by one of the devices, cause the computer-readable medium to implement the foregoing functions.

[0114] Those skilled in the art can understand that the above-mentioned modules may be distributed in the device according to the description of the embodiments, or may be correspondingly changed and distributed in one or more devices that are uniquely different from the embodiments. The modules of the above embodiments may be combined into one module, or may be further split into multiple sub-modules.

[0115] According to an embodiment of the present application, a computer program is provided, including a computer program or instruction, which when executed by a processor, can perform the method described above.

[0116] The above has introduced the embodiments of the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, based on the idea of the present application, any changes or deformations made by those skilled in the art in the specific implementation manner and application scope of the present application fall within the protection scope of the present application. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A method for position encoding of extensible text, characterized in that, Including: Determine the dimension of the text vector corresponding to the extensible text; Determine the interpolation multiple corresponding to each dimension in the dimension according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple; and Perform positional encoding on the extensible text by using the interpolation multiple.

2. The method according to claim 1, characterized in that, Before determining the dimension of the text vector corresponding to the extensible text according to the extensible text, it further includes: Perform vector quantization encoding on the extensible text to obtain the text vector.

3. The method according to claim 1, characterized in that, Determine the interpolation multiple corresponding to each dimension in the dimension according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple, including: Determine the interpolation multiple corresponding to each dimension in the dimension through a preset function according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple.

4. The method according to claim 3, wherein The preset function includes a sigmoid function.

5. The method according to claim 1, wherein In the step of performing positional encoding on the extensible text by using the interpolation multiple, positional encoding is performed through the following formula: where x is the text vector, t is the t-th vector among the N text vectors corresponding to the extensible text, j is the dimension, d is the dimension length, b is the preset fundamental frequency, α j is the scaling factor corresponding to the j-th dimension position, m is the attention scaling factor in the self-attention model, and i is the imaginary number.

6. The method according to claim 1, wherein It further includes: Calculate the evaluation index scores of the extensible text within different dimension interpolation multiple ranges; Calculate the average value of the evaluation index scores by using the evaluation index scores within different dimension interpolation multiple ranges.

7. The method according to claim 6, characterized in that, Calculate the evaluation index scores of the extensible text within different dimension interpolation multiple ranges, including: Calculate the evaluation index scores of the extensible text within different dimension interpolation multiple ranges by using the RuLES framework.

8. An apparatus for position encoding of extensible text, characterized in that, Including: A dimension determination unit for determining the dimension of the text vector corresponding to the extensible text; An interpolation multiple determination unit for determining the interpolation multiple corresponding to each dimension in the dimension according to a preset minimum dimension difference multiple and a maximum dimension interpolation multiple; and A positional encoding unit for performing positional encoding on the extensible text by using the interpolation multiple.

9. An electronic device, characterized in that, Including: A processor; A memory for storing a computer program; When the computer program is executed by the processor, the processor is caused to implement the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium having stored thereon computer-readable instructions, which when executed by a processor, cause the processor to execute the method according to any one of claims 1-7.