Data processing method and device, equipment and storage medium

By encoding the relative positions of training samples, the problem of high computational complexity in long text modeling of large models is solved, and low-cost and efficient long text processing capabilities are achieved.

CN120745871APending Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510717052.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing large models have high computational complexity and consume a lot of computing resources when processing long texts, which makes training impossible and limits the context length of long text modeling.

Method used

By determining the target sample length, selecting multiple samples to be trained based on the length, and encoding their relative positions, the target position information is obtained to construct sample data for long text modeling, reducing computational complexity and resource consumption.

Benefits of technology

It achieves low-cost and efficient long text modeling, provides rich sample data, supports the ability of large models to process long texts, and reduces computational complexity and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745871A_ABST
    Figure CN120745871A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, equipment and a storage medium, and relates to the technical field of data processing, in particular to the technical fields of artificial intelligence, deep learning, large models and the like. The specific implementation scheme is as follows: determining a target sample length; wherein the target sample length is used for constraining the sample length of a to-be-trained sample required to be used for training the initial large model; determining N to-be-trained samples based on the target sample length; n is an integer greater than 1; and performing relative position coding on each to-be-trained sample in the N to-be-trained samples based on a target context length and the target sample length expected to be processed by the initial large model to obtain target position information of each to-be-trained sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to technical fields such as artificial intelligence, deep learning, and large models. Background Art

[0002] In recent years, the demand for large models to process long texts has become increasingly urgent. These include applications such as long document comprehension, long article summarization, long-term conversational interaction, logical reasoning based on Chain of Thought (CoT), and long video comprehension. However, the processing power of large models depends on the length of the context they can support. Therefore, a long-text modeling method is urgently needed to support these long-text processing requirements. Summary of the Invention

[0003] The present disclosure provides a data processing method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, a data processing method for large model training is provided, comprising:

[0005] Determine a target sample length; wherein the target sample length is used to constrain the sample length of the training samples required for training the initial large model;

[0006] Based on the target sample length, determining N samples to be trained; N is an integer greater than 1;

[0007] Based on the target context length expected to be processed by the initial large model and the target sample length, relative position encoding is performed on each of the N samples to be trained to obtain target position information of each sample to be trained, wherein the target context length is greater than the target sample length.

[0008] According to another aspect of the present disclosure, there is provided a data processing device for large model training, comprising:

[0009] A determination unit, configured to determine a target sample length; wherein the target sample length is used to constrain the sample length of the to-be-trained samples required for training the initial large model;

[0010] A processing unit is used to determine N samples to be trained based on the target sample length; N is an integer greater than 1; based on the target context length expected to be processed by the initial large model and the target sample length, relatively position encode each sample to be trained in the N samples to be trained to obtain target position information of each sample to be trained, wherein the target context length is greater than the target sample length.

[0011] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0012] at least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0015] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0016] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0017] In this way, the disclosed solution can encode the relative position of each sample to be trained according to the target context length expected by the large model and the target sample length of the sample to be trained for training the initial large model, and thus obtain the target position information of each sample to be trained required for long text modeling. This solution is simple, efficient and practical, and thus provides rich sample data for subsequent large models to realize long text processing, and thus provides strong support for realizing low-cost training of large models.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0020] Figure 1 This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 1 ;

[0021] Figure 2 This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 2 ;

[0022] Figure 3 This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 3 ;

[0023] Figure 4This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 4 ;

[0024] Figure 5 This is a flowchart of a data processing method for large model training in a specific example according to an embodiment of the present application;

[0025] Figure 6 is a structural diagram of a data processing device for large model training according to an embodiment of the present application;

[0026] Figure 7 It is a block diagram of an electronic device used to implement the data processing method for large model training according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0028] The term "and / or" in this article is only a way to describe the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, it can mean including any one or N elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to N similar technical terms and distinguish them. They do not mean to limit the order or to limit the meaning to only two. For example, the first feature and the second feature refer to two types / two features. The first feature can be one or N, and the second feature can also be one or N.

[0029] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0030] The following describes the related technologies of the embodiments of the present disclosure. The following related technologies are optional solutions that can be combined with the technical solutions of the embodiments of the present disclosure in any way, and all of them fall within the protection scope of the embodiments of the present disclosure.

[0031] In recent years, there has been an increasingly urgent demand for large language models (LLMs) to process long texts, such as long document comprehension, long article summaries, long historical dialogue interactions, CoT-based logical reasoning, and long video comprehension. The long-text processing capabilities of large models depend on the length of the context they can support. For example, for understanding long videos (such as a one-hour video), if the video is processed at one frame per second, each frame needs to be represented by 256 visual tokens. The length of the visual token required for a one-hour video is 900k, which is much larger than the length supported by current mainstream large models (such as 32k, 64k, etc.).

[0032] Current approaches to long-text modeling based on large models typically add an additional long-text training step at the end of pre-training, allowing the model to directly accept input data with a specified context length. However, this modeling approach has high computational complexity (its complexity is the square of the context length) and consumes a large amount of computing resources (such as large amounts of video memory). This can even cause model training to fail due to computational resource constraints, severely limiting the context length that large models can support when modeling long-text text.

[0033] Based on this, the disclosed solution provides a data processing method for large model training. According to the context length (corresponding to the target context length) required for modeling the large model, the multiple samples to be trained (here, the sample length of each sample to be trained is less than the target context length) required for training the large model can be relatively encoded to obtain the target position information of each sample to be trained. In this way, the sample data for long text modeling can be obtained efficiently. Here, since the sample length of the sample to be trained is less than the context length required for modeling the large model, compared with the existing solution, the sample data constructed by the disclosed solution is used for model training, which can effectively reduce the computational complexity. Moreover, the consumption of video memory resources is low, which can effectively avoid the situation where the model training cannot run normally due to computing resource problems. In other words, the disclosed solution makes it possible to enable the trained large model to achieve the processing capability of the target context length without using the sample data of the required target context length in the subsequent model training process. In this way, the training complexity and resource consumption are effectively reduced, thereby reducing the training cost.

[0034] Specifically, Figure 1 This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0035] Furthermore, the method includes at least part of the following contents. Figure 1 Shown, including:

[0036] Step S101: Determine the target sample length.

[0037] Here, the target sample length is used to constrain the sample length of the to-be-trained samples required for training the initial large model.

[0038] For example, in one example, the sample length of the training sample required for training the initial large model (also called the actual sample length) is less than or equal to the target sample length. Furthermore, in another example, the sample length of the training sample is equal to the target sample length, thereby providing support for quickly obtaining encoding results.

[0039] Step S102: Based on the target sample length, determine N (N is an integer greater than 1) samples to be trained.

[0040] Step S103: Based on the target context length expected to be processed by the initial large model and the target sample length, relative position encoding is performed on each of the N samples to be trained to obtain target position information of each sample to be trained.

[0041] Here, the target context length is greater than the target sample length.

[0042] Furthermore, relative position encoding of each of the N samples to be trained can be understood as: based on the target position information of the previous sample to be trained among the N samples to be trained, position encoding of the next sample to be trained among the N samples to be trained is performed, so as to achieve the continuation of position information (or position splicing) through N target position information. In this way, the effect of long text is achieved through N consecutive target position information.

[0043] In this way, the disclosed solution can encode the relative position of each sample to be trained according to the target context length expected by the large model and the target sample length of the sample to be trained for training the initial large model, and thus obtain the target position information of each sample to be trained required for long text modeling. This solution is simple, efficient and practical, and thus provides rich sample data for subsequent large models to realize long text processing, and thus provides strong support for realizing low-cost training of large models.

[0044] In addition, the disclosed solution can use the target position information obtained after relative position encoding to compensate for the problem of short sample length of the samples to be trained. That is to say, the disclosed solution can replace long texts based on the obtained target position information and use samples to be trained with shorter text lengths. In this way, it is convenient to directly use shorter sample data to realize long text modeling in the future, thereby laying the foundation for quickly obtaining a large model that supports ultra-long context length processing.

[0045] In a specific example, the samples to be trained mentioned above may specifically include but are not limited to at least one of the following: text, image, video, etc., which is not limited in the present disclosure.

[0046] Furthermore, in one example, the sample to be trained may be specifically obtained by splicing a plurality of shorter initial samples, and the length of the spliced ​​sample is less than or equal to the target sample length.

[0047] In a specific example, the relative position encoding of the training sample in the above example can be specifically performed as position encoding of each token in the sample feature of the training sample (for example, obtained after feature encoding of the training sample). At this time, the target position information of the training sample can be specifically represented by a position vector, which can represent the position information of each token after encoding.

[0048] Furthermore, in a specific example, at least one batch of sample data sets can be constructed based on the N samples to be trained.

[0049] Furthermore, the difference between the interval length of the position encoding interval corresponding to the sample data set and the target context length is less than or equal to a preset threshold. For example, the absolute value of the difference is less than or equal to the preset threshold. Here, the position encoding interval corresponding to the sample data set is determined based on the target position information of all samples to be trained in the sample data set.

[0050] In this way, it can be effectively ensured that the interval length of the position coding interval determined based on all the target position information is close to or equal to the target context length, providing rich sample data for subsequent large models to realize long text processing. At the same time, it also provides strong support for effectively reducing the computational complexity and required computing resources in the model training process.

[0051] For example, in one example, a sample batch is constructed based on N samples to be trained. At this time, the interval length of the position coding interval determined by all the target position information of the sample batch obtained using the disclosed solution approaches or is equal to the target context length.

[0052] Alternatively, in another example, all batches of samples required for training the large model are constructed based on the N samples to be trained; at this time, the interval length of the position coding interval determined by all target position information of all sample batches obtained using the disclosed solution approaches or is equal to the target context length.

[0053] For example, if the target context length expected to be processed by the initial large model is 8k, then a "position splicing" method can be used to "position splice" multiple shorter training samples so that the interval length of the overall position information after "position splicing" is 8k, without the sample length of the training samples being directly 8k. "Position splicing" here means: by relative position encoding each shorter training sample, the encoded multiple target position information can cover a partial range from 0 to 8k. For example, the target position information interval of training sample 1 (whose sample length is 2k+1) after position encoding is [0, 2k], the target position information interval of training sample 2 (whose sample length is 3k) after position encoding is [2k+1, 5k], and the target position information interval of training sample 3 (whose sample length is 2k+1) after position encoding is [6k, 8k]. At this time, the interval length of the position encoding interval determined by training samples 1 to 3 is close to the target context length of 8k.

[0054] It should be noted that the above position encoding method is only an exemplary description. In actual applications, there may be other encoding results, and the present disclosure does not impose any specific restrictions on this.

[0055] In addition, it should be pointed out that the main factor affecting the long text processing of large models (such as transformer models) is the relative position information when calculating the attention weight. For example, if the large model is expected to support text with a context length of 8k, it is necessary to input a combination of position information within 8k in the attention stage. Therefore, the use of the disclosed solution can effectively realize long text modeling.

[0056] Furthermore, it should be noted that the computational complexity brought about by model (such as transformer model) training is mainly due to attention calculation, for example, the computational complexity is the square of the length of the text used; therefore, by adopting the disclosed solution, long text modeling can be achieved without increasing the sample length of the training sample, effectively reducing the text length of the required text, thereby effectively reducing the computational complexity, saving computing resources, and also improving the training effect.

[0057] Figure 2 This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 2 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0058] Furthermore, the method includes at least part of the following contents. Figure 2 Shown, including:

[0059] Step S201: Determine the target sample length.

[0060] Here, the target sample length is used to constrain the sample length of the to-be-trained samples required for training the initial large model.

[0061] Step S202: Based on the target sample length, determine N (N is an integer greater than 1) samples to be trained.

[0062] Here, for relevant examples of the samples to be trained, please refer to the above description and will not be repeated here.

[0063] Step S203: Based on the target context length expected to be processed by the initial large model and the target sample length, relative position encoding is performed on each of the N samples to be trained to obtain target position information of each sample to be trained.

[0064] Step S204: Determine the data features of each to-be-trained sample. For example, obtain a feature vector of each to-be-trained sample, and the feature vector may represent the data features of the to-be-trained sample.

[0065] Here, the execution order of step S203 and step S204 can be swapped, or they can be executed simultaneously, etc. The present disclosure does not impose any specific restrictions on the execution order of the two.

[0066] Step S205: performing attention processing based on the data features of each to-be-trained sample and the target position information of each to-be-trained sample to obtain an attention processing result.

[0067] Step S206: Based on the attention processing result, the initial large model is trained to obtain a trained target large model.

[0068] Here, the trained target large model can be used to process data with target context length.

[0069] For example, in one example, the initial large model may specifically include a preprocessing module, an attention module and an output module; further, the preprocessing module is used to obtain the target position information of each sample to be trained, and the preprocessing module is used to obtain the data features of each sample to be trained; then, the attention module is used to perform attention processing on the data features of each sample to be trained and the target position information of each sample to be trained to obtain an attention processing result; finally, the output module is used to process the attention processing result to obtain a model inference result; and then, based on the degree of difference between the model inference result and the actual inference result, some adjustable parameters in the initial large model are fine-tuned to obtain the trained target large model.

[0070] It should be noted here that the above is only an exemplary explanation. In actual applications, the target position information of each sample to be trained can be obtained first, and then each sample to be trained and the target position information of each sample to be trained can be input into the initial large model. At this time, after using the preprocessing module to obtain the data features of each sample to be trained, the attention module is used to perform attention processing on the data features of each sample to be trained and the target position information of each sample to be trained, and then the model inference result is obtained.

[0071] In this way, since the sample length of the samples to be trained in the disclosed scheme is smaller than the context length required for modeling the large model, compared with the scheme of directly using sample data of the required context length for modeling (that is, the target context length) for model training, using the sample data constructed by the disclosed scheme for model training can effectively reduce the computational complexity. Moreover, the consumption of video memory resources is low, which can effectively avoid the situation where the model training cannot run normally due to computing resource problems, thereby achieving low-cost training of large models.

[0072] Furthermore, the disclosed solution can support context modeling of any length, provided there is a sufficiently large sample data volume. Furthermore, the disclosed solution is highly compatible and can be used in conjunction with other modeling approaches, thus ensuring the efficiency and stability of the trained target large model when processing long text data.

[0073] Furthermore, in one example, after obtaining the target large model, the target large model can also be used to process data to be inferred whose text length does not exceed the target context length; specifically, the method further includes:

[0074] Obtain inference data (e.g., text or images) whose length does not exceed the target context length; further, input this inference data into the target large model to obtain the target inference result. This effectively achieves accurate inference on long texts, thereby providing fast and accurate decision support for multiple fields, laying the foundation for significantly improving information processing capabilities and decision-making efficiency, thereby enhancing user experience.

[0075] Figure 3 This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 3 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 and Figure 2 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0076] Furthermore, the method includes at least part of the following contents. Figure 3 Shown, including:

[0077] Step S301: Determine the initial context length that the initial large model can currently process.

[0078] Step S302: Determine a target sample length based on the initial context length.

[0079] Here, the target sample length is used to constrain the sample length of the to-be-trained samples required for training the initial large model.

[0080] Furthermore, in a specific example, the above-mentioned determination of the target sample length based on the initial context length (e.g., step S302) may specifically include:

[0081] Step S302-1: Determine the segmentation parameters of the initial context length.

[0082] Step S302-2: based on the segmentation parameter, the initial context length is segmented to obtain a target sample length.

[0083] For example, in one example, the above-mentioned segmentation of the initial context length based on the segmentation parameters may specifically include: based on the segmentation parameters, evenly segmenting the initial context length to obtain an average division result; and obtaining the target sample length based on the average division result.

[0084] For example, when the initial context length is SEQ_LEN, if the segmentation parameter is k, the target sample length is SEQ_LEN / k, where k is an integer greater than 1.

[0085] Thus, the disclosed solution provides a simple, efficient, and interpretable method for calculating the target sample length, thereby supporting the subsequent implementation of relative position encoding for training samples. Furthermore, the segmentation parameters can be set based on actual training requirements (e.g., loss tolerance), thus facilitating flexible adaptation to various training scenarios.

[0086] Step S303: Based on the target sample length, determine N (N is an integer greater than 1) samples to be trained.

[0087] Here, for relevant examples of the samples to be trained, please refer to the above description and will not be repeated here.

[0088] It should be noted that, in one example, the number N of samples to be trained is related to the segmentation parameter, for example, N is equal to the segmentation parameter, so as to reduce the computational complexity and the consumption of video memory resources.

[0089] Step S304: Based on the target context length expected to be processed by the initial large model and the target sample length, relative position encoding is performed on each of the N samples to be trained to obtain target position information of each sample to be trained.

[0090] In this way, the disclosed scheme provides a refined scheme for obtaining the target sample length. Here, since the target sample length is obtained based on the initial context length that the initial large model can currently process, the samples to be trained determined based on the target sample length are within the capability range of the initial large model. This provides strong support for effectively reducing computational complexity and reducing the consumption of video memory resources in the subsequent training process (that is, the training process of training the large model using the samples to be trained), thereby laying the foundation for obtaining a large model with ultra-long context processing capabilities at low cost.

[0091] Figure 4 This is a schematic flow chart of a data processing method for large model training according to an embodiment of the present application. Figure 4 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figures 1 to 3 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0092] Furthermore, the method includes at least part of the following contents. Figure 4 Shown, including:

[0093] Step S401: Determine the target sample length.

[0094] Here, the target sample length is used to constrain the sample length of the to-be-trained samples required for training the initial large model.

[0095] Step S402: Based on the target sample length, determine N (N is an integer greater than 1) samples to be trained.

[0096] It should be noted that, for the relevant contents about the target sample length and the samples to be trained, please refer to the above examples and will not be repeated here.

[0097] Step S403: obtaining target position information of the i-th sample to be trained among the N samples to be trained based at least on the target context length and the target sample length.

[0098] Here, i is an integer greater than or equal to 1 and less than N. In other words, the i-th to-be-trained sample is any one of the N to-be-trained samples except the last one.

[0099] For example, in one example, the target position information of the i-th sample to be trained may also be related to the segmentation parameters described above. In this case, obtaining the target position information of the i-th sample to be trained among the N samples to be trained based at least on the target context length and the target sample length may specifically include obtaining the target position information of the i-th sample to be trained among the N samples to be trained based on the target context length, the target sample length, and the segmentation parameters. In this way, the target position information of the sample to be trained can be accurately encoded.

[0100] Here, it can be understood that the target position information of the i-th sample to be trained is obtained based at least on the target sample length.

[0101] For example, in one example, for i taking a value of 1, the target position information of the first sample to be trained (that is, the target position information of the first sample to be trained) can be directly obtained based on the actual sample length of the first sample to be trained (for example, the actual sample length is equal to the target sample length). For example, the target position information of the first sample to be trained can be specifically [0, block-1], where block represents the actual sample length of the first sample to be trained. If the actual sample length of the first sample to be trained is equal to the target sample length, it can also represent the target sample length.

[0102] It is understandable that the interval length of the target position information of the training sample is related to the actual sample length of the training sample, for example, the two lengths are the same, so as to achieve position encoding of each feature unit (for example, token) of the training sample.

[0103] Furthermore, for i with a value of 2, the target position information of the i-th training sample can be obtained based on the target context length and the target position information of the i-1-th training sample. The specific method for obtaining the target position information can be found in the steps related to the target position information of the i+1-th training sample, and will not be repeated here.

[0104] Step S404: Based on the target context length and the target position information of the i-th sample to be trained, the starting position of the i+1-th sample to be trained is obtained.

[0105] For example, when the value of i is 1, if the target position information of the first sample to be trained is specifically [0, block1-1], the starting position of the second sample to be trained may be block1.

[0106] Step S405: based on the starting position of the i+1th sample to be trained and the actual sample length of the i+1th sample to be trained, obtain the target position information of the i+1th sample to be trained, so as to obtain the target position information of each sample to be trained.

[0107] Here, it should be noted that if the actual sample length of the i+1th sample to be trained is equal to the target sample length, step S405 can be specifically as follows: based on the starting position of the i+1th sample to be trained and the target sample length, the target position information of the i+1th sample to be trained is obtained to obtain the target position information of each sample to be trained.

[0108] That is, in one example, among N samples to be trained, the target position information of the next sample to be trained depends on the target position information of the previous sample to be trained. That is, after obtaining the target position information of the i-th sample to be trained, the starting position of the i+1-th sample to be trained is determined based on the target context length and the target position information of the i-th sample to be trained. Then, based on the starting position of the i+1-th sample to be trained and the target sample length, the target position information of the i+1-th sample to be trained is encoded. This cycle repeats until the relative position encoding of all samples to be trained is completed. In this way, the target position information of each sample to be trained can be obtained.

[0109] Here, it should be noted that the training samples in the N training samples may be out of order, but in the position coding process, the position coding of the latter depends on the position coding result of the former.

[0110] In this way, the disclosed scheme provides a specific scheme for determining the target position information of each sample to be trained, that is, first use the target position information and target context length of the previous sample to be trained to determine the starting position of the current sample to be trained, and then obtain the target position information of the current sample to be trained based on the starting position and the actual sample length of the current sample to be trained. In this way, the target position information of each sample to be trained can be efficiently encoded. The scheme is simple, practical, and interpretable. Moreover, when the sample to be trained does not need to reach the target context length, the target context length can be modeled through this encoding method, thereby effectively reducing the computational complexity and reducing the consumption of video memory resources in the subsequent training process (that is, the training process of training a large model using the sample to be trained). It provides strong support, thereby laying the foundation for obtaining a large model with ultra-long context processing capabilities at low cost.

[0111] Furthermore, in a specific example, the above-mentioned obtaining the target position information of the i+1th to-be-trained sample based on the starting position of the i+1th to-be-trained sample and the actual sample length of the i+1th to-be-trained sample may specifically include:

[0112] Starting from the starting position of the i+1th sample to be trained, encoding is performed based on the actual sample length of the i+1th sample to be trained to obtain the target position information of the i+1th sample to be trained.

[0113] For example, if the actual sample length of the i+1th sample to be trained is equal to the target sample length, the encoding can be continued from the starting position of the i+1th sample to be trained based on the target sample length to obtain the target position information of the i+1th sample to be trained.

[0114] For example, taking the second sample to be trained as an example, after determining that the starting position of the second sample to be trained is start_id2, at this time, it is possible to start from the starting position of the second sample to be trained and perform encoding extension according to the actual sample length of the second sample to be trained (which can be recorded as block2). For example, if the actual sample length of the second sample to be trained is also equal to the target sample length, encoding extension can also be performed based on the target sample length to obtain the target position information of the second sample to be trained. At this time, the target position information of the second sample to be trained can be expressed as:

[0115] [start_id2,start_id2+block2―1].

[0116] In this way, the disclosed solution can quickly and efficiently complete the position encoding of each sample to be trained, thereby providing strong support for effectively reducing the computational complexity and reducing the consumption of video memory resources in the subsequent training process (that is, the training process of training a large model using the sample to be trained), thereby laying the foundation for obtaining a large model with ultra-long context processing capabilities at low cost.

[0117] Furthermore, in a specific example, the starting position of the i+1th sample to be trained can be obtained in the following manner; specifically, the above-described method of obtaining the starting position of the i+1th sample to be trained based on the target context length and the target position information of the i-th sample to be trained (e.g., step S404) can specifically include:

[0118] Step S404 - 1 : Determine the i-th position sampling interval based on the target context length and the end position in the target position information of the i-th sample to be trained.

[0119] It should be noted that, in a specific example, in the process of determining the i-th position sampling interval, in addition to referring to the target context length and the end position in the target position information of the i-th sample to be trained, the target sample length and segmentation parameters can also be referred to.

[0120] For example, in one example, for i being an integer greater than 1, the sampling interval at the i-th position is obtained based on the following expression:

[0121] [Tp i +1,Ms―(m―i)×Tl];

[0122] Here, Tp i represents the end position of the i-th sample to be trained, Ms represents the target context length, m represents the segmentation parameter (ie, the segmentation parameter required to obtain the target sample length), and Tl represents the target sample length.

[0123] In this way, the disclosed solution provides a specific solution for obtaining the sampling interval of the i-th position. The solution is simple, efficient, and interpretable. In this way, it provides support for quickly determining the starting position of the i+1-th sample to be trained. At the same time, it provides strong support for effectively reducing the computational complexity and reducing the consumption of video memory resources in the subsequent training process (that is, the training process of training a large model using the sample to be trained), thereby laying the foundation for obtaining a large model with ultra-long context processing capabilities at low cost.

[0124] Step S404-2: In the i-th position sampling interval, sample to obtain the starting position of the i+1-th sample to be trained.

[0125] That is, for the i+1th sample to be trained, after obtaining the i-th position sampling interval, the starting position of the i+1th sample to be trained can be obtained by sampling from multiple candidate starting positions included in the i-th position sampling interval.

[0126] In this way, the disclosed solution provides a refined solution for determining the starting position of the sample to be trained, that is, the starting position of the current sample to be trained (such as the i+1th sample to be trained) can be quickly determined by using the position sampling interval. In this way, it provides strong support for quickly encoding samples that can be used for long text modeling. At the same time, it provides strong support for subsequently reducing the computational complexity brought about by the training process and the required consumption of video memory resources, thereby laying the foundation for obtaining a large model with ultra-long context processing capabilities at low cost.

[0127] Furthermore, in one example, the above-mentioned sampling in the i-th position sampling interval to obtain the starting position of the i+1-th sample to be trained may specifically include: randomly sampling in the i-th position sampling interval to obtain the starting position of the i+1-th sample to be trained. In this way, the position encoding is generalized to the target context length through random sampling, which provides strong support for effectively reducing the computational complexity and reducing the consumption of video memory resources in the subsequent training process (that is, the training process of using the sample to be trained to train a large model), thereby laying the foundation for obtaining a large model with ultra-long context processing capabilities at low cost.

[0128] It should be pointed out that since the starting position of the training sample in the disclosed solution is determined randomly, the target position information of each training sample obtained based on the disclosed solution can specifically correspond to different continuous segments within [0, target context length - 1]. At this time, when the sample size is large enough, all possible continuous segments within the target context length can be effectively obtained. In other words, all target position information will almost cover the entire [0, target context length - 1]. In this way, compared with the solution of directly using sample data of the required modeling context length (that is, the target context length) for model training, using the sample data constructed by the disclosed solution for model training can effectively reduce the computational complexity. Moreover, the consumption of video memory resources is low, which can effectively avoid the situation where the model training cannot run normally due to computing resource problems, thereby achieving low-cost training of large models.

[0129] For example, assuming that the target context length to be achieved is 64k, if the initial context length of the initial large model is 16k and the segmentation parameter is 4, then the target sample length of the samples to be trained required to constrain the training of the initial large model can be specifically 16k / 4=4k. In other words, the actual sample length of the samples to be trained is less than or equal to 4k; further, since the initial context length of the initial large model is 16k, the total sample length input to the initial large model can be specifically 16k, and since the actual sample length of the samples to be trained is less than or equal to 4k, the number of a group of samples to be trained required for position encoding is 4. At this time, the group of samples to be trained (that is, 4 samples to be trained) can be spliced ​​and used as the target samples actually used for model training.

[0130] That is to say, in one example, the number of samples N to be trained is related to the segmentation parameter, for example, the two are the same. In other words, the total sample length obtained after splicing the N samples to be trained is close to or equal to the initial context length of the initial large model, so as to ensure that the initial context length of the initial large model can be reached or approached after splicing the N samples to be trained.

[0131] Furthermore, after position encoding of the four samples to be trained, the target position information of the four samples to be trained can be approximately covered from 0 to 64k. For example, the specific encoding process is as follows:

[0132] The target position information of the first sample to be trained (its actual sample length is 4k) is encoded as [0,4k-1].

[0133] Furthermore, it is determined that the first position sampling interval is [4k, 64k-(4-1)×4k]; at this time, the starting position of the second sample to be trained can be randomly sampled from the first position sampling interval, for example, 16k, and based on the starting position of the second sample to be trained, the target position information of the second sample to be trained (whose actual sample length is also 4k) is [16k, 20k-1].

[0134] Furthermore, the second position sampling interval is determined to be [20k, 64k-(4-2)×4k]; at this time, the starting position of the third sample to be trained can be randomly sampled from the second position sampling interval, for example, 24k, and based on the starting position of the third sample to be trained, the target position information of the third sample to be trained (whose actual sample length is also 4k) is [24k, 28k-1].

[0135] Furthermore, the sampling interval of the third position is determined to be [28k, 64k-(4-3)×4k]; at this time, the starting position 36k of the fourth sample to be trained can be obtained by random sampling from the sampling interval of the third position, and according to the starting position of the fourth sample to be trained, the target position information of the fourth sample to be trained (whose actual sample length is also 4k) is obtained as [36k, 40k-1].

[0136] Based on the above encoding method, all the target position information obtained almost covers the entire [0,40k-1], exceeding the initial context length of 16k.

[0137] It should be noted that when the number of samples to be trained is large enough, the method described in the disclosed solution can cover [0, 64k].

[0138] In this way, the disclosed solution can utilize shorter training samples and achieve long text effects through relative position encoding. This provides strong support for subsequently reducing the computational complexity and memory resources required in the training process, thereby laying the foundation for obtaining a large model with ultra-long context processing capabilities at low cost.

[0139] The present disclosure scheme is further explained in detail with reference to the following specific examples; specifically, assuming that the target context length that the initial large model is expected to support is denoted as MAX_SEQ, the training context length that the existing baseline of the initial large model can support (corresponding to the initial context length described above) is denoted as SEQ_LEN, the segmentation parameter is set to 2, and the target sample length of the samples to be trained required to constrain the training of the initial large model is denoted as BLOCK, then BLOCK = SEQ_LEN / 2, and the number N of samples to be trained used for relative position encoding in the present disclosure scheme is 2.

[0140] Furthermore, if Figure 5 As shown, the core steps of the position encoding of the disclosed solution include:

[0141] Step S501: Determine two training samples that need to be encoded currently, and the actual sample length of each training sample is BLOCK.

[0142] Step S502: Encode the target position information of the first sample to be trained as [0, BLOCK-1].

[0143] Step S503: determine the position sampling interval [BLOCK, MAX_SEQ―BLOCK];

[0144] Step S504: Randomly sample from the obtained position sampling interval [BLOCK, MAX_SEQ―BLOCK] to determine the starting position of the second sample to be trained (for example, it can be recorded as start_id). Here, the starting position start_id of the second sample to be trained can be specifically:

[0145] start_id=random[BLOCK,MAX_SEQ―BLOCK].

[0146] Here, random[·,·] represents a random function.

[0147] Step S505: According to the starting position start_id of the second sample to be trained and the actual sample length (also called target sample length) BLOCK of the second sample to be trained, the target position information of the second sample to be trained is obtained, that is, [start_id, start_id+BLOCK-1].

[0148] Step S506: concatenate the two position-encoded samples to be trained to obtain actual training samples corresponding to the current sample batch.

[0149] Here, in one example, when performing sample splicing, the target position information of each sample to be trained may also be spliced ​​into one piece of position information.

[0150] At this point, repeat the above operations to obtain the actual training samples corresponding to each sample batch.

[0151] It should be noted that when the sample size is large enough, the coverage of all target position information obtained is close to [0, MAX SEQ ―1]. At this point, the initial large model is trained using the actual training samples obtained and their spliced ​​position information to obtain the target large model that can handle the target context length.

[0152] In summary, compared with the prior art, the disclosed solution has the following advantages:

[0153] First, rich sample data. The sample data processing solution provided by the disclosed solution is simple, efficient, and practical, providing rich sample data for large models to process long texts.

[0154] Second, it requires fewer computing resources. Compared to existing technologies, the disclosed solution can train large models based on the obtained target location information and utilize shorter training samples to replace long texts. This effectively reduces computational complexity and consumes less video memory resources, effectively avoiding situations where model training cannot run normally due to computing resource issues, thereby achieving low-cost training of large models.

[0155] Third, it supports modeling long texts of arbitrary length. Because the disclosed solution can model long texts of arbitrary length given sufficient sample data, and has strong compatibility and can be used in conjunction with other modeling methods, the trained target large model is guaranteed to be efficient and stable when processing long text data.

[0156] The disclosed solution also provides a data processing device for large model training, such as Figure 6 Shown, including:

[0157] The determination unit 601 is used to determine a target sample length; wherein the target sample length is used to constrain the sample length of the training sample required for training the initial large model;

[0158] Processing unit 602 is used to determine N samples to be trained based on the target sample length; N is an integer greater than 1; based on the target context length expected to be processed by the initial large model and the target sample length, relatively position encode each sample to be trained in the N samples to be trained to obtain target position information of each sample to be trained, wherein the target context length is greater than the target sample length.

[0159] In a specific example of the present disclosure, at least one batch of sample data sets can be constructed based on the N samples to be trained;

[0160] In which, the difference between the interval length of the position coding interval corresponding to the sample data set and the target context length is less than or equal to a preset threshold, and the position coding interval corresponding to the sample data set is determined based on the target position information of all samples to be trained in the sample data set.

[0161] In a specific example of the disclosed solution, the system further includes a model training unit; wherein the model training unit is configured to:

[0162] Determine the data characteristics of each sample to be trained;

[0163] Performing attention processing based on the data features of each to-be-trained sample and the target position information of each to-be-trained sample to obtain an attention processing result;

[0164] Based on the attention processing results, the initial large model is trained to obtain a trained target large model; wherein the trained target large model can be used to process data with a target context length.

[0165] In a specific example of the disclosed solution, the determining unit is specifically configured to:

[0166] Determine the initial context length that the initial large model can currently process;

[0167] Based on the initial context length, a target sample length is determined.

[0168] In a specific example of the disclosed solution, the determining unit is specifically configured to:

[0169] Determining a segmentation parameter of the initial context length;

[0170] The initial context length is segmented based on the segmentation parameter to obtain a target sample length.

[0171] In a specific example of the disclosed solution, the determining unit is specifically configured to:

[0172] Based on the segmentation parameter, the initial context length is evenly divided to obtain an evenly divided result;

[0173] Based on the average division result, the target sample length is obtained.

[0174] In a specific example of the disclosed solution, the processing unit is specifically configured to:

[0175] Obtaining target position information of an i-th to-be-trained sample among the N to-be-trained samples based at least on the target context length and the target sample length; wherein i is an integer greater than or equal to 1 and less than N;

[0176] Based on the target context length and the target position information of the i-th sample to be trained, the starting position of the i+1-th sample to be trained is obtained;

[0177] Based on the starting position of the i+1th sample to be trained and the actual sample length of the i+1th sample to be trained, the target position information of the i+1th sample to be trained is obtained.

[0178] In a specific example of the disclosed solution, the processing unit is specifically configured to:

[0179] Determine an i-th position sampling interval based on the target context length and the end position in the target position information of the i-th sample to be trained;

[0180] In the i-th position sampling interval, the starting position of the i+1-th sample to be trained is obtained by sampling.

[0181] In a specific example of the disclosed solution, the sampling interval at the i-th position is obtained based on the following expression:

[0182] [Tp i +1,Ms―(m―i)×Tl];

[0183] Among them, Tp i Represents the end position of the i-th sample to be trained, Ms represents the target context length, m represents the segmentation parameter, and Tl represents the target sample length.

[0184] In a specific example of the disclosed solution, the processing unit is specifically configured to:

[0185] In the i-th position sampling interval, the starting position of the i+1-th sample to be trained is obtained by random sampling.

[0186] In a specific example of the disclosed solution, the processing unit is specifically configured to:

[0187] Starting from the starting position of the i+1th sample to be trained, encoding is performed based on the actual sample length of the i+1th sample to be trained to obtain the target position information of the i+1th sample to be trained.

[0188] In a specific example of the disclosed solution, the system further includes an inference unit, wherein the inference unit is configured to:

[0189] Obtain the data to be inferred whose text length does not exceed the target context length;

[0190] The data to be inferred is input into the target large model to obtain the target inference result.

[0191] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0192] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0193] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0194] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0195] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0196] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0197] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the data processing method for large model training. For example, in some embodiments, the data processing method for large model training can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the data processing method for large model training described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the data processing method for large model training in any other appropriate manner (e.g., by means of firmware).

[0198] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0199] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0200] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0201] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0202] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0203] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0204] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0205] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A data processing method for large model training, comprising: Determine a target sample length; wherein the target sample length is used to constrain the sample length of the training samples required for training the initial large model; Based on the target sample length, determining N samples to be trained; N is an integer greater than 1; Based on the target context length expected to be processed by the initial large model and the target sample length, relative position encoding is performed on each of the N samples to be trained to obtain target position information of each sample to be trained, wherein the target context length is greater than the target sample length.

2. The method according to claim 1, wherein At least one batch of sample data sets can be constructed based on the N samples to be trained; In which, the difference between the interval length of the position coding interval corresponding to the sample data set and the target context length is less than or equal to a preset threshold, and the position coding interval corresponding to the sample data set is determined based on the target position information of all samples to be trained in the sample data set.

3. The method according to claim 1 or 2, further comprising: Determine the data characteristics of each sample to be trained; Performing attention processing based on the data features of each to-be-trained sample and the target position information of each to-be-trained sample to obtain an attention processing result; Based on the attention processing results, the initial large model is trained to obtain a trained target large model; wherein the trained target large model can be used to process data with a target context length.

4. The method according to any one of claims 1 to 3, wherein: Determining the target sample length includes: Determine the initial context length that the initial large model can currently process; Based on the initial context length, a target sample length is determined.

5. The method according to claim 4, wherein The determining of the target sample length based on the initial context length includes: Determining a segmentation parameter of the initial context length; The initial context length is segmented based on the segmentation parameter to obtain a target sample length.

6. The method according to claim 5, wherein: The step of segmenting the initial context length based on the segmentation parameter includes: Based on the segmentation parameter, the initial context length is evenly divided to obtain an evenly divided result; Based on the average division result, the target sample length is obtained.

7. The method according to any one of claims 1 to 6, wherein: The step of performing relative position encoding on each of the N samples to be trained based on the target context length and the target sample length expected to be processed by the initial large model to obtain target position information of each sample to be trained includes: Obtaining target position information of an i-th to-be-trained sample among the N to-be-trained samples based at least on the target context length and the target sample length; wherein i is an integer greater than or equal to 1 and less than N; Based on the target context length and the target position information of the i-th sample to be trained, the starting position of the i+1-th sample to be trained is obtained; Based on the starting position of the i+1th sample to be trained and the actual sample length of the i+1th sample to be trained, the target position information of the i+1th sample to be trained is obtained.

8. The method according to claim 7, wherein: The step of obtaining the starting position of the (i+1)th sample to be trained based on the target context length and the target position information of the (i)th sample to be trained comprises: Determine an i-th position sampling interval based on the target context length and the end position in the target position information of the i-th sample to be trained; In the i-th position sampling interval, the starting position of the i+1-th sample to be trained is obtained by sampling.

9. The method according to claim 8, wherein The sampling interval of position i is obtained based on the following expression: [Tp i +1,Ms―(m―i)×Tl]; Among them, Tp i Represents the end position of the i-th sample to be trained, Ms represents the target context length, m represents the segmentation parameter, and Tl represents the target sample length.

10. The method according to claim 8, wherein The step of sampling the starting position of the i+1th sample to be trained in the i-th position sampling interval includes: In the i-th position sampling interval, the starting position of the i+1-th sample to be trained is obtained by random sampling.

11. The method according to any one of claims 7 to 9, wherein: The step of obtaining target position information of the i+1th sample to be trained based on the starting position of the i+1th sample to be trained and the actual sample length of the i+1th sample to be trained includes: Starting from the starting position of the i+1th sample to be trained, encoding is performed based on the actual sample length of the i+1th sample to be trained to obtain the target position information of the i+1th sample to be trained.

12. The method according to claim 3, further comprising: Obtain the data to be inferred whose text length does not exceed the target context length; The data to be inferred is input into the target large model to obtain the target inference result.

13. A data processing device for large model training, comprising: A determination unit, configured to determine a target sample length; wherein the target sample length is used to constrain the sample length of the to-be-trained samples required for training the initial large model; A processing unit is used to determine N samples to be trained based on the target sample length; N is an integer greater than 1; based on the target context length expected to be processed by the initial large model and the target sample length, relatively position encode each sample to be trained in the N samples to be trained to obtain target position information of each sample to be trained, wherein the target context length is greater than the target sample length.

14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.

16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Long text processing method, related equipment and readable storage medium

    CN112527992A

  • Signal peptide prediction method, prediction model construction method, prediction model construction device and computing equipment

    CN117253545A

  • Text task processing method and model training method thereof, equipment, medium and product

    CN118586448A

  • Training method and system for enhancing long text processing capability of large language model

    CN118643883A

  • Long text generation alignment method and system based on synthetic position coding

    CN119493869A