Training method, data prediction method and device, electronic equipment and storage medium
By inserting redundant noise data in the training of data inference model and combining mixed attention mechanism and stream memory management, the problem of insufficient long-distance dependence processing capability of data inference model is solved, improving the robustness of the model and the reliability of data prediction.
Patent Information
- Application Number
- CN202510748990.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing data inference model is not ideal in long data processing, especially inadequate processing capabilities for long-distance dependence.
By inserting redundant noise data between the original training data segments, the data inference model is trained to enhance its processing power on long-distance dependence, optimizing model performance in combination with local-global hybrid attention mechanisms and stream memory management.
提高了数据推理模型的鲁棒性和长距离依赖的处理性能,优化了数据预测的可靠性,能够处理超长数据任务并减少内存占用。
Smart Images

Figure CN120278282A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data reasoning, and in particular, to a training method for a data reasoning model, a data prediction method, a data prediction device, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of language models such as LLM (Large Language Model), it is urgent to improve the long data processing performance thereof. Since long data such as non-synthetic long text data is scarce, long data with context dependence can usually be artificially synthesized to train a data reasoning model, but there are still problems with unsatisfactory data reasoning effects after the data reasoning model is actually put into use. Summary of the Invention
[0003] The present application provides a training method for a data reasoning model, a data prediction method, a data prediction device, an electronic device, and a computer-readable storage medium to at least solve the problem of unsatisfactory data reasoning effects in related technologies.
[0004] The present application provides a training method for a data reasoning model. The training method for the data reasoning model includes: obtaining original training data and dividing it into original data segments; obtaining redundant noise data and inserting it between the original data segments to form an enhanced data sequence of the original training data; inputting the enhanced data sequence into the data reasoning model to obtain an inference output of the data reasoning model; and using the inference output to update the loss of the data reasoning model.
[0005] The present application further provides a data prediction method. The data prediction method includes: obtaining data to be predicted; inputting the data to be predicted into the data reasoning model and using the data reasoning model to predict the subsequent data of the data to be predicted, and outputting the subsequent data; wherein, the data reasoning model is trained by using the training method of the data reasoning model as described above.
[0006] The present application further provides a data prediction device. The data prediction device includes: an input module and a control module; the control module is connected to the input module; and is used to implement the steps of the training method of the data reasoning model as described above, or implement the steps of the data prediction method as described above.
[0007] The present application further provides an electronic device. The electronic device includes: a memory for storing a computer program; a processor for implementing the steps of the training method of the data reasoning model as described above when executing the computer program; or implementing the steps of the data prediction method as described above.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned training method of the data inference model are implemented; or, the steps of the above-mentioned data prediction method are implemented.
[0009] Through the present application, redundant noise data is inserted between the original data segments of the original training data to form an enhanced data sequence that extends the length of the data sample. Between the original data segments with short-distance dependencies in the original training data, since redundant noise data is inserted, the distance between them is extended, so that short-distance dependencies can be converted into long-distance dependencies. In this way, when training the data inference model using the enhanced data sequence, the processing ability of the data inference model regarding long-distance dependencies can be trained, thereby improving the robustness of the data inference model, optimizing the data inference effect of the data inference model, and further improving the reliability of prediction content such as data prediction. Therefore, the technical problem of unsatisfactory data inference effect can be solved, and the technical effect of enhancing the processing performance of the data inference model for long-distance dependencies to improve the robustness of the data inference model can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic diagram of an application scenario of an embodiment of the data inference model of the present application; Figure 2 It is a schematic flowchart of an embodiment of the data prediction method of the present application; Figure 3 It is a schematic flowchart of an embodiment of the training method of the data inference model of the present application; Figure 4 It is a schematic flowchart of another embodiment of the training method of the data inference model of the present application; Figure 5 It is a schematic flowchart of an embodiment of forming an enhanced data sequence of the present application; Figure 6 It is a schematic flowchart of an embodiment of the data prediction device of the present application; Figure 7 It is a schematic flowchart of an embodiment of the electronic device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0013] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0014] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in combination with the accompanying drawings and specific embodiments.
[0015] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the training method and / or data prediction method of the data inference model depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0016] Please refer to Figure 1 , Figure 1 , which is a schematic diagram of the application scenario of an embodiment of the data inference model of the present application.
[0017] In one embodiment, the application scenario of the data inference model may include a terminal 10 and a data inference model. Among them, the data inference model may be deployed on the terminal 10; or as Figure 1 exemplified, the data inference model may be deployed on other devices 20 such as a server independent of the terminal 10.
[0018] For example, a user can input data content through the terminal 10. The terminal 10 can transmit the data content to the data inference model. The data inference model performs data inference that conforms to its function in response to the acquired data content, outputs its data output based on the inference, and feeds it back to the terminal 10. The terminal 10 can display the inference data output. For example, the function of the data inference model may be to continue data prediction, output answers to question inputs, etc., which are not strictly limited herein.
[0019] Among them, the terminal 10 can be a mobile phone, a tablet computer, a computer, a smart watch, an AR (Augmented Reality) device, a VR (Virtual Reality) device, etc.
[0020] Taking the data inference model having the function of predicting subsequent data as an example, the embodiments of the present application provide a data prediction method. Combining the execution process of the data prediction method, the data prediction method will be described in detail.
[0021] Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of an embodiment of the data prediction method of the present application.
[0022] S101: Obtain the data to be predicted.
[0023] In this embodiment, the data to be predicted represents the data content for which the subsequent data characters are to be predicted, that is, the data to be predicted can be used as the input of the data inference model.
[0024] Among them, the data to be predicted can be single-modal data, such as text, video, audio, etc. Or, the data to be predicted can be multi-modal data, that is, the data to be predicted can include information expressed, communicated, and understood in multiple different forms or perceptual channels, such as at least two of vision, hearing, text, touch, etc.
[0025] S102: Input the data to be predicted into the data inference model, use the data inference model to predict the subsequent data of the data to be predicted, and output the subsequent data; among them, the data inference model is obtained by training using the training method of the data inference model.
[0026] In this embodiment, the subsequent data represents the data predicted for the possible subsequent continuation of the data to be predicted based on the context meaning of the data to be predicted.
[0027] The data inference model is a model with data inference function, which can be trained based on models such as LLM. In this embodiment, a training method for the data inference model is also proposed to improve the robustness of the data inference model. Among them, robustness refers to the characteristic of the control system to maintain certain performance under a certain (structure, size) parameter perturbation. In this embodiment, it can represent the robustness, strength, and reliability of the data inference model.
[0028] The training method of the data inference model may at least include obtaining original training data and dividing it into original data segments; obtaining redundant noise data and inserting it between the original data segments to form an enhanced data sequence of the original training data; inputting the enhanced data sequence into the data inference model to obtain the inference output of the data inference model; and using the inference output to update the loss of the data inference model. The detailed working principle of the training method of the data inference model will be elaborated in detail later, and will not be elaborated here.
[0029] Further, when using the data inference model to predict the subsequent data of the data to be predicted, the data to be predicted can be divided into target data segments and target data blocks. Among them, the target data block includes at least one target data segment.
[0030] Still further, the performance of the data inference model can be improved by a reasonable memory method such as streaming memory. The memory method in the following embodiments will be exemplified and elaborated.
[0031] Optionally, a local attention window of the target data segment and its target association region can be formed, and the attention mask of the target data segment within the local attention window can be predicted as local information; a first number of target data blocks are selected, and the average attention score of the target data segments therein is predicted; based on the average attention score, within the first number of target data blocks, a second number of target data segments are selected as key data segments; the global attention of the key data segments to the data to be predicted is evaluated as global information; the local information and the global information are fused to predict the subsequent data. Generally speaking, a local-global hybrid attention mechanism can be used for data inference in this embodiment. When calculating the local information, that is, the local attention, each token (i.e., the target data segment) can focus on its adjacent window. For example, the adjacent window to be focused on can be calculated by the formula w = 512 / 1024 tokens to form a local attention window of a fixed size, and the attention mask can be stored using a block sparse matrix, so that the computational complexity can be reduced from O(n 2) It is reduced to O(nw); where n represents the number of tokens. When calculating the global information, i.e., the global attention, the average attention score of the [CLS] vector of each first number K of consecutive target data blocks can be calculated, and the second number M of tokens with the highest scores within the K target data blocks are selected as the global attention centers to achieve dynamic selection of key tokens to participate in enhancing the global attention evaluation of the data sequence. Other tokens that are not used as global attention centers can be considered non-key tokens, and non-key tokens can only participate in the calculation of local window attention. The input enhanced data sequence can be divided into two types of regions: local windows and global key tokens. The data inference model can calculate local attention and global attention in parallel, and the results are merged through a gating mechanism. The way of merging through the gating mechanism will be exemplified and elaborated later.
[0032] At the same time, different from the traditional sliding windows and global tokens of models such as Longformer (a deep learning model for processing long documents), the key tokens participating in the global attention calculation in this embodiment are dynamically generated according to the content. That is, in this embodiment, by calculating and selecting the M tokens with the highest scores to participate in the global calculation, the situation of manual preset global token deviation can be reduced.
[0033] And / or, optionally, the target data segment currently concerned by the inference can be used as the current data segment; the attention information of the third number of target data segments having a preset association relationship with the current data segment is retained as hot access information for access; the fourth number of other target data segments are dimensionally reduced as cold access information, and the reconstructed information matching the original information dimension is reconstructed for the cold access information. Generally speaking, taking the KV (Key-Value) matrix attention mask as an example, the last third number L1 of tokens can be used as hot access information, and their complete KV matrix is retained for high-frequency access; the historical fourth number L2 of tokens are used as non-key blocks, i.e., cold access information, and their KV matrix is dimensionally reduced by PCA (Principal Component Analysis) or the like, and the cold access information tokens with lower attention scores are clustered, and the cluster centers are retained as representations and stored. Thus, through dimensional reduction, the memory occupancy can be reduced, which is beneficial to ensuring the inference performance and reliability of the data inference model, and further improving the robustness of the data inference model. Further, when recalling and compressing the cold access information, the approximate value of the original dimension can be reconstructed by means of MLP (Multilayer Perceptron) or the like.
[0034] And / or, optionally, a target data block of new input can be obtained as an updated data block; calculate the similarity between the updated data block and the historical data block, and retain the historical data blocks with similarity higher than the similarity threshold as associated data blocks; evaluate the attention scores of the associated data blocks, and perform recombination based on the attention scores so that the attention scores of the associated data blocks and / or data segments therein are proportional to the information dimension. Generally speaking, when a new target data block is input, the similarity between its CLS vector and the target data block (i.e., the historical data block) in the cache can be calculated. Among them, CLS can represent the identifier of the target data block. Retain the historical data blocks with similarity higher than the similarity threshold, and eliminate or compress other historical data blocks. And, for the retained historical data blocks, weighted sufficiency can be performed according to the attention, so that high-score tokens can retain more information dimensions.
[0035] That is to say, when performing long data KV caching, it requires a large memory footprint. In this embodiment, to further optimize the performance of the data inference model, accelerate the data inference speed, and save memory occupancy, streaming memory management is also introduced. During the process of caching data, the KV matrix of historical blocks can be dynamically cached, and the matrix can be compressed as needed to achieve dynamic update of the KV matrix. For example, when performing a caching interaction task or the length of the data to be predicted exceeds the word processing limit of the data inference model, the streaming memory management method can be introduced to achieve operations such as the aforementioned cache compression and elimination, optimize the data caching strategy to reduce the resources consumed by data caching, which is conducive to ensuring the data inference performance and can improve the reliability and efficiency of the data inference model to output consecutive data. Further, the streaming memory management can be used in at least one of the cases of resource limitation, ultra-long data processing (data to be predicted with a length greater than the data extreme value), and unpredictable input length. In other cases, the KV matrix can be completely saved, which is not limited here.
[0036] The following gives an example to illustrate the specific implementation manner of using the gating mechanism to merge local attention and global attention.
[0037] The local attention and global attention of the input can be learned, and the gating weights can be calculated. And the gating weights are merged and output with the local attention and global attention to fuse local information and global information to obtain gating attention. The specific process expression for fusing local information and global information can be as follows: Equation 1-1 Equation 1-2 Among them, g represents the gating weight, g ∈ [0,1]; σ represents the sigmoid function (activation function); W g represents the learnable parameter, equivalent to a linear layer; h localRepresents local attention; h global Represents global attention; b g Represents learnable parameters; h out Represents gated attention.
[0038] Thus, in this embodiment, the gated mechanism can be used to suppress redundant information and enhance key long-range interactions, so as to enhance the robustness of the data inference model.
[0039] As described above, the embodiments of the present application provide a training method for a data inference model. Combining the execution process of the training method for the data inference model, the training method for the data inference model will be described in detail.
[0040] Please refer to Figure 3 , Figure 3 , which is a schematic flowchart of an embodiment of the training method for the data inference model of the present application.
[0041] S201: Obtain the original training data and divide it to form original data segments.
[0042] In this embodiment, the original training data represents data such as documents and articles used to train the data inference model. It should be noted that the original training data can be unprocessed data content, or it can be data content that has been subjected to certain preprocessing, and no strict limitation is made here.
[0043] The original data segment represents a partial data content formed by dividing and cutting the original training data.
[0044] That is to say, in this embodiment, when training the data inference model, after obtaining the original training data, the original training data can be divided into multiple original data segments, and the ability of the data inference model to extract context dependencies therein can be trained using the original data segments, so as to enable the data inference model to output inference content that conforms to context dependencies.
[0045] S202: Obtain redundant noise data and insert it between the original data segments to form an enhanced data sequence of the original training data.
[0046] In this embodiment, as the name implies, the redundant noise data represents redundant data content, and its main function is to participate in the training process of the data inference model as noise.
[0047] Inserting the redundant noise data between the original data segments in this embodiment can form an enhanced data sequence of the original training data conversion. Thus, the original data segments that are adjacent in the original training data can have their distance extended due to the insertion of the redundant noise data, so as to convert the short-distance dependencies between adjacent original data segments into long-distance dependencies.
[0048] For example, at least one redundant noise data can be inserted between every two adjacent original data segments; alternatively, at least one group of redundant noise data can be randomly selected and inserted therebetween, which will not be elaborated herein. The redundant noise data can be any data segment or a data segment having a certain correlation with the original training data, which is not limited herein.
[0049] S203: Input the enhanced data sequence into the data inference model to obtain the inference output of the data inference model; use the inference output to update the loss of the data inference model.
[0050] In this embodiment, inputting the enhanced data sequence inserted with redundant noise data into the data inference model to train the data inference model using the enhanced data sequence can enhance the data inference model's recognition and application capabilities for long-distance dependencies, facilitate the output content of the data inference model to pay attention to long-distance dependencies, and improve the long-distance consistency of the data inference content.
[0051] During the training process of the data inference model, it can perform the data inference process based on the enhanced data sequence and output the inference output obtained therefrom. Thus, during the training process of the data inference model, the inference output can be used to update the loss of the data inference model, thereby adjusting the model parameters of the data inference model and realizing the training optimization of the data inference model.
[0052] That is to say, redundant noise data is inserted between the original data segments of the original training data to form an enhanced data sequence that extends the length of the data sample. Between the original data segments with short-distance dependencies in the original training data, since the insertion of redundant noise data extends the distance therebetween, the short-distance dependencies can be converted into long-distance dependencies. Thus, when training the data inference model using the enhanced data sequence, the data inference model's processing ability for long-distance dependencies can be trained, thereby improving the robustness of the data inference model, optimizing the effect of the data inference model for data inference, and further improving the reliability of prediction content such as data prediction. Therefore, this embodiment can enhance the processing performance of the data inference model for long-distance dependencies to achieve the technical effect of improving the robustness of the data inference model.
[0053] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another embodiment of the training method of the data inference model of the present application.
[0054] S301: Obtain the original training data and divide it into original data segments.
[0055] In this embodiment, to obtain the original training data, the original training data can be divided into multiple consecutive and semantically coherent original data blocks, and one or more original data segments can be included in each original data block.
[0056] Among them, dividing the original data into blocks can make the original data blocks conform to the sequential dependence that the latter original data block depends on the former one, and can also avoid truncating sentences or key entities to facilitate ensuring semantic integrity, which is beneficial for the data inference model to perform reliable inference and training, can reduce the redundant repair operation amount or the risk of unreliable training of the data inference model due to semantic incompleteness, and thus can improve the robustness of the data inference model.
[0057] S302: Obtain the preset data set.
[0058] In this embodiment, the preset data set can be pre-constructed so that the preset data set can be obtained when training the data inference model, thus eliminating the need to immediately construct the redundant noise data required for training, which is beneficial for improving the training efficiency of the data inference model, and focusing mainly on the training process of the data inference model itself during the training process of the data inference model, and can improve the reliability of the training process.
[0059] S303: Retrieve the preset data set to obtain alternative data segments.
[0060] In this embodiment, the preset data segments that conform to the expected association relationship with the original data segments in the preset data set can be retrieved as alternative data segments. Among them, the expected association relationship means having semantic similarity with the original data segments and semantic contradiction with the original training data. As the name implies, the alternative data segments are the preset data segments that are alternative as redundant noise data.
[0061] That is to say, in this embodiment, an implementation method is exemplified in which data segments having a certain relevance to the original training data are selected as redundant noise data. Moreover, in this embodiment, the preset data segments that have a certain semantic similarity with the original data segments or the original training data but have context contradictions are selected as alternative data segments, so that difficult negative samples can be introduced during the training process of the data inference model, increasing the misleading content in the enhanced data sequence to simulate the possible interference in long data tasks in the real scenario, which is beneficial for further improving the robustness of the data inference model.
[0062] Meanwhile, data inference models generally tend to capture coarse-grained semantic differences and have difficulty in identifying subtle contradictions between key details. By introducing preset data segments with semantic similarity and semantic contradictions in this embodiment, it is beneficial to train the data inference model's ability to distinguish subtle semantic differences, enabling it to distinguish key contexts from interfering information, thereby enhancing the robustness of the data inference model.
[0063] S304: Evaluate the duplication factor between the alternative data segment and the original data segment.
[0064] In this embodiment, the duplication factor between the alternative data segment and the original data segment can be evaluated to select the alternative data segment based on the duplication factor.
[0065] For example, evaluating the duplication factor can be done through similarity calculation methods such as cosine similarity, word vector similarity, deep learning-based similarity, etc., which will not be elaborated here.
[0066] S305: Select the alternative data segments with a duplication factor lower than the preset duplication threshold as redundant noise data.
[0067] In this embodiment, the duplication factor between the alternative data segment and the original data segment can be compared with the preset duplication threshold. In response to the duplication factor being lower than the preset duplication threshold, the alternative data segment is used as redundant noise data to obtain redundant noise data. This is equivalent to de-duplicating and validating the preset data segments in this embodiment to ensure that the redundant noise data, which is a difficult negative sample, has a duplication degree lower than the preset duplication threshold with the original training data, thereby reducing the interference to the training process of the data inference model with a high duplication degree and further improving the training reliability of the data inference model.
[0068] S306: Insert the redundant noise data between the original data segments to form an enhanced data sequence of the original training data.
[0069] In this embodiment, in response to selecting redundant noise data from the preset data set, the redundant noise data can be inserted between the original data segments to extend the segment spacing between adjacent original data segments, thereby converting short-distance dependencies into long-distance dependencies, enhancing the data inference model's ability to recognize and process long-distance dependencies during training, and further enhancing the robustness of the data inference model to improve the reliability of its output results.
[0070] Optionally, the original data segments may be formed into an original data sequence according to their order within the original training data. In the original data sequence, a randomly selected redundant noise data may be inserted between adjacent original data segments to form an enhanced data sequence. That is to say, in this embodiment, the redundant noise data serving as hard negative samples may be randomly positioned in the enhanced data sequence, which helps to reduce the risk that the data inference model simply relies on the local part, thereby further improving the robustness of the data inference model.
[0071] As Figure 5 exemplified in Figure 5 is a schematic flowchart of an embodiment for forming an enhanced data sequence in this application. Hard negative samples, that is, redundant noise data, may be inserted between the original training data (Meta-chunk1~Meta-chunk3) with original order dependence to construct an enhanced training sequence. As Figure 5 the CLS in
[0072] previously described may represent the identifier of the target data block (the original data block in this embodiment), and Neg-text represents the original data segment. For another example, the original data sequence formed after the original training data is divided into original data segments may be (A, B, C), and the enhanced data sequence formed after inserting redundant noise data as hard samples may be (A, negative sample 1, B, negative sample 2, C).
[0073] S307: Perform global position encoding processing and local position encoding processing on the original data segments to obtain hierarchical position encoding.
[0074] In this embodiment, multi-level position encoding may be performed on the original data segments, which helps to locate the original data segments through position encoding, so as to solve the position perception problem during long data extrapolation and further improve the robustness of the data inference model.
[0075] Optionally, the original training data may be divided into original data blocks; wherein, the original data block includes at least one original data segment. Perform local position encoding processing on the original data segments to obtain local encoding. Perform global position encoding processing on the original data segments to obtain global encoding. Fit the local encoding and the global encoding to obtain the hierarchical position encoding of the original data segments.
[0076] Generally speaking, the relative position of a token within a data block can be perceived as its local position; taking the token length of 512 as an example, the range of relative positions is 0 to 511. It is also possible to sense the sequential number of the current original data block in the original training document or the augmented data sequence (such as the 3rd block, etc.) as the global position. In this way, the hierarchical position encoding of the token at position i is obtained by fusing its global position and local position. For example, the calculation formula for obtaining its hierarchical position encoding can be as follows: PE(i)=PE local (i mod L)+ PE global ([i / L]) Equation 2-1 Where, PE(i) represents the hierarchical position encoding of the i-th token; mod represents the remainder function; L represents the token length of the data block; PE local (i mod L) represents the local encoding; PE global ([i / L]) represents the global encoding.
[0077] In an alternative embodiment, multi-level position encoding can also be performed on redundant noise data together to obtain the hierarchical position encoding of the redundant noise data, which is not limited here.
[0078] S308: Input the augmented data sequence into the data inference model to obtain the inference output of the data inference model.
[0079] In this embodiment, inputting the augmented data sequence with inserted redundant noise data into the data inference model to train the data inference model using the augmented data sequence can enhance the data inference model's ability to recognize and utilize long-distance dependencies, which is beneficial for the output content of the data inference model to pay attention to long-distance dependencies and improve the long-distance consistency of the data inference content.
[0080] Obtaining the inference output of the data inference model based on the augmented data sequence during the training process, the inference output can be used to iteratively update the data inference model to optimize the data inference performance of the data inference model.
[0081] S309: Use the inference output to update the loss of the data inference model.
[0082] In this embodiment, the first prediction loss between the inference output and the target output can be evaluated.
[0083] By using both the similarity between original data segments and the similarity between global data segments, the training temperature coefficient is fused to form a second prediction loss. This is to make the second prediction loss mainly focus on the similarity between the original data segments originally existing in the original training data and weaken the similarity with redundant noise data, so as to further improve the learning of long-distance dependencies between original data segments. Further, a preset weighting factor can be used to weight the second prediction loss, and the weighted result is superimposed with the first prediction loss to obtain a training loss function. In this way, the data inference model can be updated using the training loss function.
[0084] Generally speaking, language modeling and contrastive learning tasks can be jointly optimized to strengthen the modeling ability of long-distance dependencies. In the main training task of the data inference model, language modeling and data inferences such as predicting the next data segment are carried out, and the standard autoregressive loss is calculated. In the training auxiliary task, contrastive learning between data blocks can also be carried out, hoping to maximize the similarity between original data blocks and minimize the similarity with redundant noise data, so as to obtain a training loss function for updating the data inference model. The specific expression formula for obtaining the training loss function can be as follows: Equation 3-1 Equation 3-2 Equation 3-3 Among them, L represents the training loss function; L ntp represents the first prediction loss between the inference output and the target output, which is used to train and extend the context length of the basic model of the data inference model; α represents a hyperparameter that adjusts the proportion of the contrastive learning loss; L contrastive represents the contrastive learning loss function; T represents the number of data segments in the enhanced data sequence or the original training data; q represents the anchor block, which can be considered as the first CLS vector; k + represents the vector of the original data segment, that is, the subsequent Meta-chunk; k i represents the i-th data segment (original data segment or redundant noise data) in the enhanced data sequence; τ represents the temperature coefficient of contrastive learning, which is used to control the distinguishability of redundant noise data.
[0085] In summary, the present application can improve the data inference ability of the data inference model, support an input length of more than 128K (kilobytes) tokens, such as for ultra-long data tasks like contract analysis, long dialogue systems, and long articles. It can also reduce the video memory occupancy by about 50% through sparsification and memory compression models, improve the training efficiency and data inference efficiency of the data inference model, and enhance the robustness of the data inference model.
[0086] Specifically, when dividing data blocks, the original position information can be retained for each sentence, and dynamic overlapping can be performed to reduce the risk of boundary effects, thereby achieving hierarchical chunking. Moreover, local semantics can be learned, and cross-chunk relationship modeling can be enhanced through contrastive learning, sequential prediction, and strengthening of data structure understanding to effectively achieve multi-task collaboration. Meanwhile, the data inference surface model's ability to capture long-distance dependencies can be enhanced by using difficult negative samples such as chunking and redundant noisy data. During the data inference process, efficient long data processing can also be achieved through dynamic sparse attention and streaming memory.
[0087] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0088] The embodiments of the present application also provide a data prediction device.
[0089] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of an embodiment of the data prediction device of the present application.
[0090] In one embodiment, the data prediction device may include an input module 21 and a control module 22.
[0091] The control module 22 is connected to the input module 21. The control module 22 can be used to implement the steps of the training method of the data inference model in any of the above embodiments; or, it can implement the steps of the data prediction method in any of the above embodiments.
[0092] For the description of the features in the corresponding embodiments of the data prediction device, reference can be made to the relevant descriptions of the corresponding embodiments of the training method of the data inference model and the data prediction method, which will not be elaborated here one by one.
[0093] The embodiments of the present application also provide an electronic device.
[0094] The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the computer program to implement the steps in any of the above embodiments of the training method of the data inference model, or to implement the steps in any of the above embodiments of the data prediction method.
[0095] For example, as Figure 7 illustrated by way of example in Figure 7 which is a schematic flowchart of an embodiment of the electronic device of the present application. The electronic device can be a server or the like.
[0096] The electronic device includes a processor, a memory, a network interface, and a database connected through a system bus.
[0097] Among them, the processor of the electronic device is used to provide computing and control capabilities.
[0098] The memory of the electronic device includes a non-volatile storage medium and an internal memory. Among them, the non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium.
[0099] The database of the electronic device is used to store data. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it is used to implement a training method or a prediction method of a data inference model.
[0100] An embodiment of the present application further provides a computer-readable storage medium.
[0101] A computer program is stored in the computer-readable storage medium. Among them, the computer program is set to implement the steps in any of the above-described training method embodiments of the data inference model or the steps in any of the above-described prediction method embodiments of the data when being executed by the processor.
[0102] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0103] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program. When the computer program is executed by the processor, it implements the steps in any of the above-described training method embodiments of the data inference model or the steps in any of the above-described prediction method embodiments of the data.
[0104] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program. When the computer program is executed by the processor, it implements the steps in any of the above-described training method embodiments of the data inference model or the steps in any of the above-described prediction method embodiments of the data.
[0105] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0106] The above has introduced in detail a method for training a data inference model, a data prediction method, a data prediction device, an electronic device, and a computer-readable storage medium provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A training method for a data inference model, characterized in that, The training method includes: Obtain the original training data and divide it to form original data segments; Obtain redundant noise data and insert it between the original data segments to form an enhanced data sequence of the original training data; Input the enhanced data sequence into a data inference model to obtain the inference output of the data inference model; use the inference output to update the loss of the data inference model.
2. The training method according to claim 1, characterized in that The obtaining of the redundant noise data includes: Obtain a preset data set; Retrieve preset data segments in the preset data set that conform to the expected association relationship with the original data segments as alternative data segments; wherein, the expected association relationship means having semantic similarity with the original data segments and semantic contradiction with the original training data; Evaluate the duplication factor between the alternative data segments and the original data segments; in response to the duplication factor being lower than a preset duplication threshold, use the alternative data segments as the redundant noise data.
3. The training method according to claim 1 or 2, characterized in that, The inserting it between the original data segments to form an enhanced data sequence of the original training data includes: Arrange the original data segments in the order within the original training data to form an original data sequence; In the original data sequence, insert a randomly selected redundant noise data between adjacent original data segments to form the enhanced data sequence.
4. The training method according to claim 1, characterized in that Before inputting the enhanced data sequence into the data inference model, it further includes: Divide the original training data to form original data blocks; wherein, the original data blocks include at least one of the original data segments; Perform local position encoding processing on the original data segments to obtain local encoding; Perform global position encoding processing on the original data segments to obtain global encoding; Fit the local encoding and the global encoding to obtain the hierarchical position encoding of the original data segments.
5. The training method according to claim 1, wherein The using the inference output to update the loss of the data inference model includes: Evaluate the first prediction loss between the inference output and the target output; Use the similarity between the original data segments and the similarity between global data segments, and fuse the training temperature coefficient to form a second prediction loss; Weight the second prediction loss using a preset weighting factor, superimpose the weighted result and the first prediction loss to obtain a training loss function; Use the training loss function to update the data inference model.
6. A data prediction method, characterized in that, The data prediction method includes: Obtain data to be predicted; Input the data to be predicted into a data inference model, and use the data inference model to predict the subsequent data of the data to be predicted and output the subsequent data; wherein, the data inference model is trained using the training method of the data inference model according to any one of claims 1 to 5.
7. The data prediction method according to claim 6, wherein The using the data inference model to predict the subsequent data of the data to be predicted includes: Divide the data to be predicted to form target data segments and target data blocks; wherein, the target data blocks include at least one of the target data segments; Form a local attention window for the target data segment and its target associated area, predict the attention mask of the target data segment within the local attention window as local information; select a first number of the target data segments, and predict the average attention score of the target data blocks therein; based on the average attention score, select a second number of the target data segments within the first number of the target data blocks as key data segments; evaluate the global attention of the key data segments to the data to be predicted as global information; fuse the local information and the global information to predict the subsequent data; and / or, Take the target data segment currently under inference attention as the current data segment; retain the attention information of a third number of target data segments having a preset association relationship with the current data segment as hot access information for access; perform dimensionality reduction processing on a fourth number of other target data segments as cold access information, and reconstruct the cold access information into reconstruction information that matches the original information dimension; and / or, Obtain a newly input target data block as an update data block; calculate the similarity between the update data block and the historical data blocks, and retain the historical data blocks with a similarity higher than the similarity threshold as associated data blocks; evaluate the attention scores of the associated data blocks, and perform recombination based on the attention scores so that the attention scores of the associated data blocks and / or the data segments therein are proportional to the information dimension.
8. A data prediction device, characterized in that, The data prediction device includes: An input module; A control module, connected to the input module; for implementing the steps of the training method of the data inference model according to any one of claims 1 to 5, or for implementing the steps of the data prediction method according to any one of claims 6 to 7.
9. An electronic device, characterized in that, The electronic device includes: A memory for storing a computer program; A processor for implementing the steps of the training method of the data inference model according to any one of claims 1 to 5 when executing the computer program; or for implementing the steps of the data prediction method according to any one of claims 6 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the training method of the data inference model according to any one of claims 1 to 5; or implements the steps of the data prediction method according to any one of claims 6 to 7.
Citation Information
Patent Citations
Text matching method based on data enhancement and graph matching network
CN115510841A
Transform model optimization method and system based on dynamic window and storage medium
CN117992567A
Smoke image detection method based on split Top-K attention mechanism
CN119107508A
Language model training method, natural language processing method and device
CN119227809A
Highly unbalanced text classification-oriented enhanced contrast learning method and device
CN119336916A