Large language model generation text continuous traceability model training method and device
The continuous traceability model training method of text generated by large language models is used to solve the frequent retraining problem caused by fixed label sets in the existing technology, and efficient and reliable text traceability is achieved.
Patent Information
- Application Number
- CN202510646849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing large language model generation text traceability method relies on a static classifier with fixed label sets and cannot automatically recognize text generated by the newly emerging large language model, resulting in frequent retraining processes, time-consuming and computing resources, and cannot keep up with the rapid development of large language models.
The training method of text continuous traceability model is used to generate large language model. Through the stage feature extraction step, the feature vectors of the current training stage are extracted, and the decorrelation prototype is generated through global and local decorrelation processing to generate a continuous traceability model for predicting the type of large language model that generates text data.
It effectively solves the problem of frequent retraining caused by the fixed label set of traditional traceability methods, improves model training efficiency, reduces resource consumption, and improves the reliability and effectiveness of traceability results.
Smart Images

Figure CN120179812A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer systems based on specific computational models, and particularly to a method and device for training a continuous traceability model for texts generated by large language models. Background Art
[0002] With the rapid growth of the content generated by large language models (LLMs), the potential risk of spreading false information also increases. Accurately identifying and tracing the source of texts generated by large language models is of great significance for ensuring responsibility division, enhancing the content verification process, and improving the transparency of information dissemination. Traditional research on tracing the source of texts generated by large language models mainly formulates this task as a binary classification problem, aiming to distinguish whether a given text is generated by a large language model or written by a human; with the in-depth research, the focus has shifted from simple detection to tracing the large language model that generates the text, usually formulating it as a multi-classification problem, making the task more practical but also more challenging.
[0003] Tracing the source of texts generated by large language models aims to determine whether a text is generated by a large language model and further identify which specific large language model generated the text. In existing methods for tracing the source of texts generated by large language models, this task is usually framed as a multi-class classification problem with a fixed label set, and the label set consists of several predefined large language models and a human category.
[0004] However, existing methods for tracing the source of texts generated by large language models usually rely on static classifiers with fixed label sets. Since the label set is fixed, the classifier cannot automatically identify texts generated by newly emerging large language models. Therefore, whenever a new large language model appears, the classifier with a fixed label set containing the new large language model must be reset and retrained using data related to existing and new categories. This repetitive retraining process is not only time-consuming and resource-consuming but also unable to keep up with the rapid development of large language models. Summary of the Invention
[0005] In view of this, embodiments of this application provide a method and device for training a continuous traceability model for texts generated by large language models to eliminate or improve one or more deficiencies existing in the prior art.
[0006] One aspect of this application provides a method for training a continuous traceability model for texts generated by large language models, including: Phase Feature Extraction Step: Input each training sample in the dataset corresponding to the current training phase during the current model training process into the feature extraction unit, so that the feature extraction unit extracts the feature vectors corresponding to the text data in each of the training samples respectively; wherein, each of the training samples further includes a label for indicating the type of the large language model that generates the text data; the release time of the large language model corresponding to the current training phase is later than the release times of the large language models corresponding to each of the historical training phases; and, according to the feature vectors of each of the text data and the labels corresponding to each of the text data, respectively obtain the initial prototypes and text feature correlation data of each of the large language models. If the current training phase is the last training phase during the current model training process, then perform global decorrelation processing and local decorrelation processing on the initial prototypes obtained in sequence for each of the historical training phases and the current training phase according to the text feature correlation data of each of the large language models obtained in sequence for each of the historical training phases and the current training phase, to obtain the decorrelated prototypes corresponding to each of the large language models, so as to generate a large language model generation text continuous traceability model for predicting the type of the large language model that generates text data and including the feature extraction unit based on the current decorrelated prototypes and the label set including the labels corresponding to each of the decorrelated prototypes.
[0007] In some embodiments of the present application, before the phase feature extraction step, it further includes: Collect each text data written by humans and each text data generated by each currently released large language model. Use each text data written by humans as the text data of the initial large language model, and set labels for indicating the type of the large language model for the initial large language model and each of the released large language models respectively. Use the label of the initial large language model as the first label, and sort the labels of each of the released large language models in sequence after the label of the initial large language model in the order of the release time from early to late, so as to obtain the corresponding label set. Construct the datasets corresponding to each of the labels in sequence according to the order of the labels in the label set, each dataset includes the text data corresponding to the label, and each text data in the dataset and the label corresponding to it respectively form different training samples.
[0008] In some embodiments of the present application, before the phase feature extraction step, it further includes: If a large language model that has not been applied during the historical training period is detected or received, set a label for uniquely representing the type to which the large language model belongs, and obtain a plurality of text data generated using the large language model to generate a dataset for the current training stage; wherein, the dataset for the current training stage contains a plurality of the training samples.
[0009] In some embodiments of the present application, the feature extraction unit includes: a pre-trained language model with fixed parameters and a randomly upward projection layer with fixed parameters; Correspondingly, the step of respectively inputting each training sample in the dataset corresponding to the current training stage in the current model training process into the feature extraction unit, so that the feature extraction unit respectively extracts the feature vectors corresponding to the text data in each of the training samples, includes: Respectively input the text data corresponding to each training sample in the dataset corresponding to the current training stage in the current model training process into the pre-trained language model, so that the pre-trained language model respectively outputs the initial feature vectors corresponding to each of the text data, and then enables the randomly upward projection layer to perform an upward projection process on each of the initial feature vectors to obtain the feature vectors corresponding to each of the text data.
[0010] In some embodiments of the present application, the step of respectively obtaining the initial prototypes and text feature correlation data of each large language model according to the feature vectors of each of the text data and the labels corresponding to each of the text data, includes: According to the labels corresponding to each of the text data, perform a summation calculation on the feature vectors of each of the text data belonging to the same large language model to respectively obtain the initial prototypes of each large language model; And, according to the labels corresponding to each of the text data, perform a Gram matrix calculation on the feature vectors of each of the text data belonging to the same large language model to respectively obtain the Gram matrices corresponding to each large language model as the text feature correlation data.
[0011] In some embodiments of the present application, the step of performing global decorrelation processing and local decorrelation processing on the initial prototypes obtained in sequence for each of the historical training stages and the current training stage according to the text feature correlation data of each large language model obtained in sequence for each of the historical training stages and the current training stage, to obtain the decorrelated prototypes corresponding to each large language model, includes: Obtain the text feature correlation data of each of the large language models obtained in sequence according to each of the historical training stages and the current training stage, and perform global decorrelation processing on each of the initial prototypes obtained in sequence according to each of the historical training stages and the current training stage, so as to obtain the prototypes with global correlation eliminated corresponding to each of the large language models; And, based on the text feature correlation data of each of the large language models obtained in sequence according to each of the historical training stages and the current training stage, perform local decorrelation processing on every two of the initial prototypes respectively, so as to obtain the prototypes with local correlation eliminated corresponding to each of the large language models; Respectively use the prototypes with global correlation eliminated and the prototypes with local correlation eliminated corresponding to each of the large language models as the decorrelated prototypes corresponding to each of the large language models.
[0012] The second aspect of the present application provides a method for continuously tracing the source of text generated by a large language model, including: Input the target text data into the large language model text generation continuous traceability model trained by the large language model text generation continuous traceability model training method described in the first aspect above, so that the feature extraction unit in the large language model text generation continuous traceability model extracts the target feature vector corresponding to the target text data, so that the large language model text generation continuous traceability model matches the target feature vector with each of the decorrelated prototypes respectively to obtain the target prototype corresponding to the target text data, so that the large language model text generation continuous traceability model searches for the label corresponding to the target prototype in the label group and outputs the label as the traceability prediction result of the target text data.
[0013] In some embodiments of the present application, the decorrelated prototype includes: a prototype with global correlation eliminated and a prototype with local correlation eliminated; The matching of the target feature vector with each of the decorrelated prototypes respectively to obtain the target prototype corresponding to the target text data includes: Match the target feature vector with each of the prototypes with global correlation eliminated respectively to obtain the global matching scores corresponding to each of the prototypes with global correlation eliminated; Sort the prototypes with global correlation eliminated in descending order of the global matching scores, and use the top two prototypes with global correlation eliminated as candidate prototypes; Match the target feature vector with the prototypes with local correlation eliminated corresponding to each of the candidate prototypes respectively to obtain the local matching scores corresponding to the two candidate prototypes; Determine the candidate prototype with the highest local matching score as the target prototype corresponding to the target text data.
[0014] The third aspect of this application provides a training device for a large language model generated text continuous traceability model, including: A stage feature extraction model for performing stage feature extraction steps: inputting each training sample in the dataset corresponding to the current training stage in the current model training process into a feature extraction unit respectively, so that the feature extraction unit extracts the feature vectors corresponding to the text data in each of the training samples respectively; wherein, each of the training samples further includes a label for indicating the type of the large language model that generates the text data; the release time of the large language model corresponding to the current training stage is later than the release times of the large language models corresponding to each historical training stage; and, according to the feature vectors of each of the text data and the labels corresponding to each of the text data respectively, obtain the initial prototypes of each of the large language models and the text feature correlation data respectively; A global and local decorrelation module, if the current training stage is the last training stage in the current model training process, then perform global decorrelation processing and local decorrelation processing on the initial prototypes obtained in each of the historical training stages and the current training stage in sequence according to the text feature correlation data of each of the large language models obtained in each of the historical training stages and the current training stage, to obtain the decorrelated prototypes corresponding to each of the large language models, so as to generate a large language model generated text continuous traceability model for predicting the type of the large language model that generates text data and including the feature extraction unit based on the current decorrelated prototypes and a label group including the labels corresponding to each of the decorrelated prototypes.
[0015] The fourth aspect of this application provides a large language model generated text continuous traceability device, including: A model prediction module for inputting target text data into the large language model generated text continuous traceability model trained by the large language model generated text continuous traceability model training method, so that the feature extraction unit in the large language model generated text continuous traceability model extracts the target feature vector corresponding to the target text data, so that the large language model generated text continuous traceability model matches the target feature vector with each of the decorrelated prototypes respectively to obtain the target prototype corresponding to the target text data, so that the large language model generated text continuous traceability model searches for the label corresponding to the target prototype in the label group and outputs the label as the traceability prediction result of the target text data.
[0016] The fifth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for training the continuous traceability model of the text generated by the large language model and / or the method for continuous traceability of the text generated by the large language model are implemented.
[0017] The sixth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for training the continuous traceability model of the text generated by the large language model and / or the method for continuous traceability of the text generated by the large language model are implemented.
[0018] The seventh aspect of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for training the continuous traceability model of the text generated by the large language model and / or the method for continuous traceability of the text generated by the large language model are implemented.
[0019] The method for training the continuous traceability model of the text generated by the large language model provided by the present application includes a step of extracting stage features: respectively inputting each training sample in the dataset corresponding to the current training stage in the current model training process into a feature extraction unit, so that the feature extraction unit respectively extracts the feature vectors corresponding to the text data in each of the training samples; wherein, each of the training samples further includes a label for indicating the type of the large language model that generates the text data; the release time of the large language model corresponding to the current training stage is later than the release times of the large language models corresponding to each of the historical training stages; and, respectively obtaining the initial prototypes and text feature correlation data of each of the large language models according to the feature vectors of each of the text data and the labels corresponding to each of the text data; if the current training stage is the last training stage in the current model training process, then perform global decorrelation processing and local decorrelation processing on the initial prototypes obtained in sequence for each of the historical training stages and the current training stage according to the text feature correlation data of each of the large language models obtained in sequence for each of the historical training stages and the current training stage, to obtain the decorrelated prototypes corresponding to each of the large language models, so as to generate a large language model continuous traceability model for predicting the type of the large language model that generates text data and including the feature extraction unit, which can solve the problem of frequent retraining caused by a fixed label set in traditional traceability methods, can effectively improve the model training efficiency and reduce resource consumption, and can improve the reliability and effectiveness of the traceability results.
[0020] Additional advantages, objects, and features of the present application will be partly set forth in the description which follows, and will partly become obvious to those of ordinary skill in the art after study of the following, or may be learned by practice of the present application. The objects and other advantages of the present application may be realized and attained by the structure particularly pointed out in the specification and the drawings.
[0021] Those skilled in the art will understand that the objects and advantages that can be achieved by the present application are not limited to those specifically described above, and the above and other objects that the present application can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and do not limit the present application. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings: Figure 1 FIG. 11 is a first schematic flowchart of a method for training a continuous traceability model for text generated by a large language model according to an embodiment of the present application.
[0023] Figure 2 FIG. 15 is a second schematic flowchart of a method for training a continuous traceability model for text generated by a large language model according to an embodiment of the present application.
[0024] Figure 3 FIG. 19 is a third schematic flowchart of a method for training a continuous traceability model for text generated by a large language model according to an embodiment of the present application.
[0025] Figure 4 FIG. 23 is a schematic flowchart of a method for continuous traceability of text generated by a large language model according to an embodiment of the present application.
[0026] Figure 5 FIG. 27 is a schematic flowchart of a prototype matching process in a method for continuous traceability of text generated by a large language model according to an embodiment of the present application.
[0027] Figure 6 FIG. 31 is an example schematic diagram of a sequential learning process corresponding to a task paradigm for continuous traceability of text generated by a large language model provided in an application example of the present application.
[0028] FIG. 7(a) is a schematic diagram of a learning stage corresponding to a global and local prototype decorrelation continuous traceability method provided in an application example of the present application.
[0029] FIG. 7(b) is a schematic diagram of an inference stage corresponding to a global and local prototype decorrelation continuous traceability method provided in an application example of the present application.
[0030] Figure 8 Schematic diagram of the GLPD algorithm provided for an application example of this application.
[0031] Figure 9 Provided for an application example of this application Figure 9 Heat map comparing the pairwise prototype similarities of ten large language models provided for an application example of this application after global and local decorrelation.
[0032] Figure 10 Schematic diagram of the structure of the training device for continuously tracing the source of text generated by a large language model in an embodiment of this application.
[0033] Figure 11 Schematic diagram of the structure of the device for continuously tracing the source of text generated by a large language model in an embodiment of this application. Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in conjunction with the implementation manners and the accompanying drawings. Herein, the illustrative implementation manners of this application and their descriptions are used to explain this application, but do not limit this application.
[0035] Herein, it should also be noted that to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution of this application are shown in the drawings, while other details less relevant to this application are omitted.
[0036] It should be emphasized that the term "including / containing" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0037] Herein, it should also be noted that if not otherwise specified, the term "connection" in this text can refer not only to a direct connection but also to an indirect connection with an intermediate.
[0038] In the following, embodiments of this application will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0039] It should be noted that the researchers explored various text features to distinguish large language models from different sources in a supervised learning manner. First, a simple way is to directly use the pre-trained language model RoBERTa to extract text features to trace the text source. Second, by using the perplexity distribution of the text in different large language models and recording the next token probabilities of significant language model n-grams, it is possible to help identify the specific model that generated the text. In addition, by fine-tuning the Text to Text Transfer Transformer (T5) model, it can implicitly identify the source of the text when making the next word prediction, thereby improving the accuracy of classification. Additionally, using writing style representation is also an effective method, which can help distinguish the generation source of the text through style features. These technical solutions provide new ideas and solutions for tracing the source of text generated by large language models.
[0040] However, although existing methods have achieved some success in the source tracing task, they usually rely on static classifiers with fixed label sets. Since the label set is fixed, the classifier cannot automatically identify text generated by newly emerging large language models. Therefore, whenever a new large language model appears, the classifier with the fixed label set containing the new large language model must be reset and retrained using data related to the existing and new categories. This repeated retraining process is not only time-consuming and resource-consuming but also unable to keep up with the rapid development of large language models.
[0041] Based on this, to solve the problem of frequent retraining caused by the fixed label set in traditional source tracing methods, the embodiments of this application respectively provide a method for training a continuous source tracing model for text generated by a large language model, a device for training a continuous source tracing model for text generated by a large language model for executing the method, an electronic device, a computer-readable storage medium, and a computer program product, which can effectively improve the model training efficiency and reduce resource consumption, and can improve the reliability and effectiveness of the source tracing results.
[0042] Specific details are described in detail through the following embodiments.
[0043] Based on this, the embodiments of this application provide a method for training a continuous source tracing model for text generated by a large language model that can be implemented by a device for training a continuous source tracing model for text generated by a large language model. Refer to Figure 1 The method for training a continuous source tracing model for text generated by a large language model specifically includes the following content: Step 100: Phase Feature Extraction Step: Input each training sample in the dataset corresponding to the current training phase in the current model training process into the feature extraction unit respectively, so that the feature extraction unit extracts the feature vectors corresponding to the text data in each of the training samples respectively; wherein, each of the training samples further includes a label for indicating the type of the large language model that generates the text data; the release time of the large language model corresponding to the current training phase is later than the release times of the large language models corresponding to each of the historical training phases; and, according to the feature vectors of each of the text data and the labels corresponding to each of the text data, respectively obtain the initial prototypes and text feature correlation data of each of the large language models.
[0044] First, it should be noted that this application provides a task paradigm for continuous traceability of text generated by large language models. In this task paradigm, it is required to train the feature extraction unit in the continuous traceability model of text generated by large language models in stages. In each training stage, learn from each data sample in a dataset, and each training stage is independent and cannot access the data of the previous historical training stage. At the same time, when detecting or receiving a large language model that has not been applied in the historical training period, a dataset corresponding to the large language model can also be generated, and then stage training can be carried out for this dataset. Then, the model parameters of the continuous traceability model of text generated by the large language model can be updated, and thus there is no need to reset the classifier including the fixed label set of the new large language model and retrain with the data related to the existing categories and new categories.
[0045] Therefore, for the above task paradigm, step 100 of this application is applicable to the initial training of the continuous traceability model of text generated by the large language model, and is also applicable to the incremental training of the continuous traceability model of text generated by the large language model, which can further improve the model training efficiency and reduce resource consumption on the basis of ensuring the effectiveness and reliability of the continuous traceability results of text generated by the large language model.
[0046] In one or more embodiments of the present application, the current model training process refers to the process of training the traceability model for the currently acquired dataset. The current training stage refers to the training stage corresponding to the dataset when the execution stage feature extraction step is performed. After the completion of the stage feature extraction step, this training stage is changed to the historical training stage. Each of the said training stages uniquely corresponds to a dataset, and each of the said datasets contains multiple training samples. Each of the said training samples contains one of the said text data and a label indicating the type of the large language model that generates the text data; wherein, there is a one-to-one correspondence between the said label and the type of the large language model. Therefore, in the present application, the said label can represent the large language model, that is, the text data corresponding to the label is the text data generated by the large language model uniquely specified by the label.
[0047] It can be understood that the text feature correlation data refers to the data used to represent the correlation between the respective feature vectors corresponding to each of the said text data generated based on the same large language model.
[0048] Therefore, in view of the limitations of traditional traceability methods and the challenges posed by the continuous emergence of new large models, the present application uses a continuous traceability method to solve the problem, aiming to solve the problem of frequent retraining caused by the fixed label set in traditional traceability methods. It is a method of the class-incremental paradigm. In the field of text traceability of large models, the present application is the first to solve the problem of text traceability in the continuous learning paradigm.
[0049] It should be noted that continuous learning (CL), also known as incremental learning (IL), enables the model to continuously accumulate knowledge in a series of tasks without catastrophic forgetting. There are three different scenarios for continuous learning: task-incremental learning, domain-incremental learning, and class-incremental learning. Among these three scenarios, class-incremental learning requires the model to retain the knowledge of the old classes and effectively distinguish them when gradually learning new class objects, so it is considered the most challenging scenario. The task of continuous traceability of text generated by the large language model provided by the embodiments of the present application is in the class-incremental learning paradigm, and the traceability model is required to retain the knowledge of the old large language models when gradually learning the text generated by new large language models.
[0050] However, there are also certain limitations in continuous learning methods for traceability tasks. Existing continuous learning methods can be roughly divided into 4 categories: (1) replay-based methods, which restore previous knowledge by replaying samples of old classes; (2) regularization-based methods, which use knowledge distillation to consolidate the knowledge of previous tasks; (3) architecture-based methods, which construct task-specific parameters to reduce interference between tasks in sequential learning; (4) prototype-based methods, which use frozen pre-trained models to extract class prototypes independent of task order.
[0051] Replay-based methods are not considered for continual learning methods because they require saving old samples, which has the same drawbacks as traditional traceability methods. For the other three continual learning methods, preliminary experiments show that the prototype-based continual learning method has strong performance advantages in the continual traceability task. However, there is a high similarity among the prototypes of large language models directly extracted from pre-trained language models. Although some collinearity among the prototypes of large language models can be eliminated through global decorrelation techniques, it cannot consider the pairwise prototype collinearity (similarity) problem at a fine-grained level. Especially when the number of new large language models that the classification model continues to learn increases, the global decorrelation technique can only capture the overall global collinearity and cannot capture the subtle collinearity between the prototypes of two large language models.
[0052] Therefore, based on the incremental learning paradigm, embodiments of the present application also propose a method based on global and local prototype decorrelation to further solve the problem that the similarity between the prototypes of large language models is too high and the correlation between the prototypes of large language models cannot be removed at a fine-grained level. The specific process is as described in step 200 below.
[0053] Step 200: If the current training stage is the last training stage in the current model training process, then, according to the text feature correlation data of each of the large language models obtained in sequence for each of the historical training stages and the current training stage, perform global decorrelation processing and local decorrelation processing on each of the initial prototypes obtained in sequence for each of the historical training stages and the current training stage to obtain the decorrelated prototypes corresponding to each of the large language models, and based on the current decorrelated prototypes and the label set containing the labels corresponding to each of the decorrelated prototypes, generate a large language model generation text continual traceability model for predicting the type of the large language model that generates text data and including the feature extraction unit.
[0054] As can be seen from the above description, the large language model generation text continual traceability model training method provided by the embodiments of the present application can not only solve the problem of frequent retraining caused by the fixed label set in traditional traceability methods, but also solve the problem that the similarity between the prototypes of large language models is too high and the correlation between the prototypes cannot be removed at a fine-grained level. It can solve the problem of frequent retraining caused by the fixed label set in traditional traceability methods, can effectively improve the model training efficiency and reduce resource consumption, and can improve the reliability and effectiveness of the traceability results.
[0055] To further improve the effectiveness and reliability of the large language model generation text continual traceability model training, in a large language model generation text continual traceability model training method provided by the embodiments of the present application, refer to Figure 2, before step 100 in the training method of the large language model generated text continuous traceability model, the following specific content is further included: Step 010: Collect each text data written by humans and each text data generated by each currently released large language model.
[0056] Step 020: Use each text data written by humans as the text data of the initial large language model, and set labels for the initial large language model and each of the released large language models to represent the type of the large language model.
[0057] Step 030: Take the label of the initial large language model as the first label, and sort the labels of each of the released large language models in ascending order of release time after the label of the initial large language model to obtain a corresponding label set.
[0058] Step 040: According to the order of each label in the label set, successively construct data sets corresponding to each label. Each data set contains the text data corresponding to the label, and each text data in the data set and its corresponding label respectively form different training samples.
[0059] That is to say, one of the triggering situations of the stage feature extraction step can be triggered by obtaining each text data written by humans and each text data generated by each currently released large language model. Steps 100 to 300 executed after step 040 can be the initial full-scale training process for the large language model generated text continuous traceability model to ensure the application reliability and effectiveness of the large language model generated text continuous traceability model.
[0060] To further improve the training efficiency of the large language model generated text continuous traceability model and reduce the resource consumption of the training device, in a training method of the large language model generated text continuous traceability model provided in an embodiment of the present application, see Figure 3 , before step 100 in the training method of the large language model generated text continuous traceability model, the following specific content is further included: Step 050: If a large language model not applied in the historical training period is detected or received, set a label uniquely representing the type of the large language model, and obtain multiple text data generated by using the large language model to generate the data set of the current training stage; wherein, the data set of the current training stage contains multiple training samples.
[0061] That is to say, the second triggering situation of the stage feature extraction step can be triggered by detecting the current existence of a newly released large language model. The steps 100 to 300 executed after step 050 can be an incremental training process for continuously tracing the source model of the text generated by the large language model. Based on the de-correlated prototypes trained in the historical stage, the training efficiency of the continuously tracing the source model of the text generated by the large language model is improved and the resource consumption of the training device is reduced.
[0062] In order to further improve the effectiveness and reliability of feature extraction in the training process of the continuously tracing the source model of the text generated by the large language model, in a method for training a continuously tracing the source model of the text generated by the large language model provided in an embodiment of the present application, the feature extraction unit includes: a pre-trained language model with fixed parameters and a randomly upward projection layer with fixed parameters, where the pre-trained language model with fixed parameters can also be referred to as a frozen pre-trained language model; the randomly upward projection layer with fixed parameters can also be referred to as a frozen randomly upward projection layer.
[0063] Correspondingly, referring to Figure 2 or Figure 3 , step 100 in the method for training a continuously tracing the source model of the text generated by the large language model specifically includes the following content: Step 110: Input the text data corresponding to each training sample in the dataset corresponding to the current training stage in the current model training process into the pre-trained language model respectively, so that the pre-trained language model outputs the initial feature vectors corresponding to each text data respectively, and then make the randomly upward projection layer perform an upward projection process on each initial feature vector to obtain the feature vectors corresponding to each text data respectively.
[0064] Specifically, the calculation formula of the feature vectors corresponding to each text data is shown in formula (1): Wherein, represents the feature vector corresponding to the k-th text data in the t-th dataset corresponding to the t-th training stage (current stage); represents the k-th text data uniquely corresponding to the k-th training sample in the t-th dataset, and the dataset contains training samples, contains training samples, represents an upward projection layer that maps the encoded output to a higher-dimensional space (M > L), represents that the characteristic vector dimension of the text data is M; , represents that the dimension of this projection layer is L × M; Denote the average pooling of the last layer hidden states obtained from the frozen pre-trained model, , Denote that the dimension after pooling is L, is an element-wise non-linear activation function.
[0065] To further improve the effectiveness and reliability of obtaining the initial prototype and text feature correlation data during the training process of the text continuous traceability model generated by the large language model, in a method for training a text continuous traceability model generated by a large language model provided in an embodiment of the present application, refer to Figure 2 or Figure 3 , the step 100 in the method for training a text continuous traceability model generated by a large language model further includes the following content executed after step 110: Step 120: According to the labels corresponding to each of the text data, sum the feature vectors of each of the text data belonging to the same large language model to respectively obtain the initial prototypes of each of the large language models.
[0066] Specifically, the calculation formula for the initial prototype of the large language model is shown in formula (2): where, Denote the initial prototype; is an indicator function; Denote the label corresponding to the k-th text data in the t-th dataset corresponding to the t-th training stage (current stage); Denote the label value.
[0067] And, step 130: According to the labels corresponding to each of the text data, perform a Gram matrix calculation on the feature vectors of each of the text data belonging to the same large language model to respectively obtain the Gram matrices corresponding to each of the large language models as text feature correlation numbers.
[0068] It can be understood that the Gram matrix refers to the Gram matrix Specifically, the calculation formula for performing a Gram matrix calculation on the feature vectors of each of the text data belonging to the same large language model is shown in formula (3): where, Denote the Gram matrix as the text feature correlation number; Denote that the dimension of Gram is a matrix of M×M; is an outer product operation.
[0069] In order to effectively solve the problem that the similarities between the prototypes of large language models are too high and the correlations between the prototypes cannot be removed at a fine-grained level, in a method for training a continuous traceability model for generating text by a large language model provided in an embodiment of the present application, if the current training stage is not the last training stage in the current model training process, then for the next data set, return to step 100 and continue to execute the stage feature extraction step; if the current training stage is the last training stage in the current model training process, then refer to Figure 2 or Figure 3 , the specific content of step 200 in the method for training a continuous traceability model for generating text by a large language model includes the following: Step 210: According to the text feature correlation data of each of the large language models obtained in sequence for each of the historical training stages and the current training stage, perform global decorrelation processing on each of the initial prototypes obtained in sequence for each of the historical training stages and the current training stage, so as to obtain the prototypes corresponding to each of the large language models with global correlation eliminated.
[0070] It can be understood that if the current training stage is the last training stage in the current model training process, then at this time, the initial prototypes of each of the large language models obtained in sequence for each of the historical training stages and the current training stage are denoted as .
[0071] Specifically, the concatenated data formed by the prototypes corresponding to each of the large language models with global correlation eliminated is calculated as shown in formula (4): where , are respectively the prototypes corresponding to each of the large language models with global correlation eliminated; indicates that the dimension of this prototype is M× ; indicates the label set corresponding to the data set ; indicates the regularization term used to ensure the numerical stability during the inverse operation of formula (4); is the concatenated data of each initial prototype, , indicates that the dimension of this prototype is M× .
[0072] Formula (4) is based on the long-established least squares error prediction theory, or more precisely, the ridge regression algorithm. This global decorrelation process enhances the distinguishability between the prototypes of all large language models.
[0073] And, step 220: Based on the text feature correlation data of each of the large language models obtained in sequence for each of the historical training stages and the current training stage, perform local decorrelation processing on every two of the initial prototypes respectively to obtain the prototypes with eliminated local correlation corresponding to each of the large language models.
[0074] Specifically, the concatenated data formed by the prototypes with eliminated local correlation corresponding to each of the large language models has the calculation formula shown in formula (5): , where is the prototype with eliminated local correlation corresponding to the i-th initial prototype ; is the prototype with eliminated local correlation corresponding to the j-th initial prototype ; represents the concatenated data of the i-th initial prototype and the j-th initial prototype; represents the regularization term used to ensure numerical stability during the inverse operation of formula (5).
[0075] Step 230: Use the prototypes with eliminated global correlation and the prototypes with eliminated local correlation corresponding to each of the large language models respectively as the decorrelated prototypes corresponding to each of the large language models.
[0076] Based on the above embodiments of the method for training a text continuous traceability model generated by a large language model, the present application also provides an embodiment of a method for text continuous traceability generated by a large language model. Refer to Figure 4 , the method for text continuous traceability generated by the large language model specifically includes the following content: Step 300: Input the target text data into the text continuous traceability model generated by the large language model trained by the method for training a text continuous traceability model generated by the large language model, so that the feature extraction unit in the text continuous traceability model generated by the large language model extracts the target feature vector corresponding to the target text data, and enables the text continuous traceability model generated by the large language model to match the target feature vector with each of the decorrelated prototypes respectively to obtain the target prototype corresponding to the target text data, and enables the text continuous traceability model generated by the large language model to search for the label corresponding to the target prototype in the label group and output the label as the traceability prediction result of the target text data.
[0077] In step 300, the target feature vector extracted by the feature extraction unit for the target text data .
[0078] As can be seen from the above description, the method for continuously tracing the text generated by the large language model provided in the embodiments of the present application can improve the reliability and effectiveness of the tracing results.
[0079] In order to effectively solve the problem that the similarity between the prototypes of the large language model is too high and the correlation between the prototypes cannot be removed at a fine-grained level, in a method for training a continuous text tracing model of a large language model provided in the embodiments of the present application, the de-correlated prototypes include: prototypes with global correlation eliminated and prototypes with local correlation eliminated; see Figure 5 , the process of matching the target feature vector with each of the de-correlated prototypes to obtain the target prototype corresponding to the target text data (which can be simply referred to as: prototype matching process) in the method for continuously tracing the text generated by the large language model specifically includes the following contents: Step 310: Match the target feature vector with each of the prototypes with global correlation eliminated to obtain the global matching score corresponding to each of the prototypes with global correlation eliminated.
[0080] The global matching score corresponding to each of the prototypes with global correlation eliminated is calculated as shown in formula (6): where represents the target feature vector corresponding to the target text data ; for , represents the global matching score with the th large language model, ; represents the prototype with global correlation eliminated corresponding to the yth large language model.
[0081] Step 320: Sort the prototypes with global correlation eliminated in descending order of the global matching score, and use the top two prototypes with global correlation eliminated as candidate prototypes.
[0082] The present application selects the two prototypes with global correlation eliminated that result in the highest global matching scores as candidate prototypes, indicating the two most likely sources predicted by the global de-correlated prototypes.
[0083] Step 330: Match the target feature vector with the prototypes with local correlation eliminated corresponding to each of the candidate prototypes to obtain the local matching scores corresponding to the two candidate prototypes.
[0084] The calculation formula for the local matching scores corresponding to the two candidate prototypes is as shown in formula (7): Among them, it is represented that the local matching scores corresponding to the two candidate prototypes respectively include , where represents the local matching score corresponding to the i-th candidate prototype; represents the local matching score corresponding to the j-th candidate prototype; represents the prototype with the local correlation eliminated corresponding to the i-th candidate prototype; represents the prototype with the local correlation eliminated corresponding to the j-th candidate prototype.
[0085] Step 340: Determine the candidate prototype with the highest local matching score as the target prototype corresponding to the target text data.
[0086] Among them, the target prototype is denoted as .
[0087] To further illustrate the above embodiments, the present application also provides a specific application example of a training method for a continuous traceability model of text generated by a large language model and a method for continuous traceability of text generated by a large language model, which relates to the field of natural language processing and can also be called a Global&Local Prototype Decorrelation (GLPD) continuous traceability method, and specifically includes the following contents: (1) Paradigm of the continuous traceability task of text generated by a large language model Specifically, assume there are large language models, namely: , sorted in chronological order according to the release order, for example . Each large language model has a set of texts generated by it. In addition, the application example of the present application also has a set of texts written by humans. Unless otherwise specified, the present application regards the human identity (Human) as a special large language model, denoted as , and its text set is denoted as . The sequence of these large language models (as well as ) is then divided into non-overlapping label sets , where the large language models in the i-th label set are always released earlier than the large language models in the j-th label set (when ), and ( ).
[0088] The application example of this application assumes that human identity belongs to the first tag set, that is For each tag set The application example of this application constructs its instance set by merging the text collections related to the large language models in the tag set That is , represents the k-th large language model. After that, the application example of this application obtains a stream data sequence containing delimited learning and training stages , are different data sets respectively, and each training stage has a data set where is a piece of text, is the large language model that generates the text.
[0089] This task requires the traceability model to learn sequentially from During the training stage the traceability model only learns from the data set of the current training stage and cannot access the data of previous training stages. After learning, given any unseen text the traceability model must predict a tag to trace the source of the text . Figure 6 Shows an example of this sequential learning process. At Figure 6Among them, Phase 1, Phase 2, and Phase 3 respectively represent different training phases executed in sequence; "Human" represents the label corresponding to the initial large language model (i.e., human identity); "LLM1" represents the label of the first type of large language model; "LLM2" represents the label of the second type of large language model; "LLM3" represents the label of the third type of large language model; "LLM4" represents the label of the fourth type of large language model; "LLM5" represents the label of the fifth type of large language model; "Human-Written" represents human-written text data; "LLM1-Generated" represents text data generated by the large language model with the label LLM1; "LLM2-Generated" represents text data generated by the large language model with the label LLM2; "LLM3-Generated" represents text data generated by the large language model with the label LLM3; "LLM4-Generated" represents text data generated by the large language model with the label LLM4; "LLM5-Generated" represents text data generated by the large language model with the label LLM5; Label Set 1 refers to the label set corresponding to Phase 1; Label Set 2 refers to the label set corresponding to Phase 2; Label Set 3 refers to the label set corresponding to Phase 4; "Model" is the abbreviation of the continuous traceability model for text generation by large language models. The continuous traceability task of text generation by large language models well simulates the real scenario, in which new large language models will continuously emerge, and the model needs to quickly adapt to and identify these new large language models without frequently retraining historical data.
[0090] (2) Global and local prototype decorrelation The application example of this application proposes a global and local prototype decorrelation continuous traceability method. This method is a prototype-based method, and its core idea is to learn the prototypes of large language models and perform source tracking through prototype matching. Figures 7(a) and 7(b) respectively show the overall processes of the learning phase and the inference phase corresponding to the global and local prototype decorrelation continuous traceability method. Among them, C0, C1, C2, and C3 respectively represent different prototypes; X represents the target text data. In the learning and training phase, a frozen pre-trained language model and a frozen upward projection layer are used to continuously extract the prototypes of newly emerging large language models. Subsequently, global and local decorrelation mechanisms are introduced to eliminate the overall and pairwise correlations between prototypes. In the inference and training phase, a two-step prototype matching mechanism is adopted to trace the source of the given text. First, the input is matched with the globally decorrelated prototype to trace the two most likely sources, and then it is matched with the corresponding locally decorrelated prototype to make a final decision. The following application example of this application details these two training phases: (1) Construct prototypes During the learning and training phase, GLPD gradually learns and decorrelates the prototypes of continuously emerging large language models. Taking the training phase as an example, a dataset containing training samples is provided . Each sample consists of a piece of text (represented as a token sequence) and a label (referring to the large language model that generates the text). The application example of this application first uses a frozen pre-trained language model, such as LLaMA, and a frozen random upward projection layer to extract the features of all training samples from , and obtains a feature vector for each according to the aforementioned formula (1).
[0091] Then, for each large language model , the application example of this application generates the feature vector of the training samples of this large language model in a summation manner through the above formula (2) to construct its initial prototype ; this initial prototype captures the overall features of the text generated by the large language model and serves as the representative embedding of this large language model.
[0092] After prototype extraction, all large language models observed so far have obtained prototypes . However, there is often a high degree of correlation between the prototypes directly extracted by the frozen pre-trained language model, which poses a challenge to accurate prototype matching, thus compromising the accuracy of source tracking. Therefore, the application example of this application introduces global and local decorrelation mechanisms to further decorrelate these prototypes. The application example of this application eliminates the correlation between prototypes through the inverse operation of the Gram matrix, which is similar to incremental linear discriminant analysis but simpler. Specifically, for each large language model , the Gram matrix of the features of all training samples generated by it is calculated according to the above formula (3), and this Gram matrix captures the feature correlation of the text generated based on the large language model.
[0093] (2) Global prototype decorrelation The application example of this application performs global decorrelation based on the aforementioned formula (4) to eliminate the overall correlation between the prototypes related to all large language models observed so far .
[0094] (3) Local prototype decorrelation In addition to global decorrelation, the application example of the present application also performs local decorrelation based on the above formula (5) to further eliminate the pairwise correlation between the prototypes associated with any two specific large language models. For any two large language models and their prototypes , the local decorrelation is performed in a similar manner to global decorrelation, and this local decorrelation process further enhances the distinguishability of the large language models from the prototypes.
[0095] (4) Prototype matching in two training stages In the inference training stage, given the globally and locally decorrelated prototypes, the goal of the application example of the present application is to predict the label of any unseen text to trace the source large language model of the text. To achieve this goal, the application example of the present application designs a two-step prototype matching process. Specifically, given , the application example of the present application first extracts the feature vector and, based on the aforementioned formula (6), matches with all globally decorrelated prototypes of the large language models to select the two globally decorrelated prototypes that result in the highest global matching score as candidate prototypes, representing the two most likely sources predicted by the globally decorrelated prototypes.
[0096] After that, the application example of the present application matches with the locally decorrelated prototypes corresponding to these two candidate prototypes based on the above formula (7) to further distinguish them.
[0097] Then, the application example of the present application selects the candidate prototype with the higher local matching score as the final prediction (i.e., the target prototype): . This two-step matching process achieves a coarse-to-fine prototype selection, thereby improving the accuracy of prototype-based source tracing. The GLPD algorithm is described in Figure 8 , where "for...do" represents a loop structure; "end for" represents the end of the loop structure. GLPD is training-independent, and the pre-trained language model and the random upper projection layer are frozen and shared in all learning and training stages.
[0098] That is to say, the application example of this application first solves the problem of frequent retraining caused by the fixed label set in traditional traceability methods. Also, aiming at the limitations of the continual learning method in the traceability task, the proposed global and local prototype decorrelation-based generated text continuous traceability method in this application example eliminates the overall correlation between all prototypes through the global decorrelation mechanism, and the local decorrelation mechanism further separates the subtle interference between paired prototypes, thus significantly improving the discrimination ability of texts generated by different large language models. After learning ten large language models, Figure 9 The similarity between the prototypes of the large language models after local decorrelation in Figure 9 . The darker the green square color, the higher the similarity. It can be seen that after globally decorrelating the prototypes of the ten large language models, the similarity between the prototypes of the large language models is still very high. Among them, Human represents the human identity and also serves as a type of large language model. GPT-2, GPT-3, Cohere, ChatGPT, Cohere-Chat, GPT-4, MPT, MPT-Chat, and LLaMA-Chat respectively represent different types of large language models.
[0099] Therefore, through the combination of global and local decorrelation, the text generated by the large language model can be traced well. At the same time, through the incremental learning mechanism, the model can effectively retain the category knowledge of the existing large language models while quickly adapting to and identifying the text generated by newly emerging large language models without frequent retraining.
[0100] The global and local prototype decorrelation-based generated text continuous traceability method in this application example can not only solve the problem of frequent retraining caused by the fixed label set in traditional traceability methods, but also solve the problem that the similarity between the prototypes of the large language models in the prototype-based continual learning method is too high and the correlation between the prototypes of the large language models cannot be removed at a fine-grained level.
[0101] In summary, the application example of this application effectively addresses the issue of tracing the text generated by large language models that people are increasingly concerned about. Different from previous studies that rely on fixed label set classifiers, the application example of this application sets this task as a category incremental learning process in continual learning, enabling the model to quickly adapt to newly emerging large language models without frequent retraining. The application example of this application designs a method without training, continuously extracts new large language model prototypes using a frozen pre-trained language model, and adopts global and local prototype decorrelation techniques to improve the accuracy of tracing the text generated by large language models.
[0102] In addition, the application example of this application constructs a COT-Bench benchmark for evaluating the performance of language models, which includes text data generated by 19 large language models from 12 suppliers in 8 different fields. The application example of this application tags the 19 large language models in chronological order according to their respective release dates, and allows the models to learn in sequence according to the order of the tags, simulating the scenario where large language models appear in sequence over time, providing a real test environment for evaluation.
[0103] The application example of this application evaluates the proposed method on the COT-Bench benchmark and compares it with a variety of existing continual learning methods. The experimental results are shown in Table 1, where LwF, Ease, SimpleCIL, Aper, RanPAC are continual learning methods, while T5-Sentinel, LeCNN, DeTeCTive, JointLearning are non-continual learning methods. Non-continual learning methods refer to retraining the classifier using historical and current data in each training stage, which can be regarded as the upper limit of the performance of continual learning methods.
[0104] The comparison table among the above methods is used to represent the final accuracy A F and the average accuracy Ā across 8 fields. "Avg." represents the arithmetic mean of 8 fields. The best performance is marked in bold. Among them, due to the large number of columns in the comparison table, the comparison table is split into Table 1-1 and Table 1-2. That is, in the complete comparison table, the column of recipes in Table 1-2 is to the right of the column of poems in Table 1-1.
[0105] Table 1-1 Table 1-2 Among them, the continual learning methods include: 1. LwF is a regularization-based method that injects prior knowledge by distilling the output of the previous model when learning new data.
[0106] 2. Ease is an architecture-based method that learns a lightweight Adapter for each training stage to adapt to the new data distribution. By concatenating the features of the pre-trained model and each adapter, joint decision-making can be carried out in multiple subspaces.
[0107] 3. SimpleCIL is a prototype-based method that uses a frozen pre-trained model to extract large language model prototypes and traces the source through simple prototype matching.
[0108] 4. Aper is a prototype-based method that fine-tunes the adapter of a pre-trained model in the first training stage and then freezes the adapter in subsequent training stages to enhance the representation ability of extracting prototypes.
[0109] 5. RanPAC is a prototype-based method that enhances the linear separability of prototypes by applying global prototype decorrelation.
[0110] Non-continuous learning methods include: 1. T5-Sentinel fine-tunes the T5-large model and uses its next-word prediction ability for source tracking.
[0111] 2. LeCNN uses LLaMA-2-7B to extract text embeddings and trains a CNN classifier on these embeddings for multi-class classification.
[0112] 3. DeTeCtive uses the BERT-based pre-trained language model RoBERTa-base model with a multi-level contrastive learning and dense retrieval framework to identify the source of text. BERT is a bidirectional encoding representation based on Transformer, and Transformer is a neural network model based on the self-attention mechanism.
[0113] 4. JointLearning directly fine-tunes the LLaMA-3.2-1B model to trace the source of the text generated by the large language model.
[0114] Experimental results show that GLPD outperforms all baselines in terms of consistency performance on datasets in all eight domains. GLPD is superior to the state-of-the-art RanPAC continuous learning method. GLPD achieves more superior average performance compared to the state-of-the-art source-tracing methods following the non-continuous learning paradigm in 8 domains. In addition, as a training-free method, the performance of GLPD is close to that of JointLearning (the overall average gap is only 2.02% A_F and 1.17% Ā), that is, joint learning. Joint learning is used as the upper limit of continuous learning methods and adopts the same backbone network as GLPD, further proving its powerful performance and showing the great potential of GLPD in new tasks. The research results emphasize the importance and feasibility of continuous source tracing for large language models, providing a promising direction for future research.
[0115] At the software level, this application also provides a device for training a continuous source tracing model for text generation by a large language model, which is used to execute all or part of the content in the method for training a continuous source tracing model for text generation by a large language model. See Figure 10 The device for training a continuous source tracing model for text generation by a large language model specifically includes the following contents: The stage feature extraction model 10 is used to perform the stage feature extraction step: respectively input each training sample in the dataset corresponding to the current training stage in the current model training process into the feature extraction unit, so that the feature extraction unit extracts the feature vectors corresponding to the text data in each of the training samples respectively; wherein, each of the training samples further includes a label for indicating the type of the large language model that generates the text data; the release time of the large language model corresponding to the current training stage is later than the release times of the large language models corresponding to each of the historical training stages; and, according to the feature vectors of each of the text data and the labels corresponding to each of the text data, respectively obtain the initial prototypes and text feature correlation data of each of the large language models. The global and local decorrelation module 20 is used to, if the current training stage is the last training stage in the current model training process, perform global decorrelation processing and local decorrelation processing on the initial prototypes obtained in sequence for each of the historical training stages and the current training stage according to the text feature correlation data of each of the large language models obtained in sequence for each of the historical training stages and the current training stage, to obtain the decorrelated prototypes corresponding to each of the large language models, so as to generate a large language model generation text continuous traceability model for predicting the type of the large language model that generates text data and including the feature extraction unit based on the current decorrelated prototypes and a label set including the labels corresponding to each of the decorrelated prototypes.
[0116] The embodiment of the large language model generation text continuous traceability model training device provided in this application can specifically be used to execute the processing flow of the embodiment of the large language model generation text continuous traceability model training method in the above embodiment, and its functions will not be elaborated here, and reference can be made to the detailed description of the embodiment of the large language model generation text continuous traceability model training method above.
[0117] The part of the large language model generation text continuous traceability model training device for training the large language model generation text continuous traceability model can be completed in a server or a client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make a limitation in this regard. If all operations are completed in the client device, the client device may further include a processor for specifically processing the training of the large language model generation text continuous traceability model.
[0118] The above-mentioned client device may have a communication module (i.e., a communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0119] Any suitable network protocol can be used for communication between the above-mentioned server and the client device, including network protocols that have not been developed as of the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may also, for example, include RPC protocol (Remote Procedure Call Protocol) and REST protocol (Representational State Transfer) used on top of the above-mentioned protocols.
[0120] As can be seen from the above description, the large language model-generated text continuous traceability model training device provided in the embodiments of this application can not only solve the problem of frequent retraining caused by a fixed tag set in traditional traceability methods, but also solve the problem that the similarity between prototypes of large language models is too high and the correlation between prototypes cannot be removed at a fine-grained level.
[0121] From a software perspective, this application also provides a large language model-generated text continuous traceability device for executing all or part of the large language model-generated text continuous traceability method. See Figure 11 , the large language model-generated text continuous traceability device specifically includes the following: A model prediction module 30, configured to input target text data into the large language model-generated text continuous traceability model trained by the large language model-generated text continuous traceability model training method, so that the feature extraction unit in the large language model-generated text continuous traceability model extracts the target feature vector corresponding to the target text data, and enables the large language model-generated text continuous traceability model to match the target feature vector with each of the decorrelated prototypes to obtain the target prototype corresponding to the target text data, and enables the large language model-generated text continuous traceability model to search for the label corresponding to the target prototype in the tag group and output the label as the traceability prediction result of the target text data.
[0122] The embodiments of the device for continuously tracing the text generated by the large language model provided in this application can be specifically used to execute the processing procedures of the embodiments of the method for continuously tracing the text generated by the large language model in the above embodiments. Its functions will not be elaborated here, and reference can be made to the detailed description of the embodiments of the method for continuously tracing the text generated by the large language model above.
[0123] The part of the device for continuously tracing the text generated by the large language model to perform continuous tracing of the text generated by the large language model can be completed in the server or the client device. Specifically, it can be selected according to the processing capabilities of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor for specific processing of continuously tracing the text generated by the large language model.
[0124] The embodiments of this application also provide an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the method for training the model for continuously tracing the text generated by the large language model and / or the method for continuously tracing the text generated by the large language model mentioned in the above embodiments. Among them, the processor and the memory can be connected through a bus or other means. Taking the connection through the bus as an example, the receiver can be connected to the processor and the memory in a wired or wireless manner.
[0125] The processor can be a Central Processing Unit (CPU). The processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.
[0126] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for training the model for continuously tracing the text generated by the large language model and / or the method for continuously tracing the text generated by the large language model in the embodiments of this application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, to implement the method for training the model for continuously tracing the text generated by the large language model and / or the method for continuously tracing the text generated by the large language model in the above method embodiments.
[0127] The memory may include a program storage area and a data storage area. Among them, the program storage area can store the operating system and application programs required for at least one function; the data storage area can store data created by the processor and the like. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0128] The one or more modules are stored in the memory and, when executed by the processor, implement the method for training the large language model to generate text continuous traceability model and / or the method for large language model to generate text continuous traceability in the embodiments.
[0129] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.
[0130] As an implementation manner, the functions of the receiver and the transmitter in the present application may be considered to be implemented through a transceiver circuit or a dedicated chip for transceiver, and the processor may be considered to be implemented through a dedicated processing chip, a processing circuit, or a general-purpose chip.
[0131] As another implementation manner, it may be considered to use a general computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, the receiver, and the transmitter are stored in the memory, and the general processor implements the functions of the processor, the receiver, and the transmitter by executing the codes in the memory.
[0132] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing method for training the large language model to generate text continuous traceability model and / or the method for large language model to generate text continuous traceability are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium well-known in the technical field.
[0133] The embodiments of the present application further provide a computer program product, including a computer program, which, when executed by a processor, implements the steps of the foregoing large language model generated text continuous traceability model training method and / or the large language model generated text continuous traceability method.
[0134] Those of ordinary skill in the art should understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0135] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0136] In the present application, the features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0137] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and variations can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for training a continuous tracing model for text generated by a large language model, characterized in that: include: Stage feature extraction step: input each training sample in the data set corresponding to the current training stage in this model training process into the feature extraction unit, so that the feature extraction unit extracts the feature vectors corresponding to the text data in each training sample; each training sample also includes a label for indicating the type of the large language model that generates the text data; the release time of the large language model corresponding to the current training stage is later than the release time of the large language model corresponding to each historical training stage; and, according to the feature vectors of each text data and the labels corresponding to each text data, respectively obtain the initial prototype and text feature correlation data of each large language model; If the current training stage is the last training stage in this model training process, then the initial prototypes obtained in sequence in each of the historical training stages and the current training stage are globally and locally decorrelated according to the text feature correlation data of each of the large language models obtained in sequence in each of the historical training stages and the current training stage to obtain decorrelated prototypes corresponding to each of the large language models, so as to generate a large language model-generated text continuous tracing model for predicting the type of the large language model to which the generated text data belongs and including the feature extraction unit based on the current decorrelated prototypes and the label group including the labels corresponding to each of the decorrelated prototypes.
2. The method for training a large language model to generate text continuous tracing model according to claim 1, characterized in that: Before the feature extraction step, the stage also includes: Collect various text data written by humans and various text data generated by various currently released large language models; Using each manually written text data as each text data of the initial large language model, and setting a label for indicating the type of the large language model for the initial large language model and each of the published large language models; The label of the initial large language model is used as the first label, and the labels of each of the large language models that have been published are sorted after the label of the initial large language model in order of publication time from earliest to latest, so as to obtain a corresponding label set; According to the order of each label in the label set, the data sets corresponding to each label are constructed in sequence, wherein the data sets contain each text data corresponding to the label, and each text data in the data set and the corresponding label constitute different training samples.
3. The method for training a large language model to generate text continuous tracing model according to claim 1, characterized in that: Before the feature extraction step, the stage also includes: If a large language model that has not been applied in the historical training period is detected or received, a label is set to uniquely represent the type to which the large language model belongs, and multiple text data generated using the large language model are obtained to generate a data set for the current training stage; wherein the data set for the current training stage includes multiple training samples.
4. The method for training a large language model to generate text continuous tracing model according to claim 1, characterized in that: The feature extraction unit includes: a pre-trained language model with fixed parameters and a random upward projection layer with fixed parameters; Correspondingly, the respective training samples in the data set corresponding to the current training stage in the model training process are input into the feature extraction unit, so that the feature extraction unit extracts the feature vectors corresponding to the text data in each of the training samples, including: The text data corresponding to each training sample in the data set corresponding to the current training stage in this model training process are respectively input into the pre-trained language model, so that the pre-trained language model outputs the initial feature vectors corresponding to each of the text data, and then the random upward projection layer performs upward projection processing on each of the initial feature vectors to obtain the feature vectors corresponding to each of the text data.
5. The method for training a large language model to generate text continuous tracing model according to claim 1, characterized in that: The step of obtaining the initial prototype and text feature correlation data of each of the large language models according to the feature vectors of each of the text data and the labels corresponding to each of the text data includes: According to the labels corresponding to the text data, the feature vectors of the text data belonging to the same large language model are added and calculated to obtain initial prototypes of the large language models; Furthermore, according to the labels corresponding to the respective text data, the feature vectors of the respective text data belonging to the same large language model are calculated by Gram matrix, so as to obtain the Gram matrix corresponding to each of the large language models as the text feature correlation number.
6. The method for training a large language model to generate text continuous tracing model according to claim 1, characterized in that: According to the text feature correlation data of each of the large language models sequentially obtained in each of the historical training stages and the current training stage, global decorrelation processing and local decorrelation processing are performed on each of the initial prototypes sequentially obtained in each of the historical training stages and the current training stage to obtain decorrelation prototypes corresponding to each of the large language models, including: According to the text feature correlation data of each of the large language models obtained in sequence in each of the historical training stages and the current training stage, a global decorrelation process is performed on each of the initial prototypes obtained in sequence in each of the historical training stages and the current training stage to obtain prototypes corresponding to each of the large language models from which global correlation has been eliminated; And, based on the text feature correlation data of each of the large language models obtained in sequence in each of the historical training stages and the current training stage, each of the two initial prototypes is locally decorrelated to obtain prototypes corresponding to each of the large language models from which local correlation has been eliminated; The prototypes from which global correlations have been eliminated and the prototypes from which local correlations have been eliminated, which correspond to each of the large language models, are respectively used as decorrelation prototypes which correspond to each of the large language models.
7. A method for continuous tracing of text generated by a large language model, characterized in that: include: The target text data is input into the large language model generated text continuous tracing model trained by the large language model generated text continuous tracing model training method according to any one of claims 1 to 6, so that the feature extraction unit in the large language model generated text continuous tracing model extracts the target feature vector corresponding to the target text data, so that the large language model generated text continuous tracing model matches the target feature vector with each of the decorrelated prototypes respectively to obtain the target prototype corresponding to the target text data, and so that the large language model generated text continuous tracing model searches for the label corresponding to the target prototype in the label group and outputs the label as the tracing prediction result of the target text data.
8. The method for continuously tracing the source of text generated by a large language model according to claim 7, characterized in that: The decorrelation prototype includes: a prototype from which global correlation has been eliminated and a prototype from which local correlation has been eliminated; The step of matching the target feature vector with each of the decorrelation prototypes to obtain a target prototype corresponding to the target text data includes: Matching the target feature vector with each of the prototypes whose global correlation has been eliminated, respectively, to obtain a global matching score corresponding to each of the prototypes whose global correlation has been eliminated; Sorting the prototypes whose global correlations have been eliminated in descending order of the global matching scores, and taking the first two prototypes whose global correlations have been eliminated as candidate prototypes; Matching the target feature vector with the prototypes whose local correlations have been eliminated corresponding to each of the candidate prototypes, respectively, to obtain local matching scores corresponding to each of the two candidate prototypes; The candidate prototype with the highest local matching score is determined as the target prototype corresponding to the target text data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the method for training a model for continuously tracing text generated by a large language model as described in any one of claims 1 to 6, and / or implements the method for continuously tracing text generated by a large language model as described in claim 7 or 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for training a model for continuously tracing a text generated by a large language model as described in any one of claims 1 to 6, and / or implements the method for continuously tracing a text generated by a large language model as described in claim 7 or 8.
Citation Information
Patent Citations
Fine adjustment method, system and equipment based on large language model and medium
CN117290480A
Large model answer tracing method based on linear regression
CN118656467A
Characterization-based large language model source tracing method
CN119357656A
Model training method and device, and computer-readable storage medium
JP2023181109A