Method and device for detecting text generated by fine-tuning large language model
By obtaining the probability list of the text to be detected on multiple base language models and performing feature extraction, combining family perception comparison learning and mixed expert classification network, the problem of fine-tuning large language models to generate text detection is solved, and effective recognition and performance improvement of generated text is achieved.
Patent Information
- Application Number
- CN202510473807.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to effectively detect text content generated by fine-tuned large language models, especially after fine-tuning of private data sets, malicious users may use these models to generate harmful texts and disseminate them.
By obtaining the probability list of the text to be detected on multiple pedestal language models, feature extraction is performed, family relationship representations of pedestal probability characteristics are enhanced using family perception comparison learning, and a hybrid expert classification network is used to generate source predictions, and the target family prediction results are generated to determine whether the text is generated text.
It realizes effective detection of text generated by fine-tuned large language models, significantly improves detection performance, and can identify the source of generated text.
Smart Images

Figure CN120353933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text detection, and in particular, to a method, apparatus, storage medium, and program product for detecting text generated by a fine-tuned large language model based on family awareness. Background Art
[0002] Since the popularization of large language model technology, although the existing methods for detecting text generated by large language models have made remarkable progress, they usually use representation learning methods to learn the common text features (such as writing styles, etc.) between different large language models, or design distinguishable metrics between human texts and the generated texts of large language models according to the internal signals of large language models (such as token probabilities, etc.). In application scenarios, the above existing technologies are usually tested on publicly available large language model-generated data, assuming that users usually use public model servers represented by ChatGPT and DeepSeek for text generation.
[0003] However, due to the active development of the open-source large model community, the above situation is changing and the assumption is difficult to hold. Thanks to the development of open-source model platforms such as HuggingFace and parameter-efficient large model training technologies represented by Low-Rank Adaptation (LoRA), it has become more convenient to fine-tune large language models using customized private datasets, and the number of such models has increased rapidly. For example, there are already more than 60,000 derivative models based on LLaMA fine-tuning on HuggingFace. After fine-tuning with a private dataset, the original features of the base language model may change, and the previous methods for detecting generated texts may become ineffective. The above problems have triggered new risks, that is, malicious users can use the fine-tuned model to generate harmful texts and spread them without being detected by the generated text detector. Therefore, it is necessary to effectively detect the text content generated by fine-tuned large language models. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention proposes a method, apparatus, storage medium, and program product for detecting text generated by a fine-tuned large language model, so as to effectively detect the text generated by the fine-tuned large language model.
[0005] On the one hand, the present invention provides a method for detecting text generated by a fine-tuned large language model, including:
[0006] Obtaining a probability list of the text to be detected on multiple base language models;
[0007] Performing feature extraction on the probability list to obtain base probability features;
[0008] Enhance the family relationship representation in the base probability features based on family-aware contrast learning to generate family prediction weights;
[0009] According to the family prediction weights, perform generation source prediction on the text to be detected through a mixture-of-experts classification network to generate a target family prediction result, and the target family prediction result is used to determine whether the text to be detected is a generated text.
[0010] In an embodiment of the present invention, input the text to be detected into multiple base language models for inference, where each base language model outputs the probability value of each token in the text to be detected;
[0011] Arrange the probability values output by each base language model in the order of the text sequence to form multiple probability lists.
[0012] In an embodiment of the present invention, perform convolutional neural network processing on the probability lists in sequence to extract the local fluctuation features of the probability lists;
[0013] Perform Transformer encoding processing on the local fluctuation features to extract the global change features and obtain the base probability features.
[0014] In an embodiment of the present invention, the family-aware contrast learning specifically includes:
[0015] Select the generated text belonging to the same family as the text to be detected as the perturbation sample, and the generated text of different families as the negative example sample;
[0016] Calculate the similarity between the current sample and the perturbation sample in the feature space through a contrast loss function to enhance the family relationship representation in the base probability features;
[0017] Input the representation vector into a multi-layer perceptron to predict the family category to which the text to be detected belongs and generate family prediction weights.
[0018] In an embodiment of the present invention, the contrast loss function is expressed as:
[0019]
[0020] Among them, represents the contrast loss function, b represents the sample set of the current batch, represents the representation vector corresponding to the i-th current sample, is the representation vector corresponding to the perturbation sample, represents the representation vectors corresponding to the other samples in the batch except the current sample, δ represents the vector dot product function, t is the temperature coefficient, represents the indicator function that takes the value of 1 if and only if j≠i.
[0021] In one embodiment of the present invention, the characterization vector is input into a multi-layer perceptron to classify the family of the corresponding generation model for the text to be detected, output the family prediction weight, and calculate the first cross-entropy loss for the text to be detected and the second cross-entropy loss for the perturbation sample:
[0022]
[0023] where, is the first cross-entropy loss of the text to be detected, represents the family prediction weight, represents the probability of being predicted as coming from the base language model θ i , y F ∈{θ1, θ2, …, θ M} represents the true family label of the text to be detected, and R F is the characterization vector corresponding to the text to be detected; is the second cross-entropy loss of the perturbation sample, is the characterization vector corresponding to the perturbation sample, and MLP F represents the multi-layer perceptron.
[0024] In one embodiment of the present invention, the mixture-of-experts classification network specifically includes multiple expert detectors, and each expert detector corresponds to a family category of the base language model;
[0025] Using the family prediction weight as a gating signal to control the weighted weights of the output results of each expert detector, a target family prediction result is obtained to determine whether the text to be detected is a generated text.
[0026] In one embodiment of the present invention, the family prediction weight is used as a gating signal to control the weighted weights of the output results of each expert detector, and a target family prediction result is obtained and the third cross-entropy loss is calculated:
[0027]
[0028] where, represents the probability of being predicted as coming from the base language model θ i ; represents the target family prediction result, that is, the probability that the text to be detected is predicted to belong to the generation of the fine-tuned large language model; y B ∈{0, 1} represents the binary classification true label of the text to be detected, 0 represents human text, and 1 represents generated text; is the second cross-entropy loss, R F is the family relationship representation, and MLP B represents the multi-layer perceptron.
[0029] In an embodiment of the present invention, the contrast loss function, the first cross-entropy loss, the second cross-entropy loss, and the third cross-entropy loss are combined to construct an objective loss for optimizing the entire network of the fine-tuned large language model. The objective loss is expressed as:
[0030]
[0031] wherein, is the objective loss, and λ1, λ2, and λ3 are hyperparameters used to control the importance of each part of the respective loss functions.
[0032] On the other hand, the present invention also provides a device for detecting texts generated by fine-tuning a large language model, including:
[0033] An acquisition module for acquiring a probability list of the text to be detected on multiple base language models;
[0034] A feature extraction module for extracting features from the probability list to obtain base probability features;
[0035] A family-aware learning module for enhancing the family relationship representation in the base probability features based on family-aware contrast learning to generate family prediction weights;
[0036] A mixture-of-experts detection module for predicting the generation source of the text to be detected through a mixture-of-experts classification network according to the family prediction weights, generating a target family prediction result, and using the target family prediction result to determine whether the text to be detected is a generated text.
[0037] On yet another aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0038] On yet another aspect, the present invention also provides a computer program product including a computer program, characterized in that when the computer program is executed by a processor, the steps of the above method are implemented.
[0039] As can be seen from the above solutions, the advantages of the present invention are:
[0040] The detection method for text generated by a fine-tuned large language model provided by the present invention is inspired by the "family characteristics" of the fine-tuned large language model and the base language model in the feature space. By obtaining the probability list of the text to be detected on multiple base language models and extracting features from the probability list to obtain the base probability features, and then based on family-aware contrast learning, enhancing the family relationship representation in the base probability features to generate family prediction weights; finally, using a mixture of experts classification network, according to the family prediction weights, to predict the generation source of the text to be detected through the mixture of experts classification network, generating the target family prediction result to determine whether the text to be detected is generated text. This method realizes the effective detection of text generated by the fine-tuned large language model and achieves a significant performance improvement compared with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 FIG. 6 shows a schematic overall flow chart of a detection method for text generated by a fine-tuned large language model provided by an embodiment of the present invention;
[0042] Figure 2 FIG. 10 shows a schematic principle diagram of the detection method for text generated by a fine-tuned large language model;
[0043] Figure 3 FIG. 14 is a schematic structural diagram of a device for detecting text generated by a fine-tuned large language model of the present invention.
[0044] Reference Signs:
[0045] 300 - Device for detecting text generated by a fine-tuned large language model;
[0046] 310 - Acquisition Module;
[0047] 320 - Feature Extraction Module;
[0048] 330 - Family-Aware Learning Module;
[0049] 340 - Mixture of Experts Detection Module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0051] Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in a process, method, article, or apparatus that comprises the element.
[0052] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0053] Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in a process, method, article, or apparatus that comprises the element.
[0054] The present invention fine-tunes an open-source large language model using data of different themes and analyzes the differences between the base language model and its fine-tuned large language model. The research finds that compared with non-related models, the probability distribution of the generated text on the fine-tuned large language model has a higher similarity to the probability distribution of its base language model, showing significant "family commonality". On this basis, the present invention provides a detection method for the generated text of a fine-tuned large language model, which uses the probability features of the base language model and combines the contrastive learning paradigm to perform family-aware learning to effectively detect the generated text of the fine-tuned large language model.
[0055] Specifically refer to Figure 1 、 Figure 2 as shown in Figure 1 FIG. shows a schematic overall flow diagram of a detection method for the generated text of a fine-tuned large language model provided by an embodiment of the present invention, Figure 2 FIG. shows a schematic diagram of the principle of the detection method for the generated text of a fine-tuned large language model.
[0056] A detection method for the generated text of a fine-tuned large language model specifically includes the following steps:
[0057] Step S1, obtain the probability list of the text to be detected on multiple base language models.
[0058] For a given text to be detected x = (x1, x2,..., x n ), it is necessary to determine whether the text to be detected x is generated by a fine-tuned large language model. Assume that a total of M base language models θ1, θ2,..., θ M(such as Llama-3, Qwen-2.5, Gemma, etc.). In an embodiment of the present invention, based on the "family characteristics" shared between the base language model and the fine-tuned large language model generated after its fine-tuning, the text content generated by the language model obtained by fine-tuning with private corpus is effectively detected.
[0059] Specifically, the text to be detected x=(x1, x2, …, x n ) is input into all the base language models θ1, θ2, …, θ M for inference. Each base language model outputs the probability value of each token in the text to be detected. The probability values output by each base language model are arranged in the order of the text sequence to form multiple probability lists p∈R M×N , where where N is the length of the text sequence and M is the number of base language models, represents the probability that the j-th word is x n ) when the first j-1 words x <j in the sequence of the text to be detected x=(x1, x2, …, x j are given and the model θ j outputs.
[0060] Step S2: Extract features from the probability lists to obtain base probability features.
[0061] In an embodiment, the probability lists are further processed by a convolutional neural network in sequence. The convolutional neural network is connected to the output ends of multiple base language models to extract the local fluctuation features of the probability lists. Then, the local fluctuation features are processed by a Transformer encoding process. The Transformer encoder network is connected to the output end of the convolutional neural network to extract the global change features, and base probability features with a dimension of M×d are obtained. In this embodiment, the number of network layers of the convolutional neural network and the Transformer encoder network is not limited and should be determined according to the actual situation.
[0062] In this embodiment, the text sequence to be tested is input into multiple base language models for inference, the probability lists of the text to be tested on multiple base models are obtained, and the local fluctuation features and global change features of the probability lists are effectively modeled by using a convolutional neural network and a Transformer network in sequence, enhancing the expression ability of the white-box features.
[0063] Step S3: Based on family-aware contrastive learning, enhance the family relationship representation in the base probability features to generate family prediction weights.
[0064] To better model the family relationship of generated texts in the feature space and enhance the detection generalization ability of the generated text detector on the same family model, family-aware contrast learning is adopted. Samples of the same family are regarded as perturbed samples. Based on the original probability features, the family relationship between generated texts is effectively modeled in the feature space through contrast learning, enhancing the generalization ability of the generated text detector on the same family model.
[0065] Specifically, referring to Figure 2 as shown in, generate texts belonging to the same family as the text to be detected and generate texts of different families During the training process, within each batch (mini-batch), the generated texts belonging to the same family as the current sample of the text to be detected are used as perturbed samples, and the generated texts of different families are used as negative example samples. The similarity between the current sample and the perturbed samples in the feature space is calculated through a contrast loss function to enhance the family relationship representation in the base probability features, obtaining a set of representation vectors, including the representation vector of the text to be detected, the representation vectors of the perturbed samples, and the representation vectors corresponding to other samples. In one embodiment, the contrast loss function is expressed as:
[0066]
[0067] Where represents the contrast loss function, b represents the set of samples in the current batch, represents the representation vector of the i-th sample (i.e., the current sample) of the text to be detected, is the representation vector of the perturbed sample, represents the representation vectors corresponding to other samples in the batch except the current sample, δ represents the vector dot product function, t is the temperature coefficient, represents the indicator function that takes the value of 1 if and only if j≠i.
[0068] Then, after learning the representation vectors, add the family recognition task as an auxiliary task during the training process. Input the obtained representation vectors into a family predictor constructed by a multi-layer perceptron (MLP F ) to classify the family of the generation model corresponding to the text to be detected, predict the family category to which the text to be detected belongs, and generate family prediction weights and calculate the first cross-entropy loss regarding the text to be detected the second cross-entropy loss regarding the perturbed samples
[0069] Where the representation vector R of the text to be detected F is input into a multi-layer perceptron (MLP F ), and the family prediction weights obtained are expressed as:
[0070] The first cross-entropy loss for the text to be detected is expressed as: and
[0071] The representation vector of the perturbed sample is input into a multi-layer perceptron (MLP F ), and the weight value corresponding to the family prediction weight is obtained and expressed as:
[0072]
[0073] The second cross-entropy loss for the perturbed sample is expressed as:
[0074]
[0075] where is the first cross-entropy loss of the text to be detected, represents the family prediction weight, represents the probability of being predicted as coming from the base language model θ i , y F ∈{θ1,θ2,…,θ M} represents the true family label of the text x to be detected, and R F is the representation vector corresponding to the text to be detected; is the second cross-entropy loss of the perturbed sample, is the representation vector corresponding to the perturbed sample, and MLP F represents the multi-layer perceptron.
[0076] Step S4: According to the family prediction weight, use a mixture-of-experts classification network to perform source prediction on the text to be detected, and generate a target family prediction result.
[0077] In one embodiment, a mixture-of-experts network is designed by fusing M expert detectors to perform binary classification for text detection, enabling each expert detector to focus on text detection of a specific family, and each expert detector corresponds to a family category of the base language model. Using the family prediction weight as a gating signal to control the weighted weights of the output results of each expert detector, a target family prediction result is obtained, and based on the target family prediction result, it is determined whether the text to be detected is generated text, thereby improving the performance of generated text detection.
[0078] Further referring to Figure 2 as shown in using the family prediction weight as a gating signal, each expert detector is designed by a multi-layer perceptron, and the weighted weights of the output results of each expert detector are controlled by the gating signal to obtain a target family prediction result And calculate the third cross-entropy loss
[0079] Among them, the prediction result of the target family is expressed as:
[0080]
[0081] The third cross-entropy loss is expressed as:
[0082]
[0083] Among them, represents the probability of being predicted as coming from the base language model θ i ; represents the prediction result of the target family, that is, the probability that the text x to be detected is predicted to belong to the generation of the fine-tuned large language model (that is, the text x to be detected is a generated text); y B ∈{0, 1} represents the binary classification true label of the text x to be detected, 0 represents human text, and 1 represents generated text; is the second cross-entropy loss, R F is the representation vector of the text to be detected, and MLP B represents a multi-layer perceptron.
[0084] In addition, in one embodiment, the contrast loss function is further fused with the first cross-entropy loss and the second cross-entropy loss and the third cross-entropy loss to construct a target loss to optimize the network of the entire fine-tuned large language model. The target loss is expressed as:
[0085]
[0086] Among them, is the target loss, and λ1, λ2, and λ3 are hyperparameters used to control the importance of each part of each loss function, and can be set according to specific situations and experience.
[0087] In summary, the detection method for texts generated by the fine-tuned large language model provided by the present invention is inspired by the "family characteristics" of the fine-tuned large language model and the base language model in the feature space. By obtaining the probability list of the text to be detected on multiple base language models and extracting features from the probability list to obtain the base probability features, and then based on family-aware contrast learning, enhancing the family relationship representation in the base probability features to generate family prediction weights; finally, using a mixture-of-experts classification network, according to the family prediction weights, to predict the generation source of the text to be detected through the mixture-of-experts classification network, generating a target family prediction result to determine whether the text to be detected is a generated text. This method realizes effective detection of texts generated by the fine-tuned large language model and achieves a significant performance improvement compared with the prior art.
[0088] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0089] The following is a corresponding apparatus embodiment for the above method embodiment. As Figure 3 shown, Figure 3 shows a schematic structural diagram of a detection apparatus for texts generated by a fine-tuned large language model. The implementation manner of this apparatus can be implemented in cooperation with the above method implementation manner. The relevant technical details mentioned in the above method implementation manner are still valid in the implementation manner of this apparatus. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in the implementation manner of this apparatus can also be applied to the above method implementation manner.
[0090] A detection apparatus 300 for texts generated by a fine-tuned large language model includes:
[0091] An acquisition module 310 for acquiring the probability list of the text to be detected on multiple base language models.
[0092] A feature extraction module 320 for extracting features from the probability list to obtain base probability features.
[0093] A family-aware learning module 330 for enhancing the family relationship representation in the base probability features based on family-aware contrast learning to generate family prediction weights.
[0094] A mixture-of-experts detection module 340 for predicting the generation source of the text to be detected through a mixture-of-experts classification network according to the family prediction weights, generating a target family prediction result, and the target family prediction result is used to determine whether the text to be detected is a generated text.
[0095] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules is only a logical functional division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present invention, each functional module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a processing unit.
[0096] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in a unit.
[0097] In addition, the above method embodiments can be implemented in whole or in part by software. When implemented using software, the above method embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the above-provided detection method for fine-tuning the large language model to generate text. When loading or executing the computer instructions or computer programs on the computer, the processes or functions described in the method embodiments of the present invention are generated in whole or in part.
[0098] In addition, if the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and is used to store a computer program for executing the detection method for fine-tuning the large language model to generate text, so that a computer device (which can be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0099] It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0100] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated examples herein.
Claims
1. A detection method for fine-tuning the text generated by a large language model, characterized in that, Including: Obtain a probability list of the text to be detected on multiple base language models; Extract features from the probability list to obtain base probability features; Based on family-aware contrast learning, enhance the family relationship representation in the base probability features to generate family prediction weights; According to the family prediction weights, use a mixture-of-experts classification network to perform a generation source prediction on the text to be detected, and generate a target family prediction result, where the target family prediction result is used to determine whether the text to be detected is a generated text.
2. The method according to claim 1, wherein: Input the text to be detected into multiple base language models for inference, where each base language model outputs the probability value of each token in the text to be detected; Arrange the probability values output by each base language model in the order of the text sequence to form multiple probability lists.
3. The method according to claim 1, wherein: Successively perform convolutional neural network processing on the probability list to extract the local fluctuation features of the probability list; Perform Transformer encoding processing on the local fluctuation features to extract global change features, and obtain base probability features.
4. The method according to claim 1, wherein: The family-aware contrast learning specifically includes: Select the generated texts belonging to the same family as the text to be detected as perturbation samples, and the generated texts of different families as negative example samples; Calculate the similarity between the current sample and the perturbation sample in the feature space through a contrast loss function, and enhance the family relationship representation in the base probability features; Input the representation vector into a multi-layer perceptron to predict the family category to which the text to be detected belongs, and generate family prediction weights.
5. The method according to claim 4, wherein: The contrast loss function is expressed as: Among them, represents the contrastive loss function, b represents the sample set of the current batch, represents the representation vector corresponding to the i-th current sample, is the representation vector corresponding to the perturbed sample, represents the representation vectors corresponding to the other samples in the batch except the current sample, δ represents the vector dot product function, and t is the temperature coefficient, represents the indicator function that takes the value of 1 if and only if j≠i.
6. The method according to claim 5, wherein: Input the representation vector into a multi-layer perceptron, classify the family of the generation model corresponding to the text to be detected, output family prediction weights, and calculate the first cross-entropy loss regarding the text to be detected and the second cross-entropy loss regarding the perturbation sample: Among them, is the first cross-entropy loss of the text to be detected, represents the family prediction weight, represents the probability of being predicted as coming from the base language model θ i y F ∈{θ1,θ2,…,θ M} represents the true family label of the text to be detected, and R F is the representation vector corresponding to the text to be detected; is the second cross-entropy loss of the perturbed sample, is the representation vector corresponding to the perturbed sample, and MLP F represents a multi-layer perceptron.
7. The method according to claim 6, wherein: The mixture-of-experts classification network specifically includes multiple expert detectors, and each expert detector corresponds to a family category of a base language model; Use the family prediction weights as gating signals to control the weighted weights of the output results of each expert detector to obtain the target family prediction result.
8. The method according to claim 7, wherein Use the family prediction weights as gating signals to control the weighted weights of the output results of each expert detector to obtain the target family prediction result and calculate the third cross-entropy loss: Among them, represents the probability predicted to be from the base language model θ i ; represents the target family prediction result, that is, the probability that the text to be detected is predicted to belong to the fine-tuned large language model; y B ∈{0, 1} represents the binary classification true label of the text to be detected, 0 represents human text, and 1 represents generated text; is the second cross-entropy loss, R F is the family relationship representation, and MLP B represents a multi-layer perceptron.
9. The method according to claim 4, wherein: Fuse the contrast loss function, the first cross-entropy loss, the second cross-entropy loss, and the third cross-entropy loss to construct a target loss to optimize the network of the entire fine-tuned large language model, and the target loss is expressed as: wherein, is the target loss, and λ1, λ2, and λ3 are hyperparameters used to control the importance of each part of each loss function.
10. A text detection device for fine-tuning large language model generated text, characterized in that, Including: An acquisition module for obtaining a probability list of the text to be detected on multiple base language models; A feature extraction module for extracting features from the probability list to obtain base probability features; A family perception learning module for enhancing the family relationship representation in the base probability features based on family perception contrast learning to generate family prediction weights; A hybrid expert detection module for predicting the generation source of the text to be detected through a hybrid expert classification network according to the family prediction weights to generate a target family prediction result, and the target family prediction result is used to determine whether the text to be detected is a generated text.
11. A computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1-9 are implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1-9 are implemented.