A sensitive text representation method based on contrastive learning
By constructing continuous templates and multi-task contrastive learning models, the problems of data augmentation destroying semantics and visual framework being inapplicable are solved, and the quality and efficiency of text representation are improved.
Patent Information
- Application Number
- CN202310517482.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-05-09
AI Technical Summary
Existing data augmentation methods destroy the semantics of sentences in text representation, and the computer vision contrast framework is not applicable to text data, resulting in low efficiency of contrastive learning training and affecting the quality of text representation.
A sensitive text representation method based on contrastive learning is adopted. By obtaining positive and hard negative sample pairs of sensitive text, a continuous template is constructed, and a multi-task contrastive learning model is constructed. The pre-trained model T5 is used for training, and multi-task generation tasks and optimization objectives are set. Weights are shared to improve model performance.
Effectively understand sensitive text representation tasks, avoid the influence of templates and random seeds, mine more semantic information, and improve the quality of text representation.
Smart Images

Figure CN116611414B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text representation, and in particular relates to a sensitive text representation method based on contrastive learning. Background Art
[0002] Text representation plays a key role in downstream tasks of natural language processing (NLP), such as text generation, translation, and question answering. In recent years, due to the successful application of contrastive learning in representation learning, many researchers have used it to obtain meaningful text representations. Generally speaking, text representation based on contrastive learning requires data augmentation to generate positive sample pairs for contrastive training. For example, CLEAR performs discrete data augmentation (such as word deletion, reordering, and synonym replacement) based on the computer vision contrastive framework to generate positive sample pairs for contrastive learning. SimCSE uses BERT's original Dropout as data augmentation to generate positive sample pairs. For contrastive training, SimCSE uses SimCLR's contrastive framework and improves the previous best results on the task of text representation. However, existing data augmentation methods have the following three problems: (1) Inappropriate data augmentation methods will destroy the original semantics of the sentence. In contrastive training, the semantic gap between positive sample pairs should be narrowed. Therefore, inappropriate data augmentation methods may change the semantics of the sentence, resulting in difficulties in contrastive learning and poor learning effects. (2) Most existing methods directly use the contrastive framework of computer vision for contrastive learning, which is not suitable for text data. Therefore, compared with image data, text data is discrete and sparse, which may hinder contrastive training. (3) Most methods adopt discrete prompt templates to help the model understand the text representation task, which is easily affected by different templates and random seeds. Summary of the Invention
[0003] In response to the above-mentioned deficiencies in the prior art, the present invention provides a sensitive text representation method based on contrastive learning, which solves the problem that the original semantics are distorted due to existing data enhancement and the contrastive learning training is inefficient due to the use of computer vision contrast framework in existing methods, which in turn affects the low quality of text representation.
[0004] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0005] This solution provides a sensitive text representation method based on contrastive learning, which includes the following steps:
[0006] S1. Obtain positive and hard negative sample pairs of sensitive text and construct continuous templates;
[0007] S2. Build a multi-task contrastive learning model based on positive samples of sensitive text, hard negative sample pairs, and continuous templates;
[0008] S3, training the multi-task contrastive learning model;
[0009] S4. Use the trained multi-task contrastive learning model to obtain the representation results of sensitive text.
[0010] The present invention has the following beneficial effects: It uses a continuous prompt mechanism to not only facilitate understanding of sensitive text representation tasks but also mitigate the performance impact of different templates and random seeds. Furthermore, it utilizes multi-task and contrastive learning for joint optimization, which exploits more inter-word and inter-sentence semantic information in samples for contrastive learning, resulting in improved performance. This invention addresses the issues of existing data augmentation causing distortion of original semantics and existing methods using computer vision contrast frameworks, resulting in inefficient contrastive learning training and, consequently, low text representation quality.
[0011] Furthermore, the step S1 includes the following steps:
[0012] S101, obtain positive samples and hard negative sample pairs of sensitive text;
[0013] S102. Based on the obtained positive sample and hard negative sample pairs, a prompt mechanism is used to construct a continuous template:
[0014] Template=[Extra_id_1,Extra_id_2,Extra_id_3]+[X]
[0015] Among them, Template represents a continuous template, Extra_id_1, Extra_id_2, and Extra_id_3 all represent unused word tags, and [X] represents the sensitive text to be represented.
[0016] The beneficial effect of the above further solution is that: through the above design, the present invention prevents the instability of obtaining positive samples and hard negative samples through data enhancement, helps understand sensitive text representation tasks and is not affected by random seeds and templates.
[0017] Furthermore, step S2 includes the following steps:
[0018] S201, inputting positive samples, negative samples and hard negative samples into the pre-trained model T5 to obtain an output vector with a fixed length for contrastive learning;
[0019] S202: Input the continuous template into the pre-trained model T5 to obtain the prediction output for multi-task learning, completing the construction of the multi-task contrastive learning model.
[0020] The beneficial effect of the above further solution is that the present invention utilizes the input of positive samples and negative samples into the same model to obtain the representation of each sample in a manner of shared weights, which can effectively reduce the number of model parameters and speed up the training speed.
[0021] Furthermore, step S3 includes the following steps:
[0022] S301. In the comparative document learning, the sentence is input into the pre-trained model T5, and the first word in its encoder is taken and marked as the sensitive text representation;
[0023] S302: Using the sensitive text representation, train a sensitive representation model based on contrastive learning, wherein the loss function expression of the sensitive representation model is:
[0024]
[0025] in, represents the loss function of the sensitive representation model, log represents the logarithm with base e, j represents the iteration subscript, M represents the sum of positive samples and negative samples, S(·) represents the cosine similarity, exp(·) represents the exponential function with base e, and f i 、f i + and f i - They represent the sensitive text representations of the original sentence, positive sample, and hard negative sample obtained after the pre-training model T5, and τ represents the temperature coefficient;
[0026] S303. Based on the training results, four generation tasks are set in multi-task learning.
[0027] S304. In multi-task learning, the following formula is used as the optimization objective of the generated task:
[0028]
[0029] in, represents the optimization goal of the generation task, t represents the iteration subscript, l represents the sentence length, q represents the input word, P represents the probability of generating the answer based on the input word, and a t represents the word generated by the first t steps of iteration, and a represents the generated word;
[0030] S305 : Based on the optimization goal of the generated task, in contrastive learning and multi-task learning, the pre-trained model T5 is optimized using weight sharing to complete the training of the multi-task contrastive learning model.
[0031] The beneficial effect of the above further scheme is that by generating tasks through samples, the multi-task contrastive learning model can utilize more sentence-level and word-level information to improve the performance of sensitive text representation tasks.
[0032] Furthermore, the four generation tasks are:
[0033] Generate Entailment by assuming Premise;
[0034] Generate a contradiction by assuming the premise;
[0035] Generate a hypothesis through entailment;
[0036] Generate a hypothesis through contradiction.
[0037] The beneficial effect of the above further solution is that the present invention helps the multi-task contrastive learning model to obtain semantic information of positive sample pairs and negative sample pairs through the above generation tasks to facilitate contrastive training.
[0038] Furthermore, the objective function of the multi-task contrastive learning model is as follows:
[0039]
[0040] in, represents the joint optimization objective function, and λ represents the weight coefficient.
[0041] The beneficial effect of the above further scheme is: while obtaining sensitive text representation through contrastive learning, auxiliary tasks are used to help the multi-task contrastive learning model use semantic information of more samples to improve the quality of sensitive text representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Flow chart of the method of the present invention.
[0043] Figure 2 Model diagram of the sensitive text representation method based on multi-task contrastive learning. DETAILED DESCRIPTION
[0044] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0045] Example
[0046] like Figure 1 As shown, the present invention provides a sensitive text representation method based on contrastive learning, and its implementation method is as follows:
[0047] S1. Obtain positive and hard negative sample pairs of sensitive text and construct continuous templates. The implementation method is as follows:
[0048] S101, obtain positive samples and hard negative sample pairs of sensitive text;
[0049] In this embodiment, the Premise hypothesis and Entailment implication in the NLI natural language inference dataset are used as positive samples, and the Premise and Contradiction contradictions are used as hard negative samples.
[0050] S102. Based on the obtained positive sample and hard negative sample pairs, a prompt mechanism is used to construct a continuous template:
[0051] Template=[Extra_id_1,Extra_id_2,Extra_id_3]+[X]
[0052] Among them, Template represents a continuous template, Extra_id_1, Extra_id_2, and Extra_id_3 all represent unused word tags, and [X] represents the sensitive text to be represented.
[0053] S2. Based on positive samples of sensitive text, hard negative sample pairs, and continuous templates, a multi-task contrastive learning model is constructed. The implementation method is as follows:
[0054] S201, inputting positive samples, negative samples, and hard negative samples into the pre-trained model T5 to obtain an output vector of fixed length for contrastive learning, and obtaining a loss function required for training the multi-task contrastive learning model based on the output vector;
[0055] S202: Input the continuous template into the pre-trained model T5 to obtain the prediction output for multi-task learning, completing the construction of the multi-task contrastive learning model.
[0056] S3. Training the multi-task contrastive learning model is implemented as follows:
[0057] S301. In the comparative document learning, the sentence is input into the pre-trained model T5, and the first word in its encoder is marked as the sensitive text representation;
[0058] S302: Using the sensitive text representation, train a sensitive representation model based on contrastive learning, wherein the loss function expression of the sensitive representation model is:
[0059]
[0060] in, represents the loss function of the sensitive representation model, log represents the logarithm with base e, j represents the iteration subscript, M represents the sum of positive samples and negative samples, S(·) represents the cosine similarity, exp(·) represents the exponential function with base e, and f i 、f i + and f i - They represent the sensitive text representations of the original sentence, positive sample, and hard negative sample obtained after the pre-training model T5, and τ represents the temperature coefficient;
[0061] S303. Based on the training results, four generation tasks are set in multi-task learning:
[0062] Generate Entailment by assuming Premise;
[0063] Generate a contradiction by assuming the premise;
[0064] Generate a hypothesis through entailment;
[0065] Generate a hypothesis through contradiction;
[0066] S304. In multi-task learning, the following formula is used as the optimization objective of the generated task:
[0067]
[0068] in, represents the optimization goal of the generation task, t represents the iteration subscript, l represents the sentence length, q represents the input word, P represents the probability of generating the answer based on the input word, and a t represents the word generated by the first t steps of iteration, and a represents the generated word;
[0069] S305. Based on the optimization objective of the generated task, in contrastive learning and multi-task learning, weight sharing is used to optimize the pre-trained model T5 to complete the training of the multi-task contrastive learning model;
[0070] The objective function of the multi-task contrastive learning model is as follows:
[0071]
[0072] in, represents the joint optimization objective function, and λ represents the weight coefficient;
[0073] S4. Use the trained multi-task contrastive learning model to obtain the sensitive text representation. After training is complete, the first token decoded by the pre-trained model T5 is used as the sensitive text representation.
[0074] In this embodiment, Figure 2 As shown, Figure 2 The left figure is the framework of contrastive learning, where x1…x N is the input of the encoder (T5 Encoder), d_Input1 is the input of the decoder (T5Decoder), <eid0>The continuous prompt constructed is specifically composed of 1 unused word token. The output of the decoder is output1, f i 、f i + and f i - They represent the sensitive text representations of the original sentence, positive sample, and hard negative sample obtained after the pre-training model. lcon is the loss function of contrastive training. Figure 2 The right picture shows the framework of multi-task learning, where x1…x N is the input of the encoder (T5Encoder), d_Input2 is the input of the decoder (T5 Decoder), <eid1> , <eid2> , <eid3>The continuous prompt constructed is composed of 3 unused word tokens. The output of the decoder is output2, a t To generate words / answers, lgen is the loss function of multi-task learning, that is, the optimization objective of the generation task. < / eid2> < / eid1>
Claims
1. A sensitive text representation method based on contrastive learning, characterized in that: The following steps are involved: S1. Obtain positive and hard negative sample pairs of sensitive text and construct continuous templates, which are as follows: S101, obtain positive samples and hard negative sample pairs of sensitive text; S102. Based on the obtained positive sample and hard negative sample pairs, a prompt mechanism is used to construct a continuous template: Template=[Extra_id_1,Extra_id_2,Extra_id_3]+[X] Among them, Template represents a continuous template, Extra_id_1, Extra_id_2, and Extra_id_3 represent unused word tags, and [X] represents the sensitive text to be represented; S2. Build a multi-task contrastive learning model based on positive samples of sensitive text, hard negative sample pairs, and continuous templates; S3. Training the multi-task contrastive learning model, specifically: S301. In contrastive learning, the sentence is input into the pre-trained model T5, and the first word in its encoder is taken and marked as a sensitive text representation; S302: Using the sensitive text representation, training a sensitive representation model based on contrastive learning; S303. Based on the training results, four generation tasks are set in multi-task learning. S304. In multi-task learning, the following formula is used as the optimization objective of the generated task: in, represents the optimization goal of the generation task, t represents the iterative subscript, l Indicates the length of the sentence, Represents the input word, represents the probability of generating an answer based on the input word, Before t The words generated by the iteration step, Indicates the generated word; S305. Based on the optimization goal of the generated task, in contrastive learning and multi-task learning, the pre-trained model T5 is optimized using weight sharing to complete the training of the multi-task contrastive learning model; S4. Use the trained multi-task contrastive learning model to obtain the representation results of sensitive text.
2. The sensitive text representation method based on contrastive learning according to claim 1 is characterized in that: The step S2 comprises the following steps: S201, inputting positive samples, negative samples and hard negative samples into the pre-trained model T5 to obtain an output vector with a fixed length for contrastive learning; S202: Input the continuous template into the pre-trained model T5 to obtain the prediction output for multi-task learning, completing the construction of the multi-task contrastive learning model.
3. The sensitive text representation method based on contrastive learning according to claim 1 is characterized in that: The loss function expression of the sensitive representation model is: in, Represents the loss function of the sensitive representation model, log represents the logarithm with base e, j represents the iterative subscript, M represents the sum of positive samples and negative samples, represents the cosine similarity, represents the exponential function with base e, 、 and They represent the sensitive text representations of the original sentence, positive sample, and hard negative sample obtained after the pre-training model T5. Represents the temperature coefficient.
4. The sensitive text representation method based on contrastive learning according to claim 3 is characterized in that: The four generation tasks are: Generate Entailment by assuming Premise; Generate a contradiction by assuming the premise; Generate a hypothesis through entailment; Generate a hypothesis through contradiction.
5. The sensitive text representation method based on contrastive learning according to claim 4 is characterized in that: The objective function of the multi-task contrastive learning model is as follows: in, represents the objective function of the multi-task contrastive learning model, Represents the weight coefficient.