Knowledge distillation-based semantic text representation method and system and storage medium

By reinforcing the learning agent dynamically allocating knowledge contribution weights of multi-teacher models, combined with comparative learning and sorting distillation, the problem of semantic expression performance limitation caused by static weight allocation is solved, and efficient text vector representation and easy convergence training process is realized.

CN120354857APending Publication Date: 2025-07-22TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417075.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, during the knowledge distillation of multi-teacher models, the static weight allocation mechanism is difficult to adapt to changes in the knowledge reliability of the dynamic teacher model, resulting in limited semantic expression performance. The student model absorbs weak teacher supervision signals in the early stage of training, making it difficult to fully expand the representation space.

Method used

The knowledge contribution weights of the multi-teacher model are dynamically allocated by reinforcement learning agents. Through the knowledge distillation method, combined with comparative learning and sorting distillation, the weight allocation is optimized using Beta distribution sampling and depth deterministic strategy gradient algorithm to realize the interaction between the teacher model and the agent and parameter update.

Benefits of technology

It improves the accuracy and processing efficiency of text vector representation, the system is easy to train, and the algorithm is easy to converge, which improves the performance of text vector representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354857A_ABST
    Figure CN120354857A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic text representation method and system based on knowledge distillation, and a storage medium. The method comprises the following steps: respectively inputting original text data into a student model and a plurality of teacher models; the student model performs embedding processing on the original text data to obtain text vector representation; respectively calculating cosine similarity between the original text and the positive and negative samples by each teacher model, and inputting the cosine similarity as a first soft label into the intelligent agent; the teacher models are associated with the agents, and the weight corresponding to each teacher model is dynamically determined through interaction between the teacher models and the agents; the intelligent agent carries out aggregation processing on the first soft label according to the weight and outputs a second soft label; the second soft label is input into the student model after being subjected to fractional distillation processing, and final text vector representation is obtained; through association of each teacher model and the agent, adaptive distribution of contribution degree weights of the teacher models in different training stages is realized, the obtained text vector representation accuracy is high, and the method is easy to converge; and the method has outstanding performance in the aspect of text vector representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic text representation, and in particular to a semantic text representation method, system and storage medium based on knowledge distillation. Background Art

[0002] General text representation learning has always been a core research topic in the field of natural language processing (NLP), providing basic support for downstream tasks such as machine translation, information retrieval, and sentiment analysis. Although the unsupervised pre-training combined with task fine-tuning paradigm based on pre-trained language models (PLMs) such as BERT has performed well in many downstream tasks, studies have shown that the text vectors generated by the original PLM have significant limitations in representation quality. Specifically, multiple empirical analyses reveal that unoptimized PLM text representations present a narrow cone distribution in the vector space, a phenomenon defined by the academic community as the anisotropy problem. This anisotropic feature leads to distortion in the calculation of semantic similarity in the vector space, which seriously restricts its semantic expression performance. The main method to solve the anisotropy problem is to use contrastive learning methods to bring text representations with similar semantics closer in the vector space and distance text representations with far semantics. These methods usually involve different data augmentation methods to generate positive samples. Among them, the unsupervised text vector model SimCSE achieved good results in various tasks. The unsupervised approach does not require a large amount of manually labeled data and is also more effective in practical use. SimCSE treats all other texts in the same training batch as negative samples and ignores the differences between negative samples.

[0003] In recent years, contrastive learning methods based on ranking distillation (such as RankCSE) have constructed more sophisticated contrastive learning objectives by introducing the ranking knowledge of the teacher model, so that the student model can not only distinguish between positive and negative samples, but also effectively learn the relative semantic ranking relationship between texts, significantly improving the fine-grained semantic modeling ability. However, in the scenario of multi-teacher collaborative distillation, the method of assigning knowledge contribution weights to each teacher model based on artificial experience or pre-defined rules has significant limitations: first, the static weight allocation mechanism is difficult to adapt to the dynamic evolution characteristics of the knowledge reliability of heterogeneous teacher models; second, the fixed linear fusion method will reduce the noise tolerance of cross-model semantic associations under complex knowledge distribution; third, the student model itself has a dynamic knowledge absorption preference - in the early stage of training, due to the insufficient expansion of the representation space, it is easy to absorb weak teacher supervision signals, and as the representation ability is enhanced, it can gradually integrate the high-order semantic guidance of strong teachers. This staged learning feature is in conflict with the static configuration of the knowledge strength of the teacher group, highlighting the necessity of developing a dynamic weight optimization strategy. Summary of the invention

[0004] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a semantic text representation method, system and storage medium based on knowledge distillation of multiple teacher models, which uses a reinforcement learning agent to dynamically allocate weights to the knowledge (soft labels) of multiple teacher models.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A semantic text representation method based on knowledge distillation, characterized by comprising the steps of:

[0007] The original text data is respectively input into a first model and a plurality of second models;

[0008] The first model performs embedding processing on the original text data to obtain a text vector representation. Specifically, an encoder based on the Transformer architecture (such as BERT, RoBERTa) is used to vectorize the text. The vectorization process can be divided into the following steps: First, the input text is converted into a Token sequence by a tokenizer and special tokens (such as [CLS] and [SEP]) are added. The [CLS] token is used to represent the semantic information of the entire text, while the [SEP] token is used to separate the two texts in the text pair. Then, these Tokens are mapped to indices (TokenIDs) in the model vocabulary, and corresponding attention masks (Attention Mask) and Segment IDs are generated to distinguish actual words and padding words and different texts in the text pair. Then, the encoder encodes the processed input. Each layer of the encoder contains a multi-head self-attention mechanism and a feed-forward neural network, which can capture the complex dependencies between words in the text. In the encoder output, the hidden state (Hidden State) corresponding to each Token is generated, and the hidden state of the [CLS] token is usually used as the vector representation of the entire text.

[0009] The second model respectively calculates the cosine similarity between the original text and positive and negative samples, and inputs it as the first soft label into the third model;

[0010] The second model is associated with the third model, and the weight corresponding to each second model is determined by the interaction between the second model and the third model;

[0011] The third model aggregates the received first soft labels according to the weights and outputs a second soft label;

[0012] Calculate the distillation loss based on the similarity prediction values of the second soft label and the two columns of text, and backpropagate the gradient of the distillation loss function. Update the parameters of the first model and the third model in sequence according to the backpropagated gradient, and finally obtain the first model after cyclic training. Input the text into the final first model to obtain the final text vector representation.

[0013] In the present invention, preferably, the step of determining the weight corresponding to each of the second models includes:

[0014] The third model performs a linear transformation and ReLU activation on the input environmental state s through specific weights W α and W β , and biases b α and b β , to obtain α and β, and output α and β for use as a parameterized Beta distribution:

[0015] α = ReLU(W α s + b α );

[0016] β = ReLU(W β s + b β );

[0017] The second model samples respective weight values from the Beta(α, β) distribution, and the formula is:

[0018] π θ (s j , a j ) = Beta(α, β).

[0019] In the present invention, preferably, the third model aggregates the received first soft labels according to the weights and then outputs the second soft label, specifically using the formula:

[0020]

[0021] In the formula, and are the weights of the corresponding two second models, and are the first soft labels output by the corresponding two second models.

[0022] In the present invention, preferably, it further includes pre-training of the model:

[0023] Input the sample data into the first model and the second model respectively;

[0024] Multiple second models respectively calculate the cosine similarity between the original text and the positive and negative samples to obtain multiple soft labels n is the number of second models;

[0025] The second model interacts with the third model in an associated manner to determine the respective weights of the second models

[0026] The third model aggregates multiple soft labels into a final soft label

[0027] In the formula, x i is the i-th text of the input, is the weight of the n-th second model corresponding thereto, is the first soft label output by the n-th second model corresponding thereto.

[0028] The soft label output by the third model contains the fine-grained sorting information in the second model, and then the first model weights and fuses the sorting knowledge from multiple second models according to the final soft label output by the third model;

[0029] Repeat the above steps until the policy (Actor) network and the value (Critic) network in the third model converge, thereby completing the training of the first model and the third model.

[0030] In the present invention, preferably, the contribution weight of the second model is redefined as a continuous variable, and the update strategy of the third model is updated. The update strategy is to copy the weight parameters of the target policy (Target Actor) network in the third model to the Actor network in the same model.

[0031] A semantic text representation system based on knowledge distillation includes a first model, a second model, a third model, and a hierarchical distillation model. The first model and the second model are both connected to the third model. The hierarchical distillation model is connected between the first model and the third model. The third model includes an environment unit, an intelligent unit, and an experience replay pool unit. The environment unit is connected to both the intelligent unit and the experience replay pool unit. The intelligent unit and the experience replay unit are connected to perform data interaction.

[0032] In the present invention, preferably, the first model and the second model adopt pre-trained language models (such as RERT, RoBERTa), and the third model adopts a reinforcement learning framework, where the first model is a student model, the second model is a teacher model, and the third model is an agent.

[0033] In the present invention, preferably, the intelligent unit adopts a deep deterministic policy gradient (DDPG) algorithm structure, and uses a double neural network model for the policy function and the value function.

[0034] In the present invention, preferably, the intelligent unit introduces an experience replay mechanism. The experience data generated by the interaction between the Actor network of the intelligent unit and the environment unit is stored in the experience replay pool, and at the same time, a batch of samples are extracted from the experience replay pool for training to remove the correlation and dependence of the samples.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] The method of the present invention realizes the adaptive allocation of the contribution degree weights of the teacher models in different training stages by associating each teacher model with the corresponding intelligent agent, and the obtained text vector representation has high accuracy and high processing efficiency; the system of the present invention is easy to train, the algorithm is easy to converge, and it has outstanding performance in text vector representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic flow chart of the method of the present invention.

[0038] Figure 2 It is a schematic structural diagram of the system of the present invention.

[0039] Figure 3 It is a schematic structural diagram of the intelligent unit of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0042] Please refer to Figure 1 , a preferred embodiment of the present invention provides a semantic text representation method based on knowledge distillation, which combines contrastive learning, ranking distillation and reinforcement learning to solve the problem of finding appropriate weights in the ranking distillation process, including the steps of:

[0043] S1: The original text data (sentences are used as processing units in this embodiment) are respectively input into a first model and a plurality of second models to obtain sentence representations. The first model performs embedding processing on the original sentence to obtain a sentence vector representation;

[0044] In this embodiment, the first model and several second models are based on pre-trained language models (such as RERT, RoBERTa). The first model preprocesses the input sentence to obtain a better sentence vector representation, converts it into a template format, and extracts the masked field from the hidden state as the representation of the sentence vector:

[0045] The sentence: “X” means [MASK].

[0046] Specifically, an encoder based on the Transformer architecture (such as BERT, RoBERTa) is used to vectorize the sentence. The vectorization process can be divided into the following steps: First, the input sentence is converted into a Token sequence by a tokenizer and special tokens (such as [CLS] and [SEP]) are added. The [CLS] token is used to represent the semantic information of the whole sentence, while the [SEP] token is used to separate the two sentences in a sentence pair. Then, these Tokens are mapped to indices (Token IDs) in the model vocabulary, and corresponding attention masks (Attention Mask) and Segment IDs are generated to distinguish actual words and padding words and different sentences in the sentence pair. Then, the encoder encodes the processed input. Each layer of the encoder contains a multi-head self-attention mechanism and a feed-forward neural network, which can capture the complex dependencies between words in the sentence. In the encoder output, the hidden state (Hidden State) corresponding to each Token is generated, and the hidden state of the [CLS] token is usually used as the vector representation of the whole sentence.

[0047] S2: The second model calculates the cosine similarity between the original sentence and positive and negative samples respectively to obtain similarity scores, which are used as the first soft labels and input to the third model;

[0048] S3: The second model is associated with the third model, and the weights corresponding to each second model are determined through the interaction between the second model and the third model;

[0049] In this embodiment, the steps of determining the weights corresponding to each second model include:

[0050] S31: The third model processes the input environmental state s, where the state s includes sentence embedding vectors, soft labels of the teacher model, and cross-entropy loss functions, through specific weights W α and W β , and biases b α and b β, perform a linear transformation and ReLU activation to obtain α and β, and output α and β for use as parameters of the Beta distribution:

[0051] α = ReLU(W α s + b α );

[0052] β = ReLU(W β s + b β );

[0053] S32: The second model randomly samples its respective weight values from the Beta(α, β) distribution. The process of Beta sampling is the process of generating random samples from the Beta distribution. The Beta distribution is a continuous probability distribution defined on the interval [0, 1], controlled by two shape parameters α and β, and its probability density function is where B(α, β) is the Beta function. Beta sampling can be implemented in various ways. One common method is sampling based on the Gamma distribution. Utilizing the relationship between the Beta distribution and the Gamma distribution, first generate two independent Gamma distribution random variables X ~ Gamma(α, 1) and Y ~ Gamma(β, 1), and then through calculate to obtain a random sample that follows the Beta distribution Beta(α, β). Another method is acceptance-rejection sampling, which is applicable to cases where it is not possible to directly generate Beta distribution samples. First, select an easily sampled proposal distribution g(x) (such as a uniform distribution), and find a constant M such that M·g(x) is always greater than the probability density function f(x) of the Beta distribution. Then generate a candidate sample x from g(x), and generate a random u from the uniform distribution U(0, 1). If then accept x as a sample of the Beta distribution, otherwise reject it and repeat the sampling process. Through Beta sampling, random samples that conform to a specific distribution can be generated, thus supporting various probability analysis and stochastic simulation tasks. The respective weight values are obtained, and the formula is:

[0054] π θ (s j , a j ) = Beta(α, β).

[0055] In this embodiment, instead of treating an action as a discrete variable to represent the selection of the second model or not, the action is defined as the weight of the second model, which is a continuous variable with a value range between [0, 1]. The present invention represents the probability distribution of the action by outputting two parameters of Beta(α, β), where α and β are the shape parameters of Beta(α, β). In the Beta(α, β) distribution, α can be regarded as the number of successes, and β represents the number of failures. Therefore, the action value sampled from Beta (ranging from 0 to 1) essentially represents the "intensity" of a certain behavior or the "possibility" of performing this behavior. In the research of the present invention, a higher action value may indicate that the teacher model tends to select this behavior, which may be because according to historical data, the success probability of this behavior is relatively high. Specifically, a higher α value may indicate an increase in the model's confidence in the success of the action, while a higher β value may indicate an increase in the model's confidence in the failure of the action. This parameterization method based on historical experience can help the model better balance the selection between different behaviors.

[0056] S4: The third model aggregates the received first soft labels according to the weight and then outputs a second soft label, specifically using the formula:

[0057]

[0058] where x i is the i-th input sentence, is the corresponding weight of the n-th second model, is the first soft label output by the corresponding n-th second model.

[0059] S5: Calculate the distillation loss between the second soft label and the similarity prediction value of two columns of sentences, and perform gradient backpropagation on the distillation loss function. Update the parameters of the first model and the third model in sequence according to the backpropagated gradient. Finally, obtain the first model after cyclic training, and input the sentence into the final first model to obtain the final sentence vector representation.

[0060] In this embodiment, it also includes pre-training of the model:

[0061] A1: Input the sample data into the first model and the second model respectively;

[0062] A2: Multiple second models calculate the cosine similarity between the original sentence and the positive and negative samples respectively, and obtain multiple soft labels n is the number of second models;

[0063] A3: The second model and the third model are associated and interacted to determine the respective weights of the second models

[0064] A4: The third model aggregates multiple soft labels into a final soft label

[0065] A5: The soft label output by the third model contains the fine-grained ranking information in the second model, and then the first model extracts and learns ranking knowledge from the second model according to the final soft label;

[0066] A6: Repeat the above steps A1 - A5, and use the Deep Deterministic Policy Gradient (DDPG) algorithm as the update policy until the Actor network and the Critic network in the third model converge, thereby completing the training of the first model and the third model.

[0067] As Figure 2 shown, another preferred embodiment of the present invention provides a semantic text representation system based on knowledge distillation, including a first model, a second model, a third model, and a hierarchical distillation model. The first model and the second model are both connected to the third model, and the hierarchical distillation model is connected between the first model and the third model. The third model includes an environment unit, an intelligent unit, and an experience replay pool unit. The environment unit is connected to both the intelligent unit and the experience replay pool unit, and the intelligent unit and the experience replay unit are connected to perform data interaction.

[0068] In this embodiment, the first model and the second model adopt pre-trained language models (such as RERT, RoBERTa), and the third model adopts a reinforcement learning framework, where the first model is a student model, the second model is a teacher model, and the third model is an agent.

[0069] As Figure 3 shown, in this embodiment, the intelligent unit adopts a Deep Deterministic Policy Gradient (DDPG) dual neural network structure, and uses a dual neural network model for the policy function and the value function. The intelligent unit includes an Actor network, a Target Actor network, a Critic network, and a Target Critic network, and the overall learning process of the intelligent unit is more stable and the convergence speed is accelerated.

[0070] In this embodiment, the intelligent unit introduces an experience replay mechanism. The experience data generated by the interaction between the Actor network of the intelligent unit and the environment unit is stored in the experience replay pool, and at the same time, a batch of samples are extracted from the experience replay pool for training to remove the correlation and dependence of the samples, making the overall third model easier to converge.

[0071] In this embodiment, the constructed system is verified:

[0072] 1. Datasets and evaluation methods

[0073] The proposed method is trained on the English Wikipedia dataset and evaluated on the Semantic Text Similarity (STS) dataset. The STS dataset includes STS2012 - STS2016 (STS12 - STS16), STSbenchmark (STS - B), and SICK - Relatedness (SICK - R). This dataset consists of sentence pairs extracted from titles, news articles, videos, images, and natural language inference tasks, with manually assigned semantic similarity scores ranging from 0 to 5. In the training phase, referring to the unsupervised training paradigm of SimCSE, 1 million sentences are randomly selected from Wikipedia for sentence representation learning. The final evaluation is achieved through the following process: First, calculate the cosine similarity scores of the sentence pairs in the STS dataset by the model, and then perform Spearman rank - correlation analysis on these scores and the manually annotated similarity scores. The obtained correlation coefficient is used as a measure of the model's performance.

[0074] 2. Establishing a baseline

[0075] The method of the present invention is compared and analyzed with several state - of - the - art unsupervised learning text representation methods at home and abroad. This includes models relying on post - processing techniques, such as BERT_flow and BERT_whitening, and contrastive learning - based methods, such as ConSERT, SimCSE, DCLR, RankEncoder, SNCSE, and RankCSE. All baseline models complete the STS benchmark test under the same experimental conditions (using BERT_base as the backbone network) to ensure that the performance differences of the methods can accurately reflect the advantages and disadvantages of the model architectures.

[0076] 3. Main results

[0077] As shown in Table 1, the method of the present invention achieves optimal performance on STS2012 to STS2016 and STS - B: it improves by 4.84% compared to SimCSE based on contrastive learning; it improves by 0.73% compared to RankCSE using contrastive learning and ranking distillation.

[0078] Table 1 Evaluation results of various sentence vector representation methods on the STS task under the unsupervised setting (the optimal performance is shown in bold)

[0079]

[0080]

[0081] 4. Ablation experiments

[0082] 4.1, Effectiveness of Adaptive Weight Allocation

[0083] To evaluate the contribution of each module of the system evaluation model, this experiment designed a three-stage ablation experiment on the STS-B dataset: First, add the Bidirectional Margin Loss (BML) loss function to the baseline model SimCSE, then add the Rank Distillation Loss (RD) loss function, and finally add the proposed Deep Deterministic Policy Gradient (DDPG) loss function, while keeping the other architecture parameters unchanged. As shown in Table 2, the complete model (BML+RD+DDPG) achieved a Spearman rank correlation coefficient of 82.85%, while gradually removing components led to a step-by-step decline in performance (BML+RD: 80.52%; only BML: 78.53%). This quantitative analysis confirmed that the three modules have a synergistic enhancement effect on semantic similarity modeling, and the DDPG loss function contributed the largest single-module gain (+2.33%).

[0084] Table 2 Influence of Different Modules on Model Performance

[0085] Model STS-B SimCSE 76.85 SimCSE + BML 78.53 SimCSE + BML + RD 80.52 SimCSE + BML + RD + DDPG 82.85

[0086] 4.2, Effectiveness of Knowledge Distillation of Multi-Teacher Model

[0087] To verify the effectiveness of the multi-teacher knowledge distillation framework, a comparative experiment was designed on the STS dataset: Set two distillation configurations of single-teacher and double-teacher models. As shown in Table 3, the multi-teacher integration strategy using rank distillation and dynamic weight allocation outperformed each single-teacher baseline score. This indicates that the multi-teacher cooperation mechanism can optimize the representation space through semantic difference complementarity and improve the generalization ability of the model. It is worth noting that when RankEncoder is used as an independent teacher, the supervised BERT-base model obtained a correlation coefficient of 80.26%, which is 0.19 percentage points higher than the performance of the teacher model itself (80.07%, see Table 1). This result shows that the adaptive weight allocation method proposed in the present invention is not only applicable to multi-teacher models, but also can effectively improve the knowledge distillation effect of single-teacher models.

[0088] Table 3 Influence of DDPG-Based Knowledge Distillation on Different Teacher Models

[0089] Teacher Model 1 Teacher Model 2 Student Model STS Average <![CDATA[SimCSE_ BERT-large > \ BERT-base 79.81 RankEncoder \ BERT-base 80.26 RankEncoder <![CDATA[SimCSE_ BERT-large > BERT-base 81.09

[0090] In some other preferred embodiments of the present invention, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the method as described in the above embodiments.

[0091] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0092] The above description is a detailed description of the preferred and feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. All equivalent changes or modifications made under the technical spirit disclosed by the present invention should fall within the scope of the patent covered by the present invention.

Claims

1. A semantic text representation method based on knowledge distillation, characterized in that Including the steps: The original text data is respectively input into the first model and several second models; The first model processes the original text data to obtain a text vector representation; The second models respectively calculate the cosine similarities between the original text and positive and negative samples, and use them as the first soft labels to be input into the third model; The second models are associated with the third model, and through the interaction between the second models and the third model, the weights corresponding to each second model are determined; The third model aggregates the received first soft labels according to the weights and then outputs the second soft labels and the similarity prediction values of two columns of texts; The second soft labels and the similarity prediction values of two columns of texts are used to calculate the distillation loss, and the gradient of the distillation loss function is backpropagated. According to the backpropagated gradient, the parameters of the first model and the third model are updated in sequence. Finally, the first model after cyclic training is obtained, and the text is input into the final first model to obtain the final text vector representation.

2. The semantic text representation method based on knowledge distillation according to claim 1, wherein The steps of determining the weights corresponding to each second model include: The third model takes the input environmental s data, which includes text embedding vectors, soft labels of the teacher model, and cross-entropy loss functions, and through weights W α and W β , as well as biases b α and b β , performs a linear transformation and ReLU activation to obtain α and β, and outputs α and β for use as parameters of the parametric Beta distribution: α = ReLU(W α s + b α ); β = ReLU(W β s + b β ); The second models sample their respective weight values from the Beta(α,β) distribution, and the formula is: π θ (s j ,a j ) = Beta(α, β).

3. The semantic text representation method based on knowledge distillation according to claim 2, wherein The specific formula for the third model to aggregate the received first soft labels according to the weights and then output the second soft labels is: where x i is the i-th input text, is the second model weight corresponding to the n-th one, is the first soft label output by the second model corresponding to the n-th one.

4. The semantic text representation method based on knowledge distillation according to claim 1, wherein It also includes the pre-training of the model: The sample data is respectively input into the first model and the second model; Multiple second models respectively calculate the cosine similarity between the original text and positive and negative samples to obtain multiple soft labels n is the number of second models; The second model is associated and interacted with the third model to determine the respective weights of the second model The third model aggregates multiple soft labels into a final soft label The soft labels output by the third model contain the fine-grained ranking information in the second model. Then, the first model extracts and learns the ranking knowledge from the second model according to the final soft labels output by the third model; Repeat the above steps until the Actor network and the Critic network in the third model converge, thereby completing the training of the first model and the third model.

5. The semantic text representation method based on knowledge distillation according to claim 1, characterized in that Redefine the weights of the second model as continuous variables, and update the update strategy of the third model. Copy the weight parameters of the Target Actor network in the third model to the Actor network in the same model.

6. A semantic text representation system based on knowledge distillation, characterized in that Including a first model, a second model, a third model, and a hierarchical distillation model. The first model and the second model are both connected to the third model. The hierarchical distillation model is connected between the first model and the third model. The third model includes an environment unit, an intelligent unit, and an experience replay pool unit. The environment unit is connected to both the intelligent unit and the experience replay pool unit. The intelligent unit and the experience replay unit are connected to perform data interaction.

7. The semantic text representation system based on knowledge distillation according to claim 6, wherein The first model and the second model adopt pre-trained language models, and the third model adopts a reinforcement learning framework.

8. The semantic text representation system based on knowledge distillation according to claim 6, characterized in that The intelligent unit adopts a deep deterministic policy gradient algorithm structure, and uses a double neural network model for the policy function and the value function.

9. The semantic text representation system based on knowledge distillation according to claim 6, characterized in that, The intelligent unit introduces an experience replay mechanism. The experience data generated by the interaction between the Actor network of the intelligent unit and the environment unit is stored in the experience replay pool. At the same time, a batch of samples are extracted from the experience replay pool for training to remove the correlation and dependence of the samples.

10. A storage medium, characterized in that, There is a computer program stored. When the computer program is executed by a processor, the processor is caused to execute the steps of the semantic text representation method based on knowledge distillation as described in any one of claims 1-5 above.