Model pre-training method based on mask automatic encoder and noise enhancement

By adopting a model pre-training method based on mask automatic encoder and noise enhancement in intensive search tasks, the problem of insufficient sentence-level representation ability of pre-trained models is solved, and higher robustness and semantic representation ability are achieved, which is suitable for zero-sample and supervised learning scenarios.

CN119940470AActive Publication Date: 2025-05-06JIANGNAN UNIV

Patent Information

Application Number
CN202510423102.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The sentence-level representation ability of pre-trained models in existing intensive search tasks is insufficient, especially in the face of input perturbations, and traditional pre-training methods are difficult to meet the semantic distinction requirements of search scenarios.

Method used

A model pre-training method based on mask automatic encoder and noise enhancement is adopted to build an asymmetric encoding-decoding model. Through differentiated masking strategies and dynamic noise injection mechanism, the model's robustness to anti-perturbation is enhanced, and noise generation is guided through KL divergence to optimize the noise generation direction.

Benefits of technology

It significantly improves the robustness and semantic representation of the model, can capture sentence-level semantic associations more effectively, enhances the anti-interference ability to input perturbations, and performs well in zero-sample and supervised learning scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940470A_ABST
    Figure CN119940470A_ABST
Patent Text Reader

Abstract

The invention provides a model pre-training method based on a mask automatic encoder and noise enhancement, and relates to the technical field of natural language processing, and the method comprises the steps: constructing an asymmetric encoding-decoding model, and improving the diversity of training signals through differentiated mask proportions and decoding mechanisms. Then, a noise injection mechanism is introduced, the robustness of disturbance is resisted by embedding and adding a noise enhancement model, and two improvements are proposed: 1, the noise amplitude is dynamically adjusted, the robustness is enhanced by using relatively large noise in the initial stage of training, and the precision is improved by reducing the noise in the later stage; secondly, noise generation is guided through KL divergence in the later training period, the distribution difference between original embedding and noise adding embedding is measured, and noise is more intelligent for the weakness of the model. And pre-training and evaluation are performed on a plurality of data sets, so that the dense retrieval performance in a zero sample and supervised learning scene is remarkably improved. Finally, according to the method, the accuracy and the stability of sentence representation in the retrieval task can be improved without extra fine adjustment of a model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a model pre-training method based on masked autoencoder and noise enhancement. Background Art

[0002] In existing dense retrieval and natural language processing systems, traditional methods mostly rely on vocabulary-based retrieval techniques, such as TF-IDF and BM25, to determine relevance by calculating word frequency statistics or probability scores between queries and documents. These methods have shown certain effectiveness in simple text matching tasks, especially for keyword-driven retrieval scenarios. However, they mainly focus on surface vocabulary matching and have difficulty capturing the deep semantic associations between queries and documents. Especially when dealing with complex sentences or cross-document information, they often fail to meet users' higher demands for semantic understanding. With the development of deep learning technology, embedding-based language models (such as BERT, RoBERTa, etc.) have been introduced into the field of dense retrieval. By converting sentences into vector representations and calculating semantic similarity, the ability to understand contextual semantics has been significantly improved. These models can identify synonymous expressions and contextual dependencies, surpassing the limitations of traditional methods. However, despite the progress made by embedding models in semantic retrieval, they still face several key challenges.

[0003] First, existing embedding models mainly rely on general semantic representations in the pre-training stage in intensive retrieval tasks. Although they can capture sentence-level semantic information, they are not robust enough when facing input perturbations (such as spelling errors, word order changes, or noisy data). This fragility causes the model to produce unstable retrieval results due to data variation in actual application scenarios, especially in large-scale real retrieval tasks. Secondly, traditional pre-training methods usually optimize models based on tag-level tasks (such as masked language modeling). The generated sentence representations are often not targeted in intensive retrieval tasks and are difficult to fully adapt to the requirements of retrieval scenarios for sentence-level semantic discrimination. In addition, existing embedding models usually do not consider the training signals of adversarial perturbations during the training process, resulting in reduced performance in complex or noisy environments. For example, when processing diverse queries on datasets such as MSMARCO, the model may not be able to effectively distinguish subtle semantic differences.

[0004] In addition, in order to improve the performance of the model in specific retrieval tasks, existing methods often require frequent fine-tuning of pre-trained models to adapt to different data sets or domain requirements. This fine-tuning process not only consumes a lot of computing resources, but also increases the complexity of deployment and is difficult to maintain efficiency in large-scale systems. Even if the model performance is improved through fine-tuning, the model may still generate inaccurate embedding representations when faced with incomplete or noisy inputs, affecting the quality of retrieval results. The traditional pre-training paradigm lacks joint optimization of robustness and semantic reconstruction, which limits the model's transferability in zero-sample or supervised learning scenarios, making it difficult to directly meet the requirements of intensive retrieval tasks for high-quality sentence representation. Therefore, how to improve the robustness and semantic representation capabilities of the model in the pre-training stage has become a technical problem that needs to be solved urgently in the current intensive retrieval field. Summary of the invention

[0005] To this end, an embodiment of the present invention provides a model pre-training method and system based on masked autoencoder and noise enhancement, which is used to solve the problem of insufficient sentence-level representation ability of pre-trained models in intensive retrieval tasks in the prior art.

[0006] In order to solve the above problems, an embodiment of the present invention provides a model pre-training method based on masked autoencoder and noise enhancement, the method comprising: S1: Construct an asymmetric encoder-decoder model, where the encoder is a full-scale deep neural network for generating an embedded representation of the input sentence, and the decoder is a lightweight single-layer neural network for reconstructing the input sentence; S2: A differentiated masking strategy is applied to the input sentence, where the encoder masking ratio is 15%-30% and the decoder masking ratio is 50%-75% to generate differentiated training signals; S3: Introducing a dynamic noise injection mechanism in the encoder embedding generation process, by adding noise of different amplitudes at different stages of the training process, the robustness of the asymmetric encoding-decoding model against perturbations is enhanced; S4: Use KL divergence to guide noise generation and optimize the noise generation direction by calculating the distribution difference between the original embedding and the noisy embedding; S5: Use multiple datasets to pre-train the asymmetric encoding-decoding model, and optimize the parameters of the asymmetric encoding-decoding model through semantic reconstruction and noise adversarial joint loss; S6: The pre-trained asymmetric encoder-decoder model is directly applied to dense retrieval tasks in zero-shot or supervised learning scenarios to output a vector representation of the sentence.

[0007] Preferably, in the asymmetric encoding-decoding model, the encoder adopts a multi-layer Transformer architecture and outputs a fixed-dimensional embedding vector, and the decoder adopts a single-layer Transformer architecture, the number of its parameters does not exceed 1 / 10 of the number of parameters of the encoder, and the parameters of the encoder and decoder are initialized through a pre-trained language model.

[0008] Preferably, step S2 specifically includes: For the encoder mask operation, first determine the mask ratio , the range is , calculate the number of words that need to be masked , then from the sentence position Randomly select positions, denoted as a set ;for Each position in , the word Replace with special token [MASK] to generate masked sentences , which is in the form of , or , ; This mask sentence As encoder input, used to generate embedding and the masked word vector ; In contrast, the decoder masking operation uses a higher masking ratio , ranging from 50%-75%, calculate the number of masked words , and from Randomly select positions, denoted as a set , May contain Similarly, Words in Replace with special marker , generate masked sentences , , or , , and as decoder input and position encoding Together they are used to reconstruct the original sentence ; in, represents the floor function, represents the length of the sentence, Encoder(·) is the encoder function, Decoder function.

[0009] Preferably, in the dynamic noise injection mechanism, the noise amplitude is dynamically adjusted according to the training progress, and its calculation formula is: ; in, represents the dynamic noise amplitude, represents the initial noise amplitude, Represents the attenuation rate, which is used to control the speed of amplitude decrease. represents the total number of training steps, Indicates the current training step number, Indicates the minimum amplitude of noise.

[0010] Preferably, in the KL divergence guided noise generation, the noise direction 𝜂 is calculated by the following formula: ; in, represents the gradient under Gaussian approximation, pointing to the direction with the largest distribution difference, represents the gradient operation on the noise vector N, represents the KL divergence, Indicates the dynamic noise amplitude; is used to measure the difference between two embedding distributions. and The original embedding and noisy embedding The probability distribution of is a random vector sampled from the unit sphere, providing an exploratory perturbation; is the dynamic weight, ranging from , Denotes L2 norm normalization.

[0011] Preferably, the noise embedding It is expressed as: ; in, represents the dynamic noise amplitude, is the optimal noise direction under the influence of KL divergence, is the regularization coefficient, adjusted through experiments, Represents the original embedding The standard deviation of Represents the minimum function.

[0012] Preferably, the semantic reconstruction and noise adversarial joint loss is: ; in, For semantic reconstruction and noise adversarial joint loss, and are weight coefficients, ranging from 0 to 1. is the semantic reconstruction loss, For noise combating losses.

[0013] Preferably, the semantic reconstruction loss is: ; in, For the mask sentence, are decoder parameters, is the sentence length, Indicates that the decoder is given a mask sentence and decoder parameters Under the condition of The words are probability.

[0014] Preferably, the noise resistance loss is: ; in, represents a masked sentence, represents the masked sentence after adding dynamic noise, represents the square of L2 norm, and Encoder(·) is the encoder function.

[0015] The embodiment of the present invention further provides a model pre-training system based on masked autoencoder and noise enhancement, which is used to implement the above-mentioned model pre-training method based on masked autoencoder and noise enhancement, and specifically includes: Asymmetric encoding-decoding model building module, used to build an asymmetric encoding-decoding model, where the encoder is a full-scale deep neural network for generating an embedded representation of the input sentence, and the decoder is a lightweight single-layer neural network for reconstructing the input sentence; A differentiated masking strategy module is used to apply a differentiated masking strategy to the input sentence, where the encoder masking ratio is 15%-30% and the decoder masking ratio is 50%-75% to generate differentiated training signals; Dynamic noise injection module, which is used to introduce a dynamic noise injection mechanism in the encoder embedding generation process. By adding noise of different amplitudes at different stages of the training process, the robustness of the asymmetric encoding-decoding model against perturbations is enhanced; The KL divergence-guided noise generation module is used to use the KL divergence to guide the noise generation and optimize the noise generation direction by calculating the distribution difference between the original embedding and the noisy embedding; Model pre-training and parameter optimization module, which is used to pre-train the asymmetric encoding-decoding model using multiple datasets and optimize the parameters of the asymmetric encoding-decoding model through semantic reconstruction and noise adversarial joint loss; The model application module is used to directly apply the pre-trained asymmetric encoding-decoding model to dense retrieval tasks in zero-shot or supervised learning scenarios, and output the vector representation of the sentence.

[0016] An embodiment of the present invention also provides an electronic device, which includes a processor, a memory and a bus system, wherein the processor and the memory are connected through the bus system, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the above-mentioned model pre-training method based on mask autoencoder and noise enhancement.

[0017] An embodiment of the present invention also provides a computer storage medium, which stores a computer software product. The computer software product includes a number of instructions for enabling a computer device to execute the above-mentioned model pre-training method based on masked autoencoder and noise enhancement.

[0018] It can be seen from the above technical solutions that the present invention has the following beneficial effects: (1) Dual improvement of robustness and semantic representation ability: This paper significantly enhances the model's ability to resist input disturbances through dynamic noise injection mechanism and adversarial training, and improves the robustness of sentence representation. The asymmetric encoding-decoding structure is combined with the semantic reconstruction task of high-ratio masking to effectively capture sentence-level semantic associations and enhance semantic representation ability.

[0019] (2) Efficient computing and strong cross-scenario adaptability: The lightweight single-layer decoder significantly reduces computational complexity, accelerates model convergence, reduces resource consumption, and facilitates efficient computing and deployment. The model performs well in both zero-shot and supervised learning scenarios, has strong cross-scenario adaptability, and its generalization ability is significantly better than traditional models.

[0020] (3) Convenient end-to-end application and innovation in theory and practice: The generated sentence vectors can be directly used for intensive retrieval tasks, which simplifies the deployment process and is suitable for large-scale real-time retrieval systems. The combination of masked autoencoders and adversarial training provides a new paradigm for retrieval-oriented language model pre-training, which combines methodological innovation with practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the implementation cases of the present invention or the technical solutions in the prior art, the following is a brief description of the drawings required for use in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be understood as limiting the present invention in any way. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them: Figure 1 A flowchart of a model pre-training method based on masked autoencoder and noise enhancement provided in an embodiment; Figure 2 It is a schematic diagram of the structure of the encoder in the embodiment; Figure 3 A schematic diagram of the structure of a decoder in an embodiment; Figure 4 A schematic diagram of a dynamic noise injection architecture in an embodiment; Figure 5 Schematic diagram of a KL divergence-guided dynamic noise injection architecture in an embodiment; Figure 6 A block diagram of a model pre-training system based on masked autoencoder and noise enhancement provided in an embodiment. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] Embodiment 1: In order to solve the problem of insufficient sentence-level representation capability of pre-trained models in intensive retrieval tasks in the prior art. Figure 1 As shown, an embodiment of the present invention proposes a model pre-training method based on masked autoencoder and noise enhancement, the method comprising: S1: Construct an asymmetric encoder-decoder model, where the encoder is a full-scale deep neural network for generating an embedded representation of the input sentence, and the decoder is a lightweight single-layer neural network for reconstructing the input sentence; S2: A differentiated masking strategy is applied to the input sentence, where the encoder masking ratio is 15%-30% and the decoder masking ratio is 50%-75% to generate differentiated training signals; S3: Introducing a dynamic noise injection mechanism in the encoder embedding generation process, by adding noise of different amplitudes at different stages of the training process, the robustness of the asymmetric encoding-decoding model against perturbations is enhanced; S4: Use KL divergence to guide noise generation and optimize the noise generation direction by calculating the distribution difference between the original embedding and the noisy embedding; S5: Use multiple datasets to pre-train the asymmetric encoding-decoding model, and optimize the parameters of the asymmetric encoding-decoding model through semantic reconstruction and noise adversarial joint loss; S6: The pre-trained asymmetric encoder-decoder model is directly applied to dense retrieval tasks in zero-shot or supervised learning scenarios to output a vector representation of the sentence.

[0024] From the above technical solution, it can be seen that the present invention proposes a model pre-training method based on masked autoencoder and noise enhancement. First, an asymmetric encoding-decoding model is constructed. The encoder uses a full-scale neural network to convert sentences into embeddings, and the decoder is a single-layer neural network to reconstruct sentences from masked inputs to balance efficiency and ability; then a differentiated masking strategy is applied to the input sentences, with the encoder mask ratio of 15%-30% and the decoder mask ratio of 50%-75% to optimize the embedding and reconstruction tasks; then a dynamic noise injection mechanism is introduced in the encoder embedding generation process, and the noise amplitude changes from large to small from the initial stage to the final stage to support the stability of the retrieval task. At the same time, the KL divergence is used to guide the noise generation, and the random vector is combined with the threshold control intensity to improve the embedding discrimination; then the model is pre-trained using multiple data sets such as MS MARCO, and the parameters of the asymmetric encoding-decoding model are optimized through semantic reconstruction and noise confrontation joint loss, and the performance is evaluated; finally, the pre-trained model is applied to dense retrieval tasks. This method can improve the accuracy and stability of sentence representation in retrieval tasks without additional fine-tuning of the model, providing a new path for efficient retrieval-oriented language model pre-training.

[0025] In step S1, an asymmetric encoding-decoding model is constructed, where the encoder is a full-scale deep neural network for generating an embedded representation of the input sentence, and the decoder is a lightweight single-layer neural network for reconstructing the input sentence.

[0026] Specifically, building an asymmetric encoder-decoder model is the basic step of the entire pre-training method, which aims to generate text representations suitable for dense retrieval tasks for input sentences, while providing additional training supervision signals through reconstruction tasks. The encoder is designed as a neural network based on a multi-layer Transformer architecture that can capture the input sentence (in Indicates words, The main function of the encoder is to convert the input sentence into an embedding vector of fixed dimension. : ; in, is the Transformer function, are encoder parameters, , is the dimension of the encoder output embedding vector, typically 512, 768, or 1024.

[0027] Furthermore, the encoder contains a multi-head self-attention mechanism, a feedforward neural network, and layer normalization operations to ensure the stability and representation ability of the deep network. At the same time, the decoder is designed as a lightweight single-layer neural network, also based on the Transformer architecture, but only uses a single-layer structure, and its parameter volume is much smaller than the encoder's parameter volume (about 1 / 10 or less of the encoder's parameter volume) to reduce computational complexity. The decoder's task is to reconstruct the original sentence from the masked input: ; in, is the masked input, is a single-layer Transformer function, are decoder parameters, For the reconstruction results.

[0028] The core of this asymmetric design is that the encoder generates high-quality embedding vectors through a full-scale network to support retrieval tasks, while the decoder assists pre-training through a lightweight structure to ensure semantic consistency. First, it is converted into a token sequence through a tokenizer and special tags (such as [CLS] and [SEP]) are added. The encoder outputs a fixed-length vector, and the decoder output is consistent with the input length. To accelerate training convergence, the parameters of the encoder and decoder are and Pre-trained language model weights (such as BERT) can be used for initialization. This initialization strategy can not only significantly improve the training efficiency of the model, but also achieve better performance with limited training data, laying the foundation for subsequent masking strategies and noise enhancement.

[0029] In step S2, a differentiated masking strategy is applied to the input sentence, where the encoder masking ratio is 15%-30% and the decoder masking ratio is 50%-75% to generate differentiated training signals.

[0030] Specifically, a differentiated masking strategy is applied to the input sentences. The purpose of this step is to generate differentiated training inputs for the encoder and decoder through different mask ratios, thereby enhancing the diversity of training signals and optimizing their respective task objectives.

[0031] Specifically, if Figure 2 and Figure 3 As shown, for the input sentence , the mask ratio on the encoder side is set to 15%-30%, while the mask ratio on the decoder side is set to 50%-75%, and the mask is randomly sampled, constraining the continuous mask not to exceed 50%. For the encoder mask operation, first determine the mask ratio (Range ), calculate the number of words that need to be masked ( represents the floor function, indicates sentence length), and then from the sentence position Randomly select positions, denoted as a set .for Each position in , the word Replace with special token [MASK] to generate masked sentences , which is in the form of (like )or (like ). This mask sentence As encoder input, used to generate embedding , where Encoder(·) is the encoder function.

[0032] In contrast, the decoder masking operation uses a higher masking ratio (Range 50%-75%), calculate the number of masked words , and from Randomly select positions, denoted as a set ( May contain ). Similarly, Words in Replace with , generate masked sentences ,(like )or (like ), and is used as the decoder input to reconstruct the original sentence ,in Decoder function.

[0033] The core of this differentiated masking strategy is that the encoder mask ratio is low to retain more semantic information and facilitate the generation of high-quality embeddings, while the decoder mask ratio is high to increase the reconstruction difficulty and improve the generalization ability of the model. The decoder loss function The optimization goal is: ; in, Indicates the first The original word at position represents the set of all masked positions in the input sentence, It is a cross entropy loss function used for the decoder to restore the masked sentence. The high mask ratio forces the decoder to generate high-quality vector representation.

[0034] In step S3, a dynamic noise injection mechanism is introduced in the encoder embedding generation process to enhance the robustness of the asymmetric encoding-decoding model against perturbations by adding noise of different amplitudes at different stages of the training process.

[0035] Specifically, after completing the entire autoencoder reconstruction step, a dynamic noise injection mechanism is introduced into the encoder embedding generation process, such as Figure 4 As shown in Figure 2. This method enhances the robustness of the model against perturbations by adding dynamically adjusted noise to the embedding layer, while balancing the accuracy requirements during training, thereby improving the stability and practicality of the pre-trained model in dense retrieval tasks. The goal of this mechanism is to enable the sentence embeddings generated by the encoder to maintain better semantic consistency in the vector space and meet the dual requirements of robustness and accuracy during the training phase.

[0036] Specifically, the encoder first constructs a masked sentence based on Generate raw embeddings : ; in, is an embedding vector of fixed dimension.

[0037] Furthermore, the task of noise injection is to embed Add random perturbations to generate noisy embeddings ,in is a noise vector that not only simulates real-world disturbances but also dynamically adjusts its characteristics according to the training progress. In the early stages of training, this task focuses on enhancing robustness by applying larger noise amplitudes (e.g. 0.3 to 0.5 times the standard deviation) forces the model to learn robust feature representations and prevent overfitting to specific patterns; in the later stages of training, the task shifts to improving accuracy, and the noise amplitude gradually decreases to lower values ​​(such as To achieve this, the dynamic noise amplitude is defined as: ; in, represents the dynamic noise amplitude, represents the initial noise amplitude, Represents the attenuation rate (floating point number), which is used to control the speed of amplitude decrease. represents the total number of training steps, Indicates the current training step number, Indicates the minimum amplitude of noise (take ), to ensure the accuracy of the later stage.

[0038] This step uses dynamic noise injection to strengthen the uncertainty and recall effect of the input sentence when the semantics are ambiguous and when encountering similar sentences. Through this dynamic noise injection, the model can not only adapt to the dense retrieval requirements under various variation conditions, but also maintain high-precision matching capabilities in undisturbed scenarios.

[0039] In step S4, KL divergence is used to guide noise generation, and the noise generation direction is optimized by calculating the distribution difference between the original embedding and the noisy embedding.

[0040] Specifically, KL divergence is used to guide noise generation, such as Figure 5 The purpose of this step is to optimize the direction of noise generation by quantifying the distribution difference between the original embedding and the noisy embedding, so that the noise can more specifically enhance the weaknesses of the model and improve its robustness and adaptability in dense retrieval tasks. The goal of this task is not to simply add random noise, but to make the noise injection "intelligent" through information theory metrics, focusing on exposing and improving the potential defects of the encoder. Specifically, for the original embedding generated by the encoder and noisy embedding (in ), KL divergence is used to measure the difference between two embedding distributions, where and They are and By minimizing or controlling To adjust the noise vector direction, so More inclined to deviate The noise direction is defined as a composite vector combining the KL divergence guide and randomness: ; in, represents the gradient under Gaussian approximation, pointing to the direction with the largest distribution difference, represents the gradient operation on the noise vector N, represents the KL divergence, Indicates the dynamic noise amplitude; is a random vector sampled from the unit sphere, providing an exploratory perturbation; is the dynamic weight, ranging from , represents L2 norm normalization. To balance the contribution of KL divergence guidance and randomness, represents L2 norm normalization, ensuring is a unit vector. KL divergence noise can be summarized as ( Can be used in conjunction with dynamic adjustment). , the noise direction is optimized to focus on the weak areas predicted by the model, while the random vector To ensure a certain degree of exploration, in order to avoid embedding distortion caused by excessive perturbation, a threshold is set. ,when When reduced , on the contrary, increase To maintain effectiveness.

[0041] Furthermore, the entire noise control direction is controlled by two parts, one is the time attenuation noise that changes over time, and the other is the noise addition in a specific direction optimized by the KL divergence. The final noise addition method is integrated as follows: ; in, represents the dynamic noise amplitude, is the optimal noise direction under the influence of KL divergence, is the regularization coefficient, adjusted experimentally (if the calculated N exceeds this limit, it is scaled to the boundary), Represents the original embedding The standard deviation of Represents the minimum function. The min method takes the minimum value to prevent excessive noise fluctuations from causing destructive interference to the vector representation itself, and sets a fixed threshold.

[0042] In step S5, the asymmetric encoding-decoding model is pre-trained using multiple datasets, and the parameters of the asymmetric encoding-decoding model are optimized through semantic reconstruction and noise adversarial joint loss.

[0043] Specifically, after completing the construction of the autoencoder and decoder and the noise addition part, the model was pre-trained using multiple datasets, and was fully trained on the MS MARCO (Microsoft Machine Reading Comprehension) dataset. The core task of the training is to enable the decoder to accurately restore the original sentence through semantic reconstruction. , and its loss function is defined as: ; in, For the mask sentence, are decoder parameters, is the sentence length, Indicates that the decoder is given a mask sentence and decoder parameters Under the condition of The words are probability.

[0044] This reconstruction task ensures that the model captures the deep semantic structure of the sentence. At the same time, the noise adversarial training task is performed by embedding Dynamic noise is applied to (As mentioned above ), optimizes the robustness of the model to disturbances, and its loss function is defined as: ; in, represents a masked sentence, represents the masked sentence after adding dynamic noise, measuring the consistency of embedding, represents the square of L2 norm, and Encoder(·) is the encoder function.

[0045] Furthermore, the overall training objective is the joint loss of semantic reconstruction and noise adversarial: ; in, For semantic reconstruction and noise adversarial joint loss, and are weight coefficients ranging from 0 to 1. Encoder parameters are optimized by gradient descent and decoder parameters In the evaluation phase, this task measures model performance through intensive retrieval tasks such as semantic similarity matching and information retrieval, using metrics including precision, recall, and mean inverse rank (MRR) to verify the generalization ability and robustness of the model under multi-dataset training.

[0046] In step S6, the pre-trained asymmetric encoding-decoding model is directly applied to dense retrieval tasks in zero-shot or supervised learning scenarios, and the retrieval task is performed through cosine similarity. The text is input into the pre-trained asymmetric encoding-decoding model to output the vector representation of the sentence, thereby achieving efficient deployment and application.

[0047] The performance of the method of the present invention on the MS_MARCO test data set is shown in Table 1 below.

[0048] Table 1 Comparison results between the method of the present invention and the traditional method

[0049] It can be seen from Table 1 that by enhancing the noise of the decoder, the method of the present invention (KL-Noise) has shown improvements in multiple indicators compared with traditional methods in the field of text representation.

[0050] Embodiment 2: like Figure 6 As shown, the present invention provides a model pre-training system based on masked autoencoder and noise enhancement, which is used to implement the model pre-training method based on masked autoencoder and noise enhancement in the above-mentioned embodiment 1, specifically comprising: An asymmetric encoding-decoding model construction module 100, for constructing an asymmetric encoding-decoding model, wherein the encoder is a full-scale deep neural network for generating an embedded representation of an input sentence, and the decoder is a lightweight single-layer neural network for reconstructing the input sentence; A differentiated masking strategy module 200 is used to apply a differentiated masking strategy to the input sentence, wherein the encoder masking ratio is 15%-30% and the decoder masking ratio is 50%-75% to generate a differentiated training signal; A dynamic noise injection module 300 is used to introduce a dynamic noise injection mechanism in the encoder embedding generation process, and enhance the robustness of the asymmetric encoding-decoding model against disturbances by adding noise of different amplitudes at different stages of the training process; A KL divergence guided noise generation module 400 is used to use KL divergence to guide noise generation and optimize the noise generation direction by calculating the distribution difference between the original embedding and the noisy embedding; A model pre-training and parameter optimization module 500 is used to pre-train an asymmetric encoding-decoding model using multiple data sets, and optimize the parameters of the asymmetric encoding-decoding model through a semantic reconstruction and noise adversarial joint loss; The model application module 600 is used to directly apply the pre-trained asymmetric encoding-decoding model to dense retrieval tasks in zero-shot or supervised learning scenarios, and output a vector representation of the sentence.

[0051] A model pre-training system based on a masked autoencoder and noise enhancement in this embodiment is used to implement the aforementioned model pre-training method based on a masked autoencoder and noise enhancement. Therefore, the specific implementation method of the model pre-training system based on a masked autoencoder and noise enhancement can be seen in the embodiment part of the model pre-training method based on a masked autoencoder and noise enhancement in the previous text. For example, an asymmetric encoding-decoding model construction module 100, a differentiated mask strategy module 200, a dynamic noise injection module 300, a KL divergence-guided noise generation module 400, a model pre-training and parameter optimization module 500, and a model application module 600 are respectively used to implement steps S1, S2, S3, S4, S5, and S6 in the aforementioned model pre-training method based on a masked autoencoder and noise enhancement. Therefore, its specific implementation method can refer to the description of the corresponding embodiments of each part. In order to avoid redundancy, it will not be repeated here.

[0052] Embodiment three: An embodiment of the present invention provides an electronic device, which includes a processor, a memory and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the above-mentioned model pre-training method based on mask autoencoder and noise enhancement.

[0053] Embodiment 4: An embodiment of the present invention provides a computer storage medium, which stores a computer software product. The computer software product includes a number of instructions for enabling a computer device to execute the above-mentioned model pre-training method based on masked autoencoder and noise enhancement.

[0054] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0055] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0056] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0057] Obviously, the above embodiments are merely examples for clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived from them are still within the protection scope of the invention.

Claims

1. A model pre-training method based on masked autoencoder and noise enhancement, characterized in that: include: S1: Construct an asymmetric encoder-decoder model, where the encoder is a full-scale deep neural network for generating an embedded representation of the input sentence, and the decoder is a lightweight single-layer neural network for reconstructing the input sentence; S2: A differentiated masking strategy is applied to the input sentence, where the encoder masking ratio is 15%-30% and the decoder masking ratio is 50%-75% to generate differentiated training signals; S3: Introducing a dynamic noise injection mechanism in the encoder embedding generation process, by adding noise of different amplitudes at different stages of the training process, the robustness of the asymmetric encoding-decoding model against perturbations is enhanced; S4: Use KL divergence to guide noise generation and optimize the noise generation direction by calculating the distribution difference between the original embedding and the noisy embedding; S5: Use multiple datasets to pre-train the asymmetric encoding-decoding model, and optimize the parameters of the asymmetric encoding-decoding model through semantic reconstruction and noise adversarial joint loss; S6: The pre-trained asymmetric encoder-decoder model is directly applied to dense retrieval tasks in zero-shot or supervised learning scenarios to output a vector representation of the sentence.

2. The model pre-training method based on masked autoencoder and noise enhancement according to claim 1, characterized in that: In the asymmetric encoding-decoding model, the encoder adopts a multi-layer Transformer architecture to output a fixed-dimensional embedding vector, and the decoder adopts a single-layer Transformer architecture, the number of its parameters does not exceed 1 / 10 of the number of parameters of the encoder, and the parameters of the encoder and decoder are initialized through a pre-trained language model.

3. The model pre-training method based on masked autoencoder and noise enhancement according to claim 1, characterized in that: Step S2 specifically includes: For the encoder mask operation, first determine the mask ratio , the range is , calculate the number of words that need to be masked , then from the sentence position Randomly select positions, denoted as a set ;for Each position in , the word Replace with special token [MASK] to generate masked sentences , which is in the form of , or , ; This mask sentence As encoder input, used to generate embeddings and the masked word vector ; In contrast, the decoder masking operation uses a higher masking ratio , ranging from 50%-75%, calculate the number of masked words , and from Randomly select positions, denoted as a set , May contain Similarly, Words in Replace with special marker , generate masked sentences , , or , , and as decoder input and position encoding Together they are used to reconstruct the original sentence ; in, represents the floor function, represents the length of the sentence, Encoder(·) is the encoder function, Decoder function.

4. The model pre-training method based on masked autoencoder and noise enhancement according to claim 1, characterized in that: In the dynamic noise injection mechanism, the noise amplitude is dynamically adjusted according to the training progress, and its calculation formula is: ; in, represents the dynamic noise amplitude, represents the initial noise amplitude, Represents the attenuation rate, which is used to control the speed of amplitude decrease. represents the total number of training steps, Indicates the current training step number, Indicates the minimum amplitude of noise.

5. The model pre-training method based on masked autoencoder and noise enhancement according to claim 1, characterized in that: In the KL divergence-guided noise generation, the noise direction 𝜂 is calculated by the following formula: ; in, represents the gradient under Gaussian approximation, pointing to the direction with the largest distribution difference, represents the gradient operation on the noise vector N, represents the KL divergence, Indicates the dynamic noise amplitude; is used to measure the difference between two embedding distributions. and The original embedding and noisy embedding The probability distribution of is a random vector sampled from the unit sphere, providing an exploratory perturbation; is the dynamic weight, ranging from , Denotes L2 norm normalization.

6. The model pre-training method based on masked autoencoder and noise enhancement according to claim 5, characterized in that: The noise embedding It is expressed as: ; in, represents the dynamic noise amplitude, is the optimal noise direction under the influence of KL divergence, is the regularization coefficient, adjusted through experiments, Represents the original embedding The standard deviation of Represents the minimum function.

7. The model pre-training method based on masked autoencoder and noise enhancement according to claim 1, characterized in that: The semantic reconstruction and noise adversarial joint loss is: ; in, For semantic reconstruction and noise adversarial joint loss, and are weight coefficients, ranging from 0 to 1. is the semantic reconstruction loss, For noise combating losses.

8. The model pre-training method based on masked autoencoder and noise enhancement according to claim 7, characterized in that: The semantic reconstruction loss is: ; in, For the mask sentence, are decoder parameters, is the sentence length, Indicates that the decoder is given a mask sentence and decoder parameters Under the condition of The words are probability.

9. The model pre-training method based on masked autoencoder and noise enhancement according to claim 7, characterized in that: The noise combat loss is: ; in, represents a masked sentence, represents the masked sentence after adding dynamic noise, represents the square of L2 norm, and Encoder(·) is the encoder function.

10. A model pre-training system based on masked autoencoder and noise enhancement, characterized in that: The system is used to implement the model pre-training method based on masked autoencoder and noise enhancement according to any one of claims 1 to 9, specifically comprising: Asymmetric encoding-decoding model building module, used to build an asymmetric encoding-decoding model, where the encoder is a full-scale deep neural network for generating an embedded representation of the input sentence, and the decoder is a lightweight single-layer neural network for reconstructing the input sentence; A differentiated masking strategy module is used to apply a differentiated masking strategy to the input sentence, where the encoder masking ratio is 15%-30% and the decoder masking ratio is 50%-75% to generate differentiated training signals; Dynamic noise injection module, which is used to introduce a dynamic noise injection mechanism in the encoder embedding generation process. By adding noise of different amplitudes at different stages of the training process, the robustness of the asymmetric encoding-decoding model against perturbations is enhanced; The KL divergence-guided noise generation module is used to use the KL divergence to guide the noise generation and optimize the noise generation direction by calculating the distribution difference between the original embedding and the noisy embedding; Model pre-training and parameter optimization module, which is used to pre-train the asymmetric encoding-decoding model using multiple datasets and optimize the parameters of the asymmetric encoding-decoding model through semantic reconstruction and noise adversarial joint loss; The model application module is used to directly apply the pre-trained asymmetric encoding-decoding model to dense retrieval tasks in zero-shot or supervised learning scenarios, and output the vector representation of the sentence.

11. An electronic device, characterized in that: The electronic device includes a processor, a memory and a bus system, the processor and the memory are connected through the bus system, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the model pre-training method based on masked autoencoder and noise enhancement as described in any one of claims 1 to 9.

12. A computer storage medium, characterized in that: The computer storage medium stores a computer software product, which includes a number of instructions for enabling a computer device to execute the model pre-training method based on masked autoencoder and noise enhancement as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Dynamic mask training method in Chinese automatic grammar error correction

    CN111062205A

  • Text generation method based on diffusion language model

    CN117610509A

  • Semantic vector model pre-training method based on multi-mask mode

    CN117952151A

  • Neural machine translation method based on word order enhancement reconstruction

    CN117973401A

  • Semantic retrieval method and device based on natural language representation learning and terminal equipment

    CN118093783A

Cited By

  • Artificial intelligence operation and maintenance method and system based on large model

    CN120494809A

  • Artificial intelligence operation and maintenance method and system based on large model

    CN120494809B

  • Plate heat exchanger surface defect detection method

    CN121213569A

  • A plate heat exchanger surface defect detection method

    CN121213569B

  • Robot action control method and device and electronic equipment

    CN121670689A