Word embedding model training method and device based on unlabeled data, equipment and storage medium
By building a vector database and similarity filtering, using unlabeled data to train the word embedding model, the problem of traditional models dependence on labeled data is solved, and the performance and generalization capabilities of the model are improved.
Patent Information
- Application Number
- CN202510101642.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional word embedding model training relies on a large number of labeled corpus. It is expensive to obtain labeled data and is difficult to scale on a large scale. There is a lack of clear positive and negative samples in the unlabeled data, making it difficult to identify challenging negative samples, which limits the improvement of model performance.
By obtaining the initial data and the initial word embedding model, the initial word embedding model is used to convert the initial data into vector representation, a vector database is built, similarity filtering is performed, the target sample data is filtered, the initial word embedding model is trained, and the loss function is optimized to improve the model performance.
It significantly improves the training effect and generalization ability of the word embedding model, reduces dependence on manual labeled data, makes full use of the potential semantic information of unlabeled data, and enhances the model's learning ability.
Smart Images

Figure CN120067293A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a method, apparatus, device and storage medium for training a word embedding model based on unlabeled data. Background Art
[0002] With the rapid development of natural language processing technology, word embedding, as a technology for converting words into vectors, has become the basis for computers to understand and process text data. High-quality word embedding models can effectively capture the semantic relationship between words, which is crucial to improving the performance of natural language processing systems. However, the training of traditional word embedding models usually relies on a large number of annotated corpora. The acquisition of these annotated data not only requires a lot of manual intervention, but is also costly and difficult to scale up on a large scale. Therefore, developing a method that can use unlabeled data to train word embedding models to reduce dependence on labeled data has become an important demand in the field of natural language processing.
[0003] Currently, researchers have tried to train word embedding models through unlabeled data. The main methods include similarity-based matching and adversarial training based on generative models. These methods attempt to use the potential structure of the data itself for learning, thereby reducing dependence on labeled data. For example, predicting the target word through contextual information, or learning the semantic information of words from unlabeled data through masked language models and next sentence prediction tasks.
[0004] Although these methods based on unlabeled data have made some progress, they still have some problems. First, the lack of clear positive and negative samples in unlabeled data makes it difficult for the model to effectively identify challenging negative samples. Second, the existing negative sample selection strategies are often too simple, making it difficult to mine hard negative examples that are truly helpful for model training, thus limiting the improvement of model performance. In addition, existing methods are insufficient in dealing with complex semantic relationships and are difficult to adapt to diverse text data. Therefore, how to improve the training quality of word embedding models has become an urgent problem to be solved. Summary of the invention
[0005] The purpose of this application is to provide a word embedding model training method, device, equipment and storage medium based on unlabeled data, aiming to solve the technical problem of how to improve the training quality of the word embedding model.
[0006] To achieve the above objectives, the present application proposes a word embedding model training method based on unlabeled data, the method comprising:
[0007] Get initial data and initial word embedding model;
[0008] Perform vector representation on the initial data according to the initial word embedding model to obtain a vector database;
[0009] Perform similarity screening based on the vector database to obtain target sample data;
[0010] Train the initial word embedding model with the target sample data to obtain a target word embedding model.
[0011] In one embodiment, performing vector representation on the initial data according to the initial word embedding model to obtain a vector database includes:
[0012] Obtain answer data according to the initial data;
[0013] Perform vector representation on the answer data through the initial word embedding model to obtain an answer vector;
[0014] Based on the answer vector, obtain a vector database.
[0015] In one embodiment, performing similarity screening based on the vector database to obtain target sample data includes:
[0016] Obtain preset question data;
[0017] Perform vector representation on the preset question data according to the initial word embedding model to obtain a question vector;
[0018] Perform cosine similarity calculation and screening based on the vector database and the question vector to obtain initial sample data;
[0019] Screen the initial sample data to obtain target sample data.
[0020] In one embodiment, screening the initial sample data to obtain target sample data includes:
[0021] Obtain a preset re-ranking model;
[0022] Perform correlation evaluation on the initial sample data through the preset re-ranking model to obtain a sample evaluation table, where the sample evaluation table includes the mapping relationship between the initial sample data and the correlation score;
[0023] Query the sample evaluation table and perform screening based on the correlation score to obtain target sample data.
[0024] In one embodiment, performing vector representation on the preset question data according to the initial word embedding model to obtain a question vector includes:
[0025] Perform word segmentation on the preset question data to obtain token data;
[0026] Based on the token data and the initial word embedding model, obtain word vectors;
[0027] Perform average pooling on the word vectors to obtain question vectors.
[0028] In one embodiment, training the initial word embedding model with the target sample data to obtain a target word embedding model includes:
[0029] Obtain a preset similarity threshold and the weight coefficients of the initial word embedding model;
[0030] Based on the target sample data and the initial word embedding model, obtain an initial positive example similarity value and an initial negative example similarity set;
[0031] Based on the initial positive example similarity value, the initial negative example similarity set, and the preset similarity threshold, adjust the weight coefficients until the initial word embedding model converges to obtain a target word embedding model.
[0032] In one embodiment, based on the initial positive example similarity value, the initial negative example similarity set, and the preset similarity threshold, adjusting the weight coefficients until the initial word embedding model converges to obtain a target word embedding model includes:
[0033] Obtain an initial similarity value based on the initial positive example similarity value and the initial negative example similarity set;
[0034] If the initial similarity value is greater than the preset similarity threshold, adjust the weight coefficients to obtain a target word embedding model.
[0035] In addition, to achieve the above object, the present application also proposes a word embedding model training device based on unlabeled data, and the device includes:
[0036] An acquisition module, configured to acquire initial data and an initial word embedding model;
[0037] A obtaining module, configured to perform vector representation on the initial data according to the initial word embedding model to obtain a vector database;
[0038] A screening module, configured to perform similarity screening based on the vector database to obtain target sample data;
[0039] A completion module, configured to train the initial word embedding model with the target sample data to obtain a target word embedding model.
[0040] In addition, to achieve the above object, the present application further provides a training device for a word embedding model based on unlabeled data, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the method for training a word embedding model based on unlabeled data as described above.
[0041] In addition, to achieve the above object, the present application further provides a storage medium, the storage medium being a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method for training a word embedding model based on unlabeled data as described above.
[0042] In addition, to achieve the above object, the present application further provides a computer program product, the computer program product comprising a computer program, and when the computer program is executed by a processor, it implements the steps of the method for training a word embedding model based on unlabeled data as described above.
[0043] One or more technical solutions proposed by the present application have at least the following technical effects:
[0044] The present application first obtains initial data and an initial word embedding model, providing a basis for subsequent vector representation and training. Then, the initial word embedding model is used to convert the initial data into vector representations, and a vector database is constructed, realizing efficient storage and fast retrieval of data, and providing support for subsequent similarity screening. Then, similarity screening is performed based on the vector database. By calculating the similarity between vectors, candidate data closest to the target sample is screened out, which can effectively identify challenging negative samples and provide more valuable training signals for model training. Finally, the initial word embedding model is trained using the screened target sample data. By optimizing the loss function of the model, the model can better distinguish between positive and negative samples, thereby improving the performance of the word embedding model. The present application not only makes full use of the potential semantic information of unlabeled data, but also enhances the learning ability of the model by screening difficult negative examples, and finally realizes significant improvement in the training effect and generalization ability of the word embedding model while reducing the dependence on manual annotation. Description of the Drawings
[0045] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 This is a schematic flowchart provided for the first embodiment of the method for training a word embedding model based on unlabeled data in this application;
[0048] Figure 2 This is a schematic flowchart provided for the second embodiment of the method for training a word embedding model based on unlabeled data in this application;
[0049] Figure 3 This is a schematic flowchart for the brief method of training a word embedding model based on unlabeled data in the embodiments of this application;
[0050] Figure 4 This is a schematic diagram of the module structure of the device for training a word embedding model based on unlabeled data in the embodiments of this application;
[0051] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the method for training a word embedding model based on unlabeled data in the embodiments of this application.
[0052] The implementation, functional features, and advantages of the purpose of this application will be further described in combination with the embodiments and with reference to the accompanying drawings. Detailed implementation manners
[0053] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0054] In order to better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.
[0055] With the rapid development of natural language processing technology, word embedding, as a technique for converting words into vectors, has become the basis for computers to understand and process text data. High-quality word embedding models can effectively capture the semantic relationships between words, which is crucial for improving the performance of natural language processing systems. However, the training of traditional word embedding models usually relies on a large amount of labeled corpora. The acquisition of these labeled data not only requires a large amount of manual intervention but also is costly and difficult to scale up. Therefore, developing a method to train word embedding models using unlabeled data to reduce the dependence on labeled data has become an important requirement in the field of natural language processing. Currently, researchers have tried to train word embedding models using unlabeled data, and the main methods include similarity-based matching and adversarial training based on generative models. These methods attempt to learn using the potential structure of the data itself to reduce the dependence on labeled data. For example, predicting the target word through context information, or learning the semantic information of words from unlabeled data through the Masked Language Model and Next Sentence Prediction tasks. Although these methods based on unlabeled data have made some progress, they still have some problems. First, there are no clear positive and negative samples in unlabeled data, making it difficult for the model to effectively identify challenging negative samples. Second, the existing negative sample selection strategies are often too simple to mine truly helpful hard negative examples for model training, thus limiting the improvement of model performance. In addition, the existing methods perform poorly in dealing with complex semantic relationships and are difficult to adapt to diverse text data.
[0056] The main solution of the embodiments of this application is as follows: The embodiments of this application first obtain initial data and an initial word embedding model to provide a basis for subsequent vector representation and training. Then, the initial word embedding model is used to convert the initial data into vector representations, and a vector database is constructed to achieve efficient storage and fast retrieval of data, providing support for subsequent similarity screening. Then, similarity screening is performed based on the vector database. By calculating the similarity between vectors, the candidate data closest to the target sample is screened out, which can effectively identify challenging negative samples and provide more valuable training signals for model training. Finally, the initial word embedding model is trained using the screened target sample data. By optimizing the loss function of the model, the model can better distinguish positive and negative samples, thereby improving the performance of the word embedding model. The embodiments of this application not only make full use of the potential semantic information of unlabeled data but also enhance the learning ability of the model by screening hard negative examples, ultimately achieving a significant improvement in the training effect and generalization ability of the word embedding model while reducing the dependence on manual annotation.
[0057] It should be noted that the execution entity of the embodiments of the present application can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a computer, etc. that can implement the above functions. Hereinafter, a computer will be taken as an example to illustrate this embodiment and the following embodiments.
[0058] Based on this, the embodiments of the present application provide a method for training a word embedding model based on unlabeled data. Referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the method for training a word embedding model based on unlabeled data of the present application.
[0059] In this embodiment, the method for training a word embedding model based on unlabeled data includes steps S10 to S40:
[0060] Step S10, obtain initial data and an initial word embedding model;
[0061] It should be noted that the initial data can be unlabeled text data for training the word embedding model. These data can be large-scale text corpora, such as news articles, web content, or social media text, etc. These data usually exist in text form and need to be preprocessed (such as word segmentation, cleaning, etc.) to adapt to the input requirements of the word embedding model. The initial word embedding model can be a predefined or pre-trained word embedding model for mapping words or phrases in text data into vector representations. The initial word embedding model can include Word2Vec (Word to Vector), GloVe (Global Vectors for Word Representation), or BERT (Bidirectional Encoder Representations from Transformers), etc. The initial data provides training materials for the model, while the initial word embedding model provides a basic framework for the vector representation of the data. The combination of the two lays a foundation for the subsequent construction of the vector database and model optimization.
[0062] It is understandable that the initial data refers to the unlabeled text data used for training, which can be obtained from Q&A pairs, news articles, social media content, or other large-scale text corpora on the Internet. They provide a rich source of semantic information for model training. The initial word embedding model, on the other hand, is a pre-defined or pre-trained model, such as Word2Vec, GloVe, or BERT, etc., which is used to convert words or phrases in the text data into vector representations. By obtaining these two parts, it lays a foundation for the subsequent construction of the vector database and model optimization, enabling the model to learn from the unlabeled data and improve the quality of word embeddings.
[0063] Step S20: Perform vector representation on the initial data according to the initial word embedding model to obtain a vector database.
[0064] It should be noted that the vector database can be an efficient data structure or system for storing and managing the vector representations obtained by converting the initial data through the word embedding model. Specifically, the functions of the vector database are as follows: storing vector representations, converting each text segment in the initial data (such as the answer in the Q&A pair) into a vector form through the initial word embedding model and storing these vectors for subsequent query and use; supporting fast similarity queries, the vector database can efficiently calculate and retrieve other vectors that are most similar to a given vector, usually achieved through cosine similarity or other distance measurement methods, so as to quickly find the candidate samples closest to the target vector; optimizing training efficiency, through the vector database, the model can quickly locate useful negative samples in the large-scale data, thereby improving the training efficiency and effect. The vector database can be built based on modern database management systems (such as Milvus, Faiss, etc.), which are specifically designed to handle high-dimensional vector data and support efficient similarity search and data management.
[0065] It is understandable that by using a pre-defined or pre-trained initial word embedding model, each text segment in the initial data (such as the answer in the Q&A pair) is converted into a high-dimensional vector form. These vectors can capture the semantic information of the text and are stored in a dedicated vector database. The vector database supports efficient similarity queries and can quickly retrieve other vectors that are most similar to the input vector, thereby providing a data basis and computational support for subsequent negative sample selection and model optimization.
[0066] As an example, performing vector representation on the initial data according to the initial word embedding model to obtain a vector database includes: obtaining answer data from the initial data; performing vector representation on the answer data through the initial word embedding model to obtain answer vectors; and obtaining a vector database based on the answer vectors.
[0067] Among them, the answer data can be text fragments extracted from the initial data. These fragments are usually associated with the questions and are used for subsequent word embedding training and similarity calculation. For example, if the initial data is a question-answer pair, then the "answer data" is the answer part in the question-answer pair. These answer data are text contents after preprocessing (such as word segmentation, cleaning, etc.) and are used to generate vector representations. The answer vectors can be vector representations obtained by converting the "answer data" through the initial word embedding model. Specifically, the initial word embedding model (such as Word2Vec, GloVe, or BERT) will map each word or phrase in the answer data into a high-dimensional vector, and these vectors can capture the semantic information of the answer. The answer vectors are the basis for subsequent similarity calculation and negative sample selection, and they are stored in the vector database for fast retrieval and matching.
[0068] Specifically, extract the answer text fragments related to the question from the initial data as the answer data. Then, use the initial word embedding model to convert these answer data into high-dimensional vector representations, which can capture the semantic information of the answer, and are called answer vectors. Finally, store all the answer vectors in a dedicated vector database for subsequent efficient similarity query and negative sample selection to support the optimization of the word embedding model.
[0069] Step S30: Based on the vector database, perform similarity screening to obtain target sample data;
[0070] It should be noted that the target sample data can be a sample set for training the word embedding model obtained through similarity screening in the vector database. Specifically, these samples are screened from the answer vectors of the initial data by calculating the similarity (such as cosine similarity or Euclidean distance) with the target vector (such as the question vector). These sample data usually include several vectors that are most similar to the target vector, and they are used as negative samples in subsequent training to help the model better learn semantic information.
[0071] It can be understood that through the answer vectors in the vector database, calculate their similarity with the current input vector (such as the question vector). According to the similarity level, screen out several vectors that are most similar to the input vector as the target sample data. These target sample data can include positive samples and hard negative examples, which are used for subsequent model training to help improve the performance and generalization ability of the word embedding model.
[0072] Step S40: Train the initial word embedding model with the target sample data to obtain a target word embedding model.
[0073] It should be noted that the target word embedding model can be the final word embedding model obtained after optimized training. It can be obtained by training and adjusting the initial word embedding model using target sample data (including positive samples and hard negative examples). The target word embedding model has better semantic representation ability and generalization performance, can capture the semantic relationships between words more accurately, and shows better performance on unlabeled data.
[0074] It can be understood that the initial word embedding model is optimized and trained using the target sample data obtained through similarity screening. By adjusting the parameters of the model, the similarity of positive samples is maximized, while the similarity of hard negative examples is minimized, thereby enhancing the model's ability to distinguish semantic information. After this training process, the initial word embedding model is improved and optimized, and finally a target word embedding model with higher performance and better generalization ability is formed, which can capture the semantic relationships between words more accurately and is suitable for subsequent natural language processing tasks.
[0075] As an example, training the initial word embedding model with the target sample data to obtain a target word embedding model includes: obtaining a preset similarity threshold and the weight coefficients of the initial word embedding model; obtaining an initial positive example similarity value and an initial negative example similarity set based on the target sample data and the initial word embedding model; adjusting the weight coefficients based on the initial positive example similarity value, the initial negative example similarity set, and the preset similarity threshold until the initial word embedding model converges to obtain the target word embedding model.
[0076] Among them, the weight coefficients can be numerical values used to adjust the model parameters during the training process of the word embedding model. These coefficients determine the importance of different features or parameters in the model. By optimizing these weight coefficients, the model can better learn the semantic information of the input data. During the training process, the weight coefficients are dynamically adjusted according to the feedback of the loss function to optimize the performance of the model. The preset similarity threshold can be a preset numerical value used to determine whether the similarity between samples reaches a certain standard during the training process. Specifically, it is used to distinguish the similarity boundary between positive samples and negative samples. For example, if the similarity difference between two vectors is higher than this threshold, the weight value is adjusted so that the similarity difference is lower than this threshold, thereby maximizing the similarity value of positive examples and minimizing the similarity value of negative examples. The initial positive example similarity value can be the similarity of positive samples calculated by the initial word embedding model during the training process. Positive samples are usually known text pairs that are semantically similar (such as questions and answers in a question-and-answer pair). The initial positive example similarity value reflects the model's estimation of the semantic similarity of positive samples at the initial stage of training. Through optimized training, the model will try to increase the similarity value of these positive examples to make it closer to 1 (or higher). The initial negative example similarity set can be a set of similarity values of all negative samples calculated by the initial word embedding model during the training process. Negative samples can be text pairs that are similar to positive samples in certain features (such as hard negative examples). The initial negative example similarity set reflects the model's estimation of the semantic similarity of negative samples at the initial stage of training. Through optimized training, the model will try to reduce the similarity value of these negative examples to make it closer to 0 (or lower).
[0077] Specifically, first obtain the preset similarity threshold and the weight coefficients of the initial word embedding model. Then, use the target sample data and the initial word embedding model to calculate the initial positive example similarity value between positive samples and the initial negative example similarity set of negative samples. Based on these similarity values and the preset similarity threshold, optimize the loss function of the model by adjusting the weight coefficients, so that the similarity value of positive examples is as close as possible to 1 (or higher), while the similarity value of negative examples is as close as possible to 0 (or lower). This process is iterated until the initial word embedding model converges, and finally an optimized target word embedding model is obtained, thereby improving the semantic discrimination ability and generalization performance of the model.
[0078] As an example, adjusting the weight coefficients based on the initial positive example similarity value, the initial negative example similarity set, and the preset similarity threshold until the initial word embedding model converges to obtain a target word embedding model includes: obtaining an initial similarity value based on the initial positive example similarity value and the initial negative example similarity set; if the initial similarity value is greater than the preset similarity threshold, adjust the weight coefficients to obtain a target word embedding model.
[0079] Among them, the initial similarity value can be the difference between the initial positive example similarity value calculated by the initial word embedding model and the maximum value in the set of initial negative example similarities. Specifically, it can be a quantitative indicator used to measure the semantic discrimination between positive and negative samples. The initial positive example similarity value reflects the semantic similarity between positive sample pairs, and the maximum value in the set of initial negative example similarities reflects the semantic similarity between the closest negative sample and the positive sample, and usually, it is expected that its value is as low as possible. By calculating the difference between the two, the initial similarity value quantifies the semantic discrimination between positive and negative samples. If the initial similarity value is large, it indicates that the semantic discrimination between positive and negative samples is relatively obvious; if the initial similarity value is small, it indicates that the model's ability to distinguish between positive and negative samples is weak. During the training process, by comparing the initial similarity value with the preset similarity threshold, the model can determine whether the current semantic discrimination ability meets the requirements. If the initial similarity value is greater than the preset similarity threshold, it means that the current model's ability to distinguish between positive and negative samples is good. At this time, the model can be further optimized by adjusting the weight coefficients until the model converges, and finally, the target word embedding model is obtained.
[0080] Specifically, first, the initial similarity value is obtained by calculating the difference between the initial positive example similarity value and the maximum value in the set of initial negative example similarities. This initial similarity value reflects the current model's semantic discrimination ability between positive and negative samples. If the initial similarity value is greater than the preset similarity threshold, it means that the current model's ability to distinguish between positive and negative samples is already good, but there is still room for optimization. Therefore, the model will adjust the weight coefficients according to this condition and further optimize the model parameters until the model converges, and finally, a target word embedding model with better performance is obtained, ensuring that the model can continuously improve its ability to capture semantic information during training, so as to better process unlabeled data.
[0081] This embodiment provides a method for training a word embedding model based on unlabeled data. First, this embodiment obtains initial data and an initial word embedding model, providing a basis for subsequent vector representation and training. Then, the initial word embedding model is used to convert the initial data into a vector representation, and a vector database is constructed, achieving efficient storage and fast retrieval of data and providing support for subsequent similarity screening. Next, similarity screening is performed based on the vector database. By calculating the similarity between vectors, candidate data closest to the target sample is screened out, which can effectively identify challenging negative samples and provide more valuable training signals for model training. Finally, the initial word embedding model is trained using the screened target sample data. By optimizing the loss function of the model, the model can better distinguish between positive and negative samples, thereby improving the performance of the word embedding model. This embodiment not only makes full use of the potential semantic information of unlabeled data but also enhances the learning ability of the model by screening difficult negative examples, ultimately achieving a significant improvement in the training effect and generalization ability of the word embedding model while reducing the dependence on manual annotation.
[0082] Based on the first embodiment of the present application, in the second embodiment of the present application, for the same or similar content as in the above-mentioned first embodiment, reference can be made to the above introduction and will not be elaborated hereinafter. On this basis, please refer to Figure 2 , Figure 2 which is a schematic flowchart of the second embodiment of the method for training a word embedding model based on unlabeled data of the present application. Step S30 of the method for training a word embedding model based on unlabeled data includes steps S31 to S34:
[0083] Step S31, obtain preset problem data;
[0084] It should be noted that the preset question data can be a set of pre - defined or prepared question texts, which are the key inputs used to guide the model for similarity query and training. They are usually extracted from the initial data or designed according to specific task requirements to simulate the query needs in actual application scenarios. The preset question data can be used to guide the similarity query. As the starting point of the query, by calculating the similarity between it and the answer vectors stored in the vector database, the model can find the answer or text segment most relevant to the question. These answers or segments serve as candidate negative samples for subsequent training. The preset question data can also be used to provide training signals. The preset question data and their corresponding answers (positive samples) provide clear semantic matching targets for model training. By comparing the similarity between the question vector and the answer vector, the model can learn which text segments are semantically relevant and which are not. The preset question data can also support the mining of hard negative examples. The preset question data is not only used to find positive samples but also to mine hard negative examples (i.e., text segments that are semantically similar to the question but are not the correct answers). These hard negative examples are the key to improving the model's performance because they can enhance the model's ability to distinguish complex semantic relationships. The preset question data can be extracted from the initial data. For example, the question part can be extracted from the question - answer pair data, or key questions can be extracted from text data such as news titles and conversation histories. The preset question data can also be designed manually. According to specific application scenarios or training objectives, some representative and challenging questions can be designed to guide the model to learn specific semantic patterns.
[0085] It can be understood that a set of representative and targeted question texts are extracted from the initial data or pre - designed as the input for model training and similarity query. These preset question data are used to guide the model to find the answer vector most relevant to them in the vector database, helping the model identify positive samples and mine hard negative examples, so as to provide clear semantic matching signals for the subsequent training of the word - embedding model and enhance the model's learning and discrimination ability of semantic relationships.
[0086] Step S32, represent the preset question data in vector form according to the initial word - embedding model to obtain question vectors;
[0087] It should be noted that the question vector can be the result of converting the preset question data into a high - dimensional vector form through the initial word - embedding model. Each question text in the preset question data is processed by the word - embedding model and mapped to a vector that can represent its semantic information. This vector captures the core semantic features of the question, enabling the model to judge the semantic relevance between different texts by calculating the similarity between vectors. The question vector is the basis for subsequent similarity query and model training, and is used to compare with the answer vectors in the vector database to find the answer most relevant to the question or mine hard negative examples.
[0088] It is understandable that each question text in the preset question data is converted into a high-dimensional vector form by using an initial word embedding model (such as Word2Vec, GloVe, or BERT, etc.). These vectors can capture the core semantic features of the questions, map the text information into a mathematical space, enabling the model to quantify the semantic relevance between different texts by calculating the similarity between vectors. The question vectors are the basis for subsequent similarity queries and model training, used to compare with the answer vectors in the vector database, thereby achieving the matching of positive samples and the mining of hard negative examples, providing key semantic information for the optimization of the word embedding model.
[0089] As an example, performing vector representation on the preset question data according to the initial word embedding model to obtain question vectors includes: performing word segmentation processing on the preset question data to obtain token data; obtaining word vectors based on the token data and the initial word embedding model; and obtaining question vectors through average pooling processing based on the word vectors.
[0090] Among them, the token data can be a set of words or sub-word units obtained by performing word segmentation processing on the preset question data. Word segmentation is an important step in natural language processing, which divides a continuous text string into individual tokens, and these tokens can be words, phrases, or character sequences. For example, for the sentence "Natural language processing is very important", the possible token data obtained after word segmentation may be "natural", "language", "processing", "very", "important". The word vectors can be the results of converting each token into a high-dimensional vector representation through the initial word embedding model. The word vectors can capture the semantic and syntactic features of the tokens, enabling the model to understand the relationships between tokens through mathematical operations between vectors. For example, through the model, the word "natural" may be converted into a vector with a length of 300, and this vector can reflect the position of the word "natural" in the semantic space. When processing the preset question data, first perform word segmentation on the question text to obtain token data. Then use the initial word embedding model to convert each token into the corresponding word vector. These word vectors can be further combined or aggregated, such as through average pooling processing, to obtain the vector representation of the entire question, that is, the question vector. The question vector can capture the semantic information of the entire question and is used for subsequent similarity calculations and model training.
[0091] Specifically, first, perform word segmentation on the preset question data, splitting it into individual tokens to obtain token data. Then, use the initial word embedding model to map each token to a high-dimensional word vector, which can capture the semantic information of the token. Finally, through average pooling, these word vectors are combined into a vector that can represent the semantics of the entire question, namely the question vector. This process realizes the conversion from text to vector, enabling the model to process and understand the semantics of the question through the calculation of the similarity between vectors.
[0092] Step S33: Calculate the cosine similarity based on the vector database and the question vector and perform screening to obtain initial sample data.
[0093] It should be noted that the initial sample data can be a set of vectors and their corresponding text data that are most similar to the question vector screened from the vector database through cosine similarity calculation. These vectors are sorted from high to low according to the similarity scores after calculating the similarity with the question vector in the vector database, and the top k (top-k) most similar vectors and their corresponding text segments are selected. These text segments are usually answers or text content related to the semantics of the question, but have not been further screened and verified, so they are called initial sample data. The initial sample data can provide candidate samples. As a candidate sample set, the initial sample data provides a basis for subsequent negative sample selection and model training. These samples may include positive samples (answers highly relevant to the semantics of the question) and negative samples (text segments with lower semantic relevance). The initial sample data can also support hard negative example mining. By further analysis and sorting, hard negative examples (i.e., text segments similar to the semantics of the question but not the correct answer) can be screened out from the initial sample data. These hard negative examples are crucial for improving the semantic discrimination ability of the model. The screening process of the initial sample data reduces the amount of data to be processed, making subsequent model training more efficient and ensuring the quality of training samples.
[0094] It can be understood that by calculating the cosine similarity between the question vector and all answer vectors stored in the vector database, the semantic relevance between them is quantified. The results are sorted from high to low according to the similarity scores, and the top k vectors with the highest scores and their corresponding text segments are selected as the initial sample data. These initial sample data, as a candidate set, provide a basis for further negative sample screening and model training, helping the model better learn the semantic discrimination ability.
[0095] Step S34: Screen the initial sample data to obtain target sample data.
[0096] It should be noted that the target sample data can be a set of high-quality samples obtained by further screening and evaluating the initial sample data. These samples are finally determined to be the positive samples and hard negative examples that are most valuable for model training through evaluation. The target sample data is highly semantically related to the question vector and is the positive sample in model training, which is used to strengthen the model's learning of correct semantic relationships. By screening out the most representative positive samples and hard negative examples, the target sample data helps the model learn more efficiently, improve its generalization ability and semantic discrimination ability, and is used to optimize the performance of the word embedding model. By using these screened samples, the model can better learn semantic information, distinguish the differences between positive and negative samples, and thus show better training effects on unlabeled data.
[0097] It can be understood that by using the re-ranking model to evaluate each sample in the initial sample data, the positive samples and hard negative examples that are most semantically related to the question vector and most representative are screened out. This process further optimizes the quality of the samples, ensuring that the target sample data can provide more effective semantic information for the training of the word embedding model, thereby improving the model's ability to distinguish between positive and negative samples and its overall performance.
[0098] As an example, screening the initial sample data to obtain the target sample data includes: obtaining a preset re-ranking model; performing a relevance evaluation on the initial sample data through the preset re-ranking model to obtain a sample evaluation table, where the sample evaluation table includes the mapping relationship between the initial sample data and the relevance score; and screening according to querying the sample evaluation table and based on the relevance score to obtain the target sample data.
[0099] Among them, the preset re-ranking model can be a model used to perform semantic relevance evaluation and ranking on the initial sample data. It can be based on deep learning techniques (such as cross-encoders), which can output similarity scores for query and document pairs and re-rank the samples according to these scores. Its purpose is to perform more precise screening on candidate samples according to semantic relevance, thereby improving the effect of model training. The sample evaluation table can be a data structure used to store the initial sample data and its corresponding semantic relevance scores. It is a mapping relationship table, where each sample corresponds to a relevance score, which is used to represent the semantic matching degree between the sample and the query question. The mapping relationship can be the process of establishing a corresponding relationship between the elements in one set and the elements in another set. The mapping relationship is specifically manifested as the corresponding relationship between the initial sample data and the relevance score in the sample evaluation table. This mapping is achieved through the evaluation results of the re-ranking model and is used to quickly query and screen out the most representative samples.
[0100] Specifically, first, a pre-set re-ranking model is obtained. This model can evaluate the semantic relevance of each sample in the initial sample data and output a relevance score. These scores reflect the semantic matching degree between the sample and the question vector and are stored in the sample evaluation table, forming a mapping relationship between the initial sample data and the relevance scores. Subsequently, according to the relevance scores in the sample evaluation table, the most representative samples, including positive samples and hard negative examples, are selected, and finally, the target sample data is obtained. Through precise evaluation and screening, it is ensured that the target sample data can provide high-quality semantic information for the training of the word embedding model, thereby improving the performance and generalization ability of the model.
[0101] In this embodiment, first, a set of representative and targeted question texts is extracted from the initial data or pre-designed as the input of the model. These pre-set question data provide clear semantic goals for subsequent similarity queries and training. Then, the initial word embedding model is used to convert the pre-set question data into a high-dimensional vector form to obtain the question vector. This process maps the text information into a mathematical space, enabling the model to quantify the semantic relevance between different texts by calculating the similarity between vectors. Next, the cosine similarity between the question vector and the answer vectors stored in the vector database is calculated to quantify their semantic relevance. The results are sorted from high to low according to the similarity scores, and the top k vectors with the highest scores and their corresponding text fragments are selected as the initial sample data. These initial sample data serve as a candidate set, providing a basis for further screening and verification. Finally, the re-ranking model evaluates the semantic relevance of each sample in the initial sample data, and the positive samples and hard negative examples that are most semantically relevant and representative to the question vector are selected, and finally, the target sample data is obtained. This process further optimizes the quality of the samples, ensuring that the target sample data can provide the most effective semantic information for model training. This embodiment not only realizes the effective conversion from text to vector but also ensures, through precise screening and optimization, that the target sample data can provide high-quality semantic information for the training of the word embedding model, thereby significantly improving the performance and generalization ability of the model.
[0102] Exemplarily, to facilitate understanding of the implementation process of the word embedding model training method based on unlabeled data obtained by combining the above Embodiment 1, please refer to Figure 3 , Figure 3 A brief flow schematic diagram of a word embedding model training method based on unlabeled data is provided. Specifically:
[0103] First, obtain the initial data and the initial word embedding model, and use the initial word embedding model to convert the answer part in the initial data into an answer vector to construct a vector database. Then, generate question vectors through preset question data, and based on the cosine similarity calculation between the vector database and the question vectors, filter out the initial sample data. Next, use the preset re-ranking model to evaluate the relevance of the initial sample data, generate a sample evaluation table, and further filter out the target sample data according to the relevance score. Finally, train the initial word embedding model with the target sample data, and combine the preset similarity threshold and weight coefficient adjustment to optimize the model until convergence to obtain the target word embedding model.
[0104] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method for training a word embedding model based on unlabeled data in this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0105] This application also provides a device for training a word embedding model based on unlabeled data. Please refer to Figure 4 , the device for training a word embedding model based on unlabeled data includes:
[0106] An acquisition module 10, configured to acquire initial data and an initial word embedding model;
[0107] A obtaining module 20, configured to perform vector representation on the initial data according to the initial word embedding model to obtain a vector database;
[0108] A screening module 30, configured to perform similarity screening based on the vector database to obtain target sample data;
[0109] A completion module 40, configured to train the initial word embedding model with the target sample data to obtain a target word embedding model.
[0110] The device for training a word embedding model based on unlabeled data provided by this application adopts the method for training a word embedding model based on unlabeled data in the above embodiment, and can solve the technical problem of how to improve the training quality of the word embedding model. Compared with the prior art, the beneficial effects of the device for training a word embedding model based on unlabeled data provided by this application are the same as those of the method for training a word embedding model based on unlabeled data provided by the above embodiment, and other technical features in the device for training a word embedding model based on unlabeled data are the same as the features disclosed in the method of the above embodiment, and will not be elaborated here.
[0111] In one embodiment, the obtaining module 20 is further configured to obtain answer data according to the initial data; perform vector representation on the answer data through the initial word embedding model to obtain an answer vector; and obtain a vector database based on the answer vector.
[0112] In one embodiment, the screening module 30 is further configured to obtain preset question data; perform vector representation on the preset question data according to the initial word embedding model to obtain a question vector; perform cosine similarity calculation and screening based on the vector database and the question vector to obtain initial sample data; and screen the initial sample data to obtain target sample data.
[0113] In one embodiment, the screening module 30 is further configured to obtain a preset re-ranking model; perform a relevance evaluation on the initial sample data through the preset re-ranking model to obtain a sample evaluation table, where the sample evaluation table includes a mapping relationship between the initial sample data and the relevance score; and screen according to querying the sample evaluation table and based on the relevance score to obtain target sample data.
[0114] In one embodiment, the screening module 30 is further configured to perform word segmentation processing on the preset question data to obtain word element data; obtain word vectors based on the word element data and the initial word embedding model; and perform average pooling processing on the word vectors to obtain a question vector.
[0115] In one embodiment, the completion module 40 is further configured to obtain a preset similarity threshold and a weight coefficient of the initial word embedding model; obtain an initial positive example similarity value and an initial negative example similarity set based on the target sample data and the initial word embedding model; and adjust the weight coefficient based on the initial positive example similarity value, the initial negative example similarity set, and the preset similarity threshold until the initial word embedding model converges to obtain a target word embedding model.
[0116] In one embodiment, the completion module 40 is further configured to obtain an initial similarity value based on the initial positive example similarity value and the initial negative example similarity set; if the initial similarity value is greater than the preset similarity threshold, adjust the weight coefficient to obtain a target word embedding model.
[0117] The present application provides a word embedding model training device based on unlabeled data. The word embedding model training device based on unlabeled data includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the word embedding model training method based on unlabeled data in the first embodiment above.
[0118] Refer to the following Figure 5, which shows a schematic structural diagram of a word embedding model training device suitable for implementing the embodiments of the present application based on unlabeled data. The word embedding model training device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The shown word embedding model training device based on unlabeled data is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0119] As Figure 5 shown, the word embedding model training device based on unlabeled data may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM) 1004. In the RAM 1004, various programs and data required for the operation of the word embedding model training device based on unlabeled data are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the word embedding model training device based on unlabeled data to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a word embedding model training device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.
[0120] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0121] The word embedding model training device based on unlabeled data provided by the present application adopts the word embedding model training method based on unlabeled data in the above embodiments, and can solve the technical problem of how to improve the training quality of the word embedding model. Compared with the prior art, the beneficial effects of the word embedding model training device based on unlabeled data provided by the present application are the same as those of the word embedding model training method based on unlabeled data provided in the above embodiments, and other technical features in the word embedding model training device based on unlabeled data are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0122] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0123] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0124] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the word embedding model training method based on unlabeled data in the above embodiments.
[0125] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0126] The above computer-readable storage medium can be included in the word embedding model training device based on unlabeled data; it can also exist independently without being assembled into the word embedding model training device based on unlabeled data.
[0127] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the word embedding model training device based on unlabeled data, the word embedding model training device based on unlabeled data is enabled to: obtain initial data and an initial word embedding model; perform vector representation on the initial data according to the initial word embedding model to obtain a vector database; perform similarity screening based on the vector database to obtain target sample data; and train the initial word embedding model through the target sample data to obtain a target word embedding model.
[0128] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0130] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0131] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned method for training a word embedding model based on unlabeled data, and can solve the technical problem of how to improve the training quality of the word embedding model. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the method for training a word embedding model based on unlabeled data provided by the above embodiments, and will not be elaborated here.
[0132] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-described method for training a word embedding model based on unlabeled data.
[0133] The computer program product provided by the present application can solve the technical problem of how to improve the training quality of the word embedding model. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the method for training a word embedding model based on unlabeled data provided in the above embodiments, and will not be elaborated here.
[0134] The foregoing are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A word embedding model training method based on unlabeled data, characterized in that: The method comprises: Get initial data and initial word embedding model; Performing vector representation on the initial data according to the initial word embedding model to obtain a vector database; Perform similarity screening based on the vector database to obtain target sample data; The initial word embedding model is trained using the target sample data to obtain a target word embedding model.
2. The method according to claim 1, characterized in that The step of performing vector representation on the initial data according to the initial word embedding model to obtain a vector database includes: Obtain answer data according to the initial data; The answer data is represented by a vector through the initial word embedding model to obtain an answer vector; Based on the answer vector, a vector database is obtained.
3. The method according to claim 1, characterized in that The performing similarity screening based on the vector database to obtain target sample data includes: Get preset question data; Performing vector representation on the preset question data according to the initial word embedding model to obtain a question vector; Calculate and screen the cosine similarity based on the vector database and the question vector to obtain initial sample data; The initial sample data is screened to obtain target sample data.
4. The method according to claim 3, characterized in that The screening of the initial sample data to obtain target sample data includes: Get the preset reordering model; Performing relevance evaluation on the initial sample data by using the preset reordering model to obtain a sample evaluation table, wherein the sample evaluation table includes a mapping relationship between the initial sample data and the relevance score; The target sample data is obtained by querying the sample evaluation table and screening based on the relevance score.
5. The method according to claim 3, characterized in that The step of performing vector representation on the preset question data according to the initial word embedding model to obtain a question vector includes: Performing word segmentation processing on the preset question data to obtain word metadata; Obtaining a word vector based on the word metadata and the initial word embedding model; Average pooling is performed based on the word vector to obtain a question vector.
6. The method according to claim 1, characterized in that The step of training the initial word embedding model by using the target sample data to obtain a target word embedding model includes: Obtaining a preset similarity threshold and a weight coefficient of the initial word embedding model; Based on the target sample data and the initial word embedding model, obtaining an initial positive example similarity value and an initial negative example similarity set; Based on the initial positive example similarity value, the initial negative example similarity set and the preset similarity threshold, the weight coefficient is adjusted until the initial word embedding model converges to obtain a target word embedding model.
7. The method according to claim 6, characterized in that The adjusting the weight coefficient based on the initial positive example similarity value, the initial negative example similarity set and the preset similarity threshold until the initial word embedding model converges to obtain a target word embedding model includes: Obtaining an initial similarity value based on the initial positive example similarity value and the initial negative example similarity set; If the initial similarity value is greater than the preset similarity threshold, the weight coefficient is adjusted to obtain a target word embedding model.
8. A word embedding model training device based on unlabeled data, characterized in that: The device comprises: The acquisition module is used to obtain initial data and initial word embedding model; An obtaining module, used for performing vector representation on the initial data according to the initial word embedding model to obtain a vector database; A screening module, used for performing similarity screening based on the vector database to obtain target sample data; The completion module is used to train the initial word embedding model through the target sample data to obtain a target word embedding model.
9. A word embedding model training device based on unlabeled data, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the word embedding model training method based on unlabeled data as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the word embedding model training method based on unlabeled data are implemented as described in any one of claims 1 to 7.