Evaluation Method, Device, Electronic Device and Medium of Word Vector Model

By calculating the word similarity in the search request and target response, the quality of the word vector model is automatically evaluated, and the evaluation accuracy problems caused by relying on manual annotation in the prior art are solved, and the accuracy and automation of the evaluation results are improved.

CN114021552BActive Publication Date: 2025-06-13SHANDONG KURUI TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111302051.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-06-13
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

The evaluation methods of existing word vector models rely on manual labeling data, which makes it difficult to guarantee the accuracy of the evaluation results. Especially for a large number of word vector data, manual labeling errors will directly affect the evaluation results.

Method used

By obtaining the words corresponding to the search request and the words in the target response, the cosine similarity of the words is calculated based on the target word vector model, and the quality evaluation results of the word vector model are automatically determined to avoid manual labeling errors.

Benefits of technology

It improves the accuracy of the quality evaluation of word vector model, establishes a strong correlation between search requests and user click content, and reduces the dependence of manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114021552B_ABST
    Figure CN114021552B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, electronic device, and medium for evaluating a word vector model. The method includes: obtaining at least one first word corresponding to a first search request; obtaining at least one second word corresponding to a target response of the first search request, where the target response is used to describe a search result corresponding to the first search request; determining a similarity between at least one first word and at least one second word based on a target word vector model; and determining a quality evaluation result of the target word vector model according to the similarity between at least one first word and at least one second word. The embodiments of the present disclosure can avoid the problem that the quality evaluation result is affected by manual annotation errors and effectively improve the accuracy of the quality evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of model evaluation, and particularly to a method, device, electronic device and medium for evaluating a word vector model. Background Art

[0002] The corpus training model is the most common underlying model in the field of Neuro-Linguistic Programming (NLP) and can be used in almost all NLP upstream tasks. Word vectors can transform words in the training corpus into low-dimensional dense vector forms. The quality of corpus word vector training determines whether it can help users recall the desired content in the search engine and display the recalled text at the front of the list. Therefore, the quality evaluation of the word vector model becomes extremely important.

[0003] The current evaluation method is mainly achieved through relevance evaluation, that is, directly measuring the syntactic and semantic relationships between two given words and evaluating using the similarity between the labeled data and the trained word vectors.

[0004] However, the labeled data needs to be manually labeled, and the labeling accuracy will affect the evaluation effect. For a large amount of word vector data, once the manual labeling is incorrect, it will directly affect the evaluation result, and it is difficult to ensure the accuracy of the evaluation result. Summary of the Invention

[0005] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, device, electronic device and medium for evaluating a word vector model.

[0006] In a first aspect, the present disclosure provides a method for evaluating a word vector model, including:

[0007] Obtaining at least one first word corresponding to a first search request;

[0008] Obtaining at least one second word corresponding to the target response of the first search request, where the target response is used to describe the search result corresponding to the first search request;

[0009] Based on the target word vector model, determining the similarity between the at least one first word and the at least one second word;

[0010] Determining the quality evaluation result of the target word vector model according to the similarity between the at least one first word and the at least one second word.

[0011] Optionally, the determining the similarity between the at least one first word and the at least one second word based on the target word vector model includes:

[0012] Based on the target word vector model, calculate the cosine similarity between each of the first words and the at least one second word respectively to obtain a first similarity set;

[0013] Based on the target word vector model, calculate the cosine similarity between each of the second words and the at least one first word respectively to obtain a second similarity set;

[0014] Determine the similarity between the at least one first word and the at least one second word according to the first similarity set and the second similarity set.

[0015] Optionally, the step of calculating the cosine similarity between each of the first words and the at least one second word respectively based on the target word vector model to obtain a first similarity set includes:

[0016] Based on the target word vector model, calculate the cosine similarity between each of the first words and the at least one second word respectively to obtain at least one cosine similarity corresponding to each of the first words;

[0017] Determine the first similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each of the first words.

[0018] Optionally, the step of calculating the cosine similarity between each of the second words and the at least one first word respectively based on the target word vector model to obtain a second similarity set includes:

[0019] Based on the target word vector model, calculate the cosine similarity between each of the second words and the at least one first word respectively to obtain at least one cosine similarity corresponding to each of the second words;

[0020] Determine the second similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each of the second words.

[0021] Optionally, the step of determining the similarity between the at least one first word and the at least one second word according to the first similarity set and the second similarity set includes:

[0022] Determine the mean value of all the cosine similarities included in the first similarity set as the first average similarity;

[0023] Determine the mean value of all the cosine similarities included in the second similarity set as the second average similarity;

[0024] Determine the similarity between the at least one first word and the at least one second word according to the first average similarity and the second average similarity.

[0025] Optionally, determining the quality evaluation result of the target word vector model according to the similarity between the at least one first word and the at least one second word includes:

[0026] When it is detected that the similarity between the at least one first word and the at least one second word is greater than or equal to a preset similarity threshold, the at least one second word is marked as an associated word; or, when it is detected that the similarity between the at least one first word and the at least one second word is less than the preset similarity threshold, the at least one second word is marked as a non-associated word;

[0027] Determine the quality evaluation result of the target word vector model according to the marking results of other words corresponding to other responses of the first search request.

[0028] Optionally, determining the quality evaluation result of the target word vector model according to the marking results of other words corresponding to other responses of the first search request includes:

[0029] Determine the association accuracy of the target word vector model according to the marking results of other words corresponding to other responses of the first search request;

[0030] Determine the quality evaluation result of the target word vector model according to the association accuracy and a preset accuracy threshold.

[0031] In a second aspect, the present disclosure provides an evaluation device for a word vector model, including:

[0032] A first acquisition module, configured to acquire at least one first word corresponding to a first search request;

[0033] A second acquisition module is further configured to acquire at least one second word corresponding to a target response of the first search request, where the target response is used to describe a search result corresponding to the first search request;

[0034] A first determination module, configured to determine the similarity between the at least one first word and the at least one second word based on a target word vector model;

[0035] A second determination module, configured to determine the quality evaluation result of the target word vector model according to the similarity between the at least one first word and the at least one second word.

[0036] Optionally, the first determination module includes: a first calculation unit, a second calculation unit, and a first determination unit;

[0037] A first calculation unit, configured to calculate the cosine similarity between each of the first words and the at least one second word based on a target word vector model, so as to obtain a first similarity set;

[0038] A second calculation unit, further configured to calculate the cosine similarity between each of the second words and the at least one first word based on the target word vector model, so as to obtain a second similarity set;

[0039] A first determination unit, configured to determine the similarity between the at least one first word and the at least one second word according to the first similarity set and the second similarity set.

[0040] Optionally, the first calculation unit is specifically configured to:

[0041] Calculate the cosine similarity between each of the first words and the at least one second word based on the target word vector model, so as to obtain at least one cosine similarity corresponding to each of the first words;

[0042] Determine the first similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each of the first words.

[0043] Optionally, the second calculation unit is specifically configured to:

[0044] Calculate the cosine similarity between each of the second words and the at least one first word based on the target word vector model, so as to obtain at least one cosine similarity corresponding to each of the second words;

[0045] Determine the second similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each of the second words.

[0046] Optionally, the first determination unit is specifically configured to:

[0047] Determine the mean value of all the cosine similarities included in the first similarity set as the first average similarity;

[0048] Determine the mean value of all the cosine similarities included in the second similarity set as the second average similarity;

[0049] Determine the similarity between the at least one first word and the at least one second word according to the first average similarity and the second average similarity.

[0050] Optionally, the second determination module includes: a detection unit and a second determination unit;

[0051] A detection unit, configured to mark the at least one second word as an associated word when the similarity between the at least one first word and the at least one second word is greater than or equal to a preset similarity threshold; or mark the at least one second word as a non-associated word when the similarity between the at least one first word and the at least one second word is less than the preset similarity threshold.

[0052] A second determination unit, configured to determine a quality evaluation result of the target word vector model according to a marking result of other words corresponding to other responses of the first search request.

[0053] Optionally, the second determination unit is specifically configured to:

[0054] Determine an association accuracy of the target word vector model according to a marking result of other words corresponding to other responses of the first search request;

[0055] Determine a quality evaluation result of the target word vector model according to the association accuracy and a preset accuracy threshold.

[0056] In a third aspect, the present disclosure further provides an electronic device, including:

[0057] One or more processors;

[0058] A storage device, configured to store one or more programs,

[0059] When the one or more programs are executed by the one or more processors, the one or more processors implement any one of the word vector model evaluation methods in the embodiments of the present invention.

[0060] In a fourth aspect, the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the word vector model evaluation methods in the embodiments of the present invention.

[0061] The technical solutions provided in the embodiments of the present disclosure have the following advantages compared with the prior art: By obtaining at least one first word corresponding to a first search request and at least one second word corresponding to a target response of the first search request, where the target response is used to describe a search result corresponding to the first search request, and the target response can effectively reflect the user's feedback result to directly record the content that the user is really interested in, thereby establishing a strong correlation between the search request and the user-clicked content; and based on the target word vector model, determining the similarity between at least one first word and at least one second word to determine a quality evaluation result of the target word vector model, avoiding the problem that manual annotation errors affect the quality evaluation result, and effectively improving the accuracy of the quality evaluation result. Description of the Drawings

[0062] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0063] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings required for use in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative efforts.

[0064] Figure 1 It is a schematic flowchart of a method for evaluating a word vector model provided by an embodiment of the present disclosure;

[0065] Figure 2 It is a schematic flowchart of another method for evaluating a word vector model provided by an embodiment of the present disclosure;

[0066] Figure 3 It is a schematic diagram of a process for evaluating a word vector model provided by an embodiment of the present disclosure;

[0067] Figure 4 It is a schematic structural diagram of an apparatus for evaluating a word vector model provided by an embodiment of the present disclosure;

[0068] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0069] In order to be able to more clearly understand the above objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0070] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0071] Currently, the models for training word vectors include fast text classification models (such as fasttext), neural network models (such as word2vec), glove models, language representation models (such as bert), etc. However, the evaluation methods for the trained word vectors, especially Chinese word vectors, are not complete. Therefore, it is difficult to effectively identify the advantages and disadvantages of the word vector models.

[0072] Among them, in the prior art, evaluating the quality of word vectors by determining whether there is a relatively high correlation between the cosine similarity of two words calculated by a word vector model and the human subjective evaluation score, and whether the words selected by testers from a word set are consistent with the added irrelevant words requires a large amount of manual annotation. Moreover, the accuracy of the annotation has a very significant impact on the evaluation index. For a large amount of word vector data, the above methods are relatively difficult to implement.

[0073] Based on this, the present disclosure proposes to avoid using manual labor and be able to automatically and intelligently evaluate a word vector model. After a user performs a search action, according to the feedback results, the user will click on the content that they are truly interested in. Therefore, a strong correlation can be established between the search request and the content clicked by the user. In terms of words, the keywords in the search request and the keywords in the text of the clicked content should also be strongly semantically related, and the distance between these words in the pre-trained model should be closer. Thus, automatic annotation is achieved while intelligently evaluating the quality of the word vector model.

[0074] Exemplarily, the present disclosure provides a method, apparatus, electronic device, and medium for evaluating a word vector model. By obtaining at least one first word corresponding to a first search request and at least one second word corresponding to the target response of the first search request, where the target response is used to describe the search result corresponding to the first search request, and the target response can effectively reflect the feedback result of the user, so as to directly record the content that the user is truly interested in, thereby establishing a strong correlation between the search request and the content clicked by the user; and based on the target word vector model, determining the similarity between at least one first word and at least one second word to determine the quality evaluation result of the target word vector model, avoiding the problem that manual annotation errors affect the quality evaluation result and effectively improving the accuracy of the quality evaluation result.

[0075] Among them, the method for evaluating the word vector model of the present disclosure is executed by an electronic device or a client installed in the electronic device. The electronic device can be a tablet computer, a mobile phone, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart TV, a smart screen, a high-definition TV, a 4K TV, a smart speaker, a smart projector, etc. The present disclosure does not impose any restrictions on the specific type of the electronic device.

[0076] Among them, the present disclosure does not limit the type of the operating system of the electronic device. For example, Android system, Linux system, Windows system, iOS system, etc.

[0077] For details, please refer to Figure 1 .

[0078] Figure 1 FIG. is a schematic flow chart of a method for evaluating a word vector model provided by an embodiment of the present disclosure. The method of this embodiment can be executed by an evaluation device of the word vector model. The device can be implemented in a hardware / or software manner and can be configured in an electronic device. The method for evaluating the word vector model described in any embodiment of the present application can be implemented. As Figure 1 shown, the method specifically includes the following:

[0079] S110. Obtain at least one first word corresponding to the first search request.

[0080] Among them, at least one first word can be obtained by collecting the user's historical search requests and segmenting the request statements included in the historical search requests.

[0081] Among them, the first search request is any one of the search requests in the user's historical search requests.

[0082] S120. Obtain at least one second word corresponding to the target response of the first search request.

[0083] Among them, the target response is used to describe the search results corresponding to the first search request.

[0084] Among them, the target response is the item title in the search results corresponding to the first search request. Keyword extraction can be performed on the item title to obtain at least one second word corresponding to the target response.

[0085] For example, the first search request is "Shaanxi cuisine", the item title corresponding to the first search request can be "Top Ten Shaanxi Cuisine Rankings", and the specific content under the item title can be used as the feedback content of the first search request.

[0086] It should be noted that the keyword extraction method can include but is not limited to: term frequency–inverse document frequency (tf-idf) algorithm or text ranking (such as TextRank) algorithm.

[0087] S130. Based on the target word vector model, determine the similarity between at least one first word and at least one second word.

[0088] Among them, the similarity between at least one first word and at least one second word may include: the similarity between each first word included in at least one first word and each second word included in at least one second word.

[0089] S140. Determine the quality evaluation result of the target word vector model according to the similarity between at least one first word and at least one second word.

[0090] Among them, a threshold set in advance can be used to judge the similarity between at least one first word and at least one second word, so as to determine the quality of the quality evaluation result of the target word vector model.

[0091] In this embodiment, optionally, determining the quality evaluation result of the target word vector model according to the similarity between at least one first word and at least one second word includes:

[0092] When it is detected that the similarity between at least one first word and at least one second word is greater than or equal to the preset similarity threshold, mark at least one second word as an associated word; or, when it is detected that the similarity between at least one first word and at least one second word is less than the preset similarity threshold, mark at least one second word as a non-associated word;

[0093] Determine the quality evaluation result of the target word vector model according to the marking results of other words corresponding to other responses of the first search request.

[0094] Among them, when it is detected that the similarity between at least one first word and at least one second word is greater than or equal to the preset similarity threshold, the data corresponding to at least one second word can be marked to effectively distinguish the recognition rate of the target word vector model.

[0095] For example, mark the label corresponding to at least one second word as 1, indicating that at least one second word is an associated word, or mark the label corresponding to at least one second word as 0, indicating that at least one second word is a non-associated word.

[0096] Thus, based on the multiple marking results of other words corresponding to other responses of the first search request, the quality evaluation result of the target word vector model can be effectively measured.

[0097] In this embodiment, optionally, determining the quality evaluation result of the target word vector model according to the marking results of other words corresponding to other responses of the first search request includes:

[0098] Determine the association accuracy of the target word vector model according to the marking results of other words corresponding to other responses of the first search request.

[0099] Determine the quality evaluation result of the target word vector model according to the correlation accuracy and a preset accuracy threshold.

[0100] Among them, the correlation accuracy of the target word vector model can be determined according to the ratio of the number of words marked as related words in the marking results of other words corresponding to other responses of the first search request to all the marked data.

[0101] For example, when it is determined that the correlation accuracy is greater than or equal to the preset accuracy threshold, it is determined that the quality evaluation result of the target word vector model is good; or, when it is determined that the correlation accuracy is less than the preset accuracy threshold, it is determined that the quality evaluation result of the target word vector model is poor.

[0102] Among them, the preset accuracy threshold can be adaptively adjusted based on different evaluation requirements, and the present disclosure does not make specific limitations thereto.

[0103] Thus, by comparing the preset accuracy threshold with the marking result, qualitatively judge the quality evaluation result of the target word vector model.

[0104] In addition, the method of this embodiment can also judge the quality of two word vector models. For example, the quality of the word vector model is determined by its corresponding correlation accuracy. The specific implementation manner of the correlation accuracy is the same as the above manner and will not be elaborated here.

[0105] The evaluation method of the word vector model provided in this embodiment obtains at least one first word corresponding to the first search request and at least one second word corresponding to the target response of the first search request. The target response is used to describe the search result corresponding to the first search request. Among them, the target response can effectively reflect the user's feedback result to directly record the content that the user is really interested in. Thus, a strong correlation between the search request and the content clicked by the user is established; and based on the target word vector model, determine the similarity between at least one first word and at least one second word to determine the quality evaluation result of the target word vector model, avoiding the problem that the manual annotation error affects the quality evaluation result, and effectively improving the accuracy of the quality evaluation result.

[0106] Figure 2 It is a schematic flowchart of another evaluation method of the word vector model provided by the embodiments of the present disclosure. This embodiment is based on the above embodiment. Among them, a possible implementation manner of S130 is as follows:

[0107] S1301. Based on the target word vector model, calculate the cosine similarity between each first word and at least one second word respectively to obtain a first similarity set.

[0108] Among them, the first similarity set includes the cosine similarities between each first word and all or part of the second words included in at least one second word.

[0109] In this embodiment, optionally, based on the target word vector model, calculate the cosine similarities between each first word and at least one second word respectively to obtain a first similarity set, including:

[0110] Based on the target word vector model, calculate the cosine similarities between each first word and at least one second word respectively to obtain at least one cosine similarity corresponding to each first word;

[0111] Determine the first similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each first word.

[0112] For example, if the number of second words is 5, then each first word will correspond to 5 cosine similarities. In the present disclosure, each first word only takes the maximum one cosine similarity from its corresponding cosine similarity set to form the first similarity set, that is, respectively take the maximum one cosine similarity from the 5 cosine similarities and add it to the first similarity set.

[0113] Thus, it is possible to construct the first similarity set based on the maximum cosine similarity between each first word and each second word, so that the correlation between the first words and the second words corresponding to the first similarity set is relatively high.

[0114] S1302. Based on the target word vector model, calculate the cosine similarities between each second word and at least one first word respectively to obtain a second similarity set.

[0115] Among them, the second similarity set includes the cosine similarities between each second word and all or part of the first words included in at least one first word.

[0116] In this embodiment, optionally, based on the target word vector model, calculate the cosine similarities between each second word and at least one first word respectively to obtain a second similarity set, including:

[0117] Based on the target word vector model, calculate the cosine similarities between each second word and at least one first word respectively to obtain at least one cosine similarity corresponding to each second word;

[0118] Determine the second similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each second word.

[0119] For example, if the number of the first words is six, each second word will correspond to six cosine similarities. In the present disclosure, only the largest cosine similarity in the corresponding cosine similarity set of each second word is taken to form a second similarity set. That is, the largest one of the six cosine similarities is respectively added to the second similarity set.

[0120] Therefore, a second similarity set can be constructed based on the largest cosine similarity between each second word and each first word, so that the correlation between the second words corresponding to the second similarity set and the first words is relatively high.

[0121] S1303. Determine the similarity between at least one first word and at least one second word according to the first similarity set and the second similarity set.

[0122] Among them, a similarity can be determined by weighted value taking, mean value taking, median value taking, etc. for each cosine similarity in the first similarity set and the second similarity set respectively, so as to effectively measure the similarity between at least one first word and at least one second word.

[0123] In this embodiment, optionally, determining the similarity between at least one first word and at least one second word according to the first similarity set and the second similarity set includes:

[0124] Determine the first average similarity as the mean value of all cosine similarities included in the first similarity set;

[0125] Determine the second average similarity as the mean value of all cosine similarities included in the second similarity set;

[0126] Determine the similarity between at least one first word and at least one second word according to the first average similarity and the second average similarity.

[0127] Among them, the mean values of all cosine similarities in the first similarity set and the second similarity set can be calculated to obtain two average similarities, and then the mean value of the two average similarities is used to effectively measure the similarity between at least one first word and at least one second word.

[0128] Based on the description of the above embodiments, the present disclosure also provides a schematic diagram of the evaluation process of a word vector model, as Figure 3 exemplarily shown.

[0129] Among them, A1, A2, ..., AM respectively represent different first words, B1, B2, ..., BN respectively represent different second words, A1_MAX represents the maximum cosine similarity among the cosine similarities of A1 with B1, B2, ..., BN, to obtain a first similarity set, B1_MAX represents the maximum cosine similarity among the cosine similarities of B1 with A1, A2, ..., AM, to obtain a second similarity set, determine the mean value as A_AGV from the first similarity set, determine the mean value as B_AGV from the second similarity set, then calculate the mean value of the two as score, when it is determined that score is greater than or equal to the preset threshold T, label 1 for the corresponding group of data, so as to calculate the accuracy rate of the target word vector model, and determine the quality evaluation result of the target word vector model based on the accuracy rate.

[0130] Figure 4 It is a schematic structural diagram of an evaluation device for a word vector model provided by an embodiment of the present disclosure; this device is configured in an electronic device and can implement the evaluation method of the word vector model described in any embodiment of the present application. This device specifically includes the following:

[0131] The first acquisition module 410 is used to acquire at least one first word corresponding to the first search request;

[0132] The second acquisition module 420 is further used to acquire at least one second word corresponding to the target response of the first search request, and the target response is used to describe the search result corresponding to the first search request;

[0133] The first determination module 430 is used to determine the similarity between the at least one first word and the at least one second word based on the target word vector model;

[0134] The second determination module 440 is used to determine the quality evaluation result of the target word vector model according to the similarity between the at least one first word and the at least one second word.

[0135] In this embodiment, optionally, the first determination module 430 includes: a first calculation unit, a second calculation unit, and a first determination unit;

[0136] The first calculation unit is used to calculate the cosine similarity between each of the first words and the at least one second word based on the target word vector model, to obtain a first similarity set;

[0137] The second calculation unit is further used to calculate the cosine similarity between each of the second words and the at least one first word based on the target word vector model, to obtain a second similarity set;

[0138] A first determination unit, configured to determine the similarity between the at least one first word and the at least one second word according to the first similarity set and the second similarity set.

[0139] In this embodiment, optionally, the first calculation unit is specifically configured to:

[0140] Based on the target word vector model, calculate the cosine similarity between each of the first words and the at least one second word respectively, to obtain at least one cosine similarity corresponding to each of the first words;

[0141] Determine the first similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each of the first words.

[0142] In this embodiment, optionally, the second calculation unit is specifically configured to:

[0143] Based on the target word vector model, calculate the cosine similarity between each of the second words and the at least one first word respectively, to obtain at least one cosine similarity corresponding to each of the second words;

[0144] Determine the second similarity set according to the maximum cosine similarity among the at least one cosine similarity corresponding to each of the second words.

[0145] In this embodiment, optionally, the first determination unit is specifically configured to:

[0146] Determine the mean value of all the cosine similarities included in the first similarity set as the first average similarity;

[0147] Determine the mean value of all the cosine similarities included in the second similarity set as the second average similarity;

[0148] Determine the similarity between the at least one first word and the at least one second word according to the first average similarity and the second average similarity.

[0149] In this embodiment, optionally, the second determination module 440 includes: a detection unit and a second determination unit;

[0150] The detection unit is configured to, when detecting that the similarity between the at least one first word and the at least one second word is greater than or equal to a preset similarity threshold, mark the at least one second word as an associated word; or, when detecting that the similarity between the at least one first word and the at least one second word is less than the preset similarity threshold, mark the at least one second word as a non-associated word;

[0151] A second determination unit, configured to determine a quality evaluation result of the target word vector model according to tagging results of other words corresponding to other responses of the first search request.

[0152] In this embodiment, optionally, the second determination unit is specifically configured to:

[0153] Determine an association accuracy of the target word vector model according to tagging results of other words corresponding to other responses of the first search request;

[0154] Determine a quality evaluation result of the target word vector model according to the association accuracy and a preset accuracy threshold.

[0155] Through the word vector model evaluation device of the embodiment of the present invention, by obtaining at least one first word corresponding to a first search request and at least one second word corresponding to a target response of the first search request, where the target response is used to describe a search result corresponding to the first search request, and the target response can effectively reflect the user's feedback result to directly record the content that the user is really interested in, thereby establishing a strong correlation between the search request and the content clicked by the user; and based on the target word vector model, determine the similarity between at least one first word and at least one second word to determine a quality evaluation result of the target word vector model, avoiding the problem that manual annotation errors affect the quality evaluation result and effectively improving the accuracy of the quality evaluation result.

[0156] The word vector model evaluation device provided by the embodiment of the present invention can execute the word vector model evaluation method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0157] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 5 shown, the electronic device includes a processor 510, a memory 520, an input device 530, and an output device 540; the number of processors 510 in the electronic device can be one or more, Figure 5 taking one processor 510 as an example; the processor 510, the memory 520, the input device 530, and the output device 540 in the electronic device can be connected through a bus or other means, Figure 5 taking connection through a bus as an example.

[0158] The memory 520, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the evaluation method of the word vector model in the embodiments of the present invention. The processor 510 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 520, that is, implements the evaluation method of the word vector model provided by the embodiments of the present invention.

[0159] The memory 520 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 520 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 520 may further include a memory remotely set relative to the processor 510, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0160] The input device 530 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function control of the electronic device, and may include a keyboard, a mouse, etc. The output device 540 may include a display device such as a display screen.

[0161] The embodiments of the present disclosure also provide a storage medium containing computer-executable instructions, and the computer-executable instructions are used to implement the evaluation method of the word vector model provided by the embodiments of the present invention when executed by a computer processor.

[0162] Of course, for a storage medium containing computer-executable instructions provided by the embodiments of the present invention, the computer-executable instructions are not limited to the method operations as described above, and can also execute the related operations in the evaluation method of the word vector model provided by any embodiment of the present invention.

[0163] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0164] It should be noted that in the embodiments of the above search device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0165] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0166] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for evaluating a word vector model, characterized in that, it includes: Obtain at least one first word corresponding to a first search request; Obtain at least one second word corresponding to the target response of the first search request, where the target response is used to describe the search result corresponding to the first search request; Based on the target word vector model, determine the similarity between the at least one first word and the at least one second word; According to the similarity between the at least one first word and the at least one second word, determine the quality evaluation result of the target word vector model; The determining the similarity between the at least one first word and the at least one second word based on the target word vector model includes: Based on the target word vector model, calculate the cosine similarity between each of the first words and the at least one second word respectively to obtain a first similarity set; the first similarity set includes the maximum cosine similarity among the at least one cosine similarities corresponding to each of the first words; Based on the target word vector model, calculate the cosine similarity between each of the second words and the at least one first word respectively to obtain a second similarity set; the second similarity set includes the maximum cosine similarity among the at least one cosine similarities corresponding to each of the second words; According to the first similarity set and the second similarity set, determine the similarity between the at least one first word and the at least one second word; The determining the similarity between the at least one first word and the at least one second word according to the first similarity set and the second similarity set includes: Determine the mean value of all the cosine similarities included in the first similarity set as the first average similarity; Determine the mean value of all the cosine similarities included in the second similarity set as the second average similarity; According to the first average similarity and the second average similarity, determine the similarity between the at least one first word and the at least one second word.

2. The method according to claim 1, characterized in that, the calculating the cosine similarity between each of the first words and the at least one second word based on the target word vector model to obtain a first similarity set includes: Based on the target word vector model, calculate the cosine similarity between each of the first words and the at least one second word respectively to obtain at least one cosine similarity corresponding to each of the first words; According to the maximum cosine similarity among the at least one cosine similarities corresponding to each of the first words, determine the first similarity set.

3. The method according to claim 1, characterized in that, the calculating the cosine similarity between each of the second words and the at least one first word based on the target word vector model to obtain a second similarity set includes: Based on the target word vector model, calculate the cosine similarity between each of the second words and the at least one first word respectively to obtain at least one cosine similarity corresponding to each of the second words; According to the maximum cosine similarity among the at least one cosine similarities corresponding to each of the second words, determine the second similarity set.

4. The method according to claim 1, wherein, determining the quality evaluation result of the target word vector model according to the similarity between the at least one first word and the at least one second word includes: when it is detected that the similarity between the at least one first word and the at least one second word is greater than or equal to a preset similarity threshold, marking the at least one second word as an associated word; or, when it is detected that the similarity between the at least one first word and the at least one second word is less than the preset similarity threshold, marking the at least one second word as a non-associated word; determining the quality evaluation result of the target word vector model according to the marking results of other words corresponding to other responses of the first search request.

5. The method according to claim 4, wherein, determining the quality evaluation result of the target word vector model according to the marking results of other words corresponding to other responses of the first search request includes: determining the association accuracy of the target word vector model according to the marking results of other words corresponding to other responses of the first search request; determining the quality evaluation result of the target word vector model according to the association accuracy and a preset accuracy threshold.

6. An evaluation device for a word vector model, wherein, it includes: a first acquisition module, configured to acquire at least one first word corresponding to a first search request; a second acquisition module, further configured to acquire at least one second word corresponding to a target response of the first search request, where the target response is used to describe a search result corresponding to the first search request; a first determination module, configured to determine the similarity between the at least one first word and the at least one second word based on a target word vector model; a second determination module, configured to determine the quality evaluation result of the target word vector model according to the similarity between the at least one first word and the at least one second word; the first determination module includes: a first calculation unit, a second calculation unit, and a first determination unit; the first calculation unit is configured to calculate the cosine similarity between each of the first words and the at least one second word based on the target word vector model to obtain a first similarity set; the first similarity set includes the maximum cosine similarity among at least one cosine similarity corresponding to each of the first words; the second calculation unit is further configured to calculate the cosine similarity between each of the second words and the at least one first word based on the target word vector model to obtain a second similarity set; the second similarity set includes the maximum cosine similarity among at least one cosine similarity corresponding to each of the second words; the first determination unit is configured to determine the similarity between the at least one first word and the at least one second word according to the first similarity set and the second similarity set; The first determining unit is specifically configured to determine the mean value of all cosine similarities included in the first similarity set as the first average similarity; determine the mean value of all cosine similarities included in the second similarity set as the second average similarity; and determine the similarity between the at least one first word and the at least one second word according to the first average similarity and the second average similarity.

7. An electronic device characterized in that it includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the evaluation method of the word vector model according to any one of claims 1 to 5.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the evaluation method of the word vector model according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text retrieval method and device

    CN110019669A