A method and apparatus for model training and semantic similarity determination
By generating feature vectors from historical search results and using semantic similarity to determine model training, the problem of insufficient semantic similarity calculation time and accuracy in existing technologies is solved, achieving a more efficient information search effect.
Patent Information
- Application Number
- CN202210577430.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-05-25
AI Technical Summary
While existing technologies can reduce the time required to calculate semantic similarity, they cannot guarantee the accuracy of the calculations, thus affecting the effectiveness of information retrieval.
By acquiring the semantic similarity between historical search results and search terms, a semantic vector is generated using the first sub-model. This vector is then combined with the feature extraction and feature fusion layers of the second sub-model to generate feature vectors for historical search results. This process trains the semantic similarity determination model, thereby improving the accuracy of semantic similarity.
By introducing the interaction between the semantic vectors of historical search terms and search results, the accuracy of semantic similarity determination is improved, thereby enhancing the effectiveness of information retrieval.
Smart Images

Figure CN114970545B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method and apparatus for model training and semantic similarity determination. Background Technology
[0002] Currently, with the development of computer technology, content-based search results have become one of people's sources of information. Information search platforms can display several search results to users based on the search terms they input. Among these, the semantic similarity between the search results and the search terms directly affects the information search effect.
[0003] Because information search scenarios require real-time return of search results to users, current similarity calculation methods focus more on reducing the time required to calculate the semantic similarity between search terms and search results. However, this does not guarantee the accuracy of semantic similarity calculation.
[0004] Improving the accuracy of semantic similarity calculations while reducing the computation time has become an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a method and apparatus for model training and semantic similarity determination, in order to partially solve the aforementioned problems existing in the prior art.
[0006] The following technical solution is adopted in this specification:
[0007] This manual provides a model training method, including:
[0008] Obtain historical search results, historical search terms, and the actual semantic similarity between the historical search results and the historical search terms;
[0009] The historical search terms are input into the first sub-model of the semantic similarity determination model to be trained, and the semantic vector of the historical search terms is obtained through the first sub-model.
[0010] The historical search results are used as input to the second sub-model of the semantic similarity determination model to be trained, and the first feature vector of the historical search results is determined by the feature extraction layer in the second sub-model.
[0011] The semantic vector of the historical search term is input into the second sub-model. The second sub-model performs feature fusion on the semantic vector of the historical search term and the first feature vector of the historical search result. The fused feature vector is used as the second feature vector of the historical search result.
[0012] Based on the similarity between the semantic vector of the historical search term and the second feature vector of the historical search result, the predicted semantic similarity between the historical search result and the historical search term is determined;
[0013] The semantic similarity determination model to be trained is trained based on the predicted semantic similarity and the actual semantic similarity.
[0014] Optionally, the first sub-model includes a first semantic vector layer;
[0015] The historical search terms are input into the first sub-model of the semantic similarity determination model to be trained, and the semantic vector of the historical search terms is obtained through the first sub-model, specifically including:
[0016] The historical search terms are input into the first sub-model to obtain the semantic vector of the historical search terms output by the first semantic vector layer in the first sub-model.
[0017] Optionally, before the historical search results are input into the second sub-model of the semantic similarity determination model to be trained, and before the feature extraction layer in the second sub-model determines the first feature vector of the historical search results, the method further includes:
[0018] Obtain preset search intents of various types;
[0019] The preset search intents of each type are input into the third sub-model of the semantic similarity determination model to be trained, and feature vectors of the preset search intents of each type are output by the third sub-model.
[0020] The historical search results are used as input to the second sub-model of the semantic similarity determination model to be trained. The feature extraction layer in the second sub-model determines the first feature vector of the historical search results, specifically including:
[0021] Based on the feature vectors of the preset search intents of each type and the historical search results, the feature extraction layer in the second sub-model outputs the first feature vector of the historical search results.
[0022] Optionally, based on the feature vectors of the preset search intents for each type and the historical search results, the feature extraction layer in the second sub-model outputs a first feature vector of the historical search results, specifically including:
[0023] The historical search results are input into the second sub-model to obtain the semantic vector of the historical search results output by the second semantic vector layer in the second sub-model;
[0024] For each type of search intent, the semantic vector of the historical search results and the feature vector of the search intent of that type are input into the feature extraction layer corresponding to the search intent of that type in the second sub-model to obtain the first feature vector output by the feature extraction layer corresponding to the search intent of that type; the first feature vector output by the feature extraction layer corresponding to the search intent of that type is used to characterize the semantics of the historical search results under the search intent of that type.
[0025] Optionally, before inputting the semantic vector of the historical search terms into the second sub-model, and performing feature fusion on the semantic vector of the historical search terms and the first feature vector of the historical search results through the second sub-model, and using the fused feature vector as the second feature vector of the historical search results, the method further includes:
[0026] The historical search terms are input into a pre-trained search intent recognition model to obtain the actual search intent of the historical search terms output by the search intent recognition model.
[0027] Based on the first feature vectors of the historical search results and the actual search intent of the historical search terms, a specified feature vector of the historical search results is determined; the specified feature vector of the historical search results is used to characterize the semantics of the historical search results under the actual search intent.
[0028] The semantic vector of the historical search terms is input into the second sub-model. The second sub-model performs feature fusion on the semantic vector of the historical search terms and the first feature vector of the historical search results. The fused feature vector is used as the second feature vector of the historical search results. Specifically, this includes:
[0029] The semantic vector of the historical search terms and the specified feature vector of the historical search results are input into the feature fusion layer in the second sub-model. The feature fusion layer performs feature fusion on the semantic vector of the historical search terms and the specified feature vector of the historical search results, and the fused feature vector is used as the second feature vector of the historical search results.
[0030] Optionally, before training the semantic similarity determination model based on the predicted semantic similarity and the actual semantic similarity, the method further includes:
[0031] The feature vectors of the preset search intents of each type are input into the first sub-model, and the predicted search intent of the historical search terms output by the first sub-model is obtained based on the feature vectors of the preset search intents of each type and the semantic vectors of the historical search terms output by the first semantic vector layer in the first sub-model.
[0032] Based on the predicted semantic similarity and the actual semantic similarity, the semantic similarity determination model to be trained is trained, specifically including:
[0033] Based on the difference between the actual search intent of the historical search terms and the predicted search intent of the historical search terms, as well as the difference between the predicted semantic similarity and the actual semantic similarity, at least one of the first sub-model, the second sub-model, and the third sub-model in the semantic similarity determination model to be trained is trained.
[0034] This specification provides a method for determining semantic similarity, including:
[0035] Get the user's input search terms and search results;
[0036] The search term is input into the first sub-model of the semantic similarity determination model trained according to the above model training method, and the semantic vector of the search term is obtained through the first sub-model;
[0037] Each search result is used as input to the second sub-model of the semantic similarity determination model trained by the above model training method. The feature extraction layer in the second sub-model determines the first feature vector of each search result.
[0038] The semantic vector of the search term is input into the second sub-model. The second sub-model performs feature fusion on the semantic vector of the search term and the first feature vector of each search result. The fused feature vector is used as the second feature vector of each search result.
[0039] The semantic similarity between each search result and the search term is determined based on the similarity between the semantic vector of the search term and the second feature vector of each search result.
[0040] Based on the semantic similarity between the search term and each search result, the search results are returned to the user.
[0041] Optionally, before obtaining the semantic vector of the search term through the first sub-model, the method further includes:
[0042] Determine that the semantic vectors of each pre-stored historical search term do not contain the semantic vector of the search term;
[0043] The method further includes:
[0044] If the semantic vector of each pre-stored historical search term contains the semantic vector of the search term, the semantic vector of the pre-stored historical search term is used as the semantic vector of the search term; wherein, the semantic vector of the pre-stored historical search term is obtained by inputting the historical search term into the first sub-model and by the first sub-model.
[0045] Optionally, after obtaining the semantic vector of the search term through the first sub-model, the method further includes:
[0046] Obtain the search frequency of the search term;
[0047] When the search frequency of the search term is higher than a preset search frequency threshold, the search term and its corresponding semantic vector are stored.
[0048] Optionally, before determining the first feature vector of each search result by the feature extraction layer in the second sub-model, the method further includes:
[0049] Determine that the first feature vector of each pre-stored historical search result does not contain the first feature vector of the search result;
[0050] The method further includes:
[0051] If the first feature vector of each pre-stored historical search result contains the semantic vector of the search result, the first feature vector of the pre-stored historical search result is used as the first feature vector of the search result; wherein, the first feature vector of the pre-stored historical search result is obtained by inputting the historical search result into the second sub-model and by the feature extraction layer in the first sub-model.
[0052] Optionally, after the feature extraction layer in the second sub-model determines the first feature vector of each search result, the method further includes:
[0053] Obtain the frequency of the search results;
[0054] When the frequency of obtaining the search result is higher than a preset frequency threshold, the search result and the first feature vector corresponding to the search result are stored.
[0055] This specification provides a model training apparatus, including:
[0056] The first acquisition module is used to acquire historical search results, historical search terms, and the actual semantic similarity between the historical search results and the historical search terms.
[0057] The first semantic vector determination module is used to input the historical search terms into the first sub-model of the semantic similarity determination model to be trained, and obtain the semantic vector of the historical search terms through the first sub-model.
[0058] The first feature vector determination module is used to input the historical search results as input to the second sub-model of the semantic similarity determination model to be trained, and the feature extraction layer in the second sub-model determines the first feature vector of the historical search results.
[0059] The second feature vector determination module is used to input the semantic vector of the historical search term into the second sub-model, and to perform feature fusion on the semantic vector of the historical search term and the first feature vector of the historical search result through the second sub-model, and use the fused feature vector as the second feature vector of the historical search result;
[0060] The predicted semantic similarity determination module is used to determine the predicted semantic similarity between the historical search results and the historical search terms based on the similarity between the semantic vector of the historical search terms and the second feature vector of the historical search results.
[0061] The training module is used to train the semantic similarity determination model to be trained based on the predicted semantic similarity and the actual semantic similarity.
[0062] This specification provides a semantic similarity determination device, including:
[0063] The second acquisition module is used to acquire the user's input search terms and search results;
[0064] The second semantic vector determination module is used to input the search term into the first sub-model of the semantic similarity determination model trained according to any of the above model training methods, and obtain the semantic vector of the search term through the first sub-model;
[0065] The third feature vector determination module is used to take each search result as input and input it into the second sub-model of the semantic similarity determination model trained by any of the above model training methods, and the feature extraction layer in the second sub-model determines the first feature vector of each search result respectively.
[0066] The fourth feature vector determination module is used to input the semantic vector of the search term into the second sub-model, and to perform feature fusion on the semantic vector of the search term and the first feature vector of each search result through the second sub-model, and use the fused feature vector as the second feature vector of each search result;
[0067] A semantic similarity determination module is used to determine the semantic similarity between each search result and the search term based on the similarity between the semantic vector of the search term and the second feature vector of each search result;
[0068] The search results return module is used to return each search result to the user based on the semantic similarity between the search term and each search result.
[0069] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described model training or semantic similarity determination method.
[0070] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described model training or semantic similarity determination method.
[0071] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0072] This method inputs historical search results into the second sub-model of the semantic similarity determination model to be trained. The feature extraction layer of the second sub-model obtains the first feature vector of the historical search results. Then, the semantic vector of the historical search terms and the first feature vector of the historical search results are fused to obtain the second feature vector of the historical search results. This determines the predicted semantic similarity between the historical search results and the historical search terms, and the semantic similarity determination model is trained accordingly. It is evident that by introducing the semantic vector of the historical search terms during the generation of the second feature vector of the historical search results, the crossover between the semantics of the historical search terms and the semantics of the historical search results is considered, allowing for sufficient interaction between the historical search terms and historical search results during model training, thus improving the accuracy of semantic similarity determination. Attached Figure Description
[0073] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0074] Figure 1 This is a flowchart illustrating one model training method described in this specification.
[0075] Figure 2 This is a schematic diagram of a semantic similarity determination model in this specification;
[0076] Figure 3 This is a flowchart illustrating one semantic similarity determination method described in this specification.
[0077] Figure 4 This is a schematic diagram of a semantic similarity determination model in this specification;
[0078] Figure 5 This is a schematic diagram of a model training device provided in this specification;
[0079] Figure 6 This is a schematic diagram of a semantic similarity determination device provided in this specification;
[0080] Figure 7 The corresponding information provided in this specification Figure 1 or Figure 2 A schematic diagram of an electronic device. Detailed Implementation
[0081] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0082] Additionally, it should be noted that all actions involving the acquisition of signals, information, or data in this invention are carried out in compliance with the relevant data protection laws and regulations of the country where the invention is located, and with authorization from the owner of the corresponding device.
[0083] Search platforms can recommend several search results to users based on the search terms they input. Information search platforms can display a list of search results related to the user's input terms. The semantic similarity between search terms and search results is one of the core factors in information search scenarios, and the accuracy of semantic similarity directly affects the effectiveness of information search. Determining the semantic relevance of text can be considered as calculating the degree of similarity between search terms and search content at the semantic level, that is, determining whether the content of the search results meets the user's search needs. Determining the relevance between search terms and search content is one of the essential functions of a search system.
[0084] Currently, commonly used methods for calculating semantic relevance mainly include representation-based matching and interaction-based matching. The key difference lies in their approaches: interaction-based matching involves pre-interacting the semantic features of the search term and the search document at the neural network level, resulting in a better semantic feature vector representation and a semantic matching score; while representation-based matching uses a deep learning model to separately determine the semantic vectors of the search term and the search result, and then determines the matching score based on these vectors. Each method has its advantages. Representation-based matching allows for offline pre-computation of the search result's semantic vector, requiring only recalculation of the search term's semantic vector during online prediction. However, it suffers from a lack of interaction between the search term and the search result during model learning, failing to fully utilize the fine-grained matching signals between them. Interaction-based matching, on the other hand, enables thorough interaction and matching between the search term and the search result during model training, resulting in better semantic matching performance. Its disadvantage is higher deployment costs.
[0085] Fine-grained matching of search terms and search results refers to the semantically related matching relationship between search terms and search results. For example, consider the relationship between the search term "causes of pneumonia" and the search results "what causes lung inflammation" and "what are the symptoms of lung inflammation." In a coarse-grained matching scenario, the search term is about medical information retrieval, and both search results are also medical-related. In this scenario, the search term is semantically related to both search results. However, in a fine-grained scenario, the search term is about "pneumonia" and "causes," while the search result "what are the symptoms of lung inflammation" is about symptoms, and only the search result "what causes lung inflammation" is about causes. In this case, the search term "causes of pneumonia" is not semantically related to the search result "what are the symptoms of lung inflammation," but is semantically related to the search result "what causes lung inflammation."
[0086] In order to combine the advantages of the two existing methods while overcoming their disadvantages, this specification proposes a method for determining the semantic vector of search results using the semantic vector of search terms. This method enables the semantic vector of search results to interact not only with the semantic features of search terms, but also to improve the accuracy of determining the semantic similarity between search terms and search results.
[0087] Furthermore, the determination of semantic similarity between search terms and search results can be applied not only to information search but also to intelligent question-answering scenarios, where intelligent question-answering systems provide answers matching user questions, and information flow recommendation scenarios, where recommendation systems recommend information matching user characteristics (profiles). This specification does not limit the specific scenarios in which the semantic similarity model is applied in the embodiments. For ease of explanation, the application of the semantic similarity model to determining the semantic similarity between search terms and search results is used as an example to elaborate on the specific technical solutions.
[0088] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0089] Figure 1 This is a flowchart illustrating one model training method described in this specification, which specifically includes the following steps:
[0090] S100: Obtain historical search results, historical search terms, and the actual semantic similarity between the historical search results and the historical search terms.
[0091] Generally, information search platforms can determine the semantic similarity between search terms and search results using a semantic similarity determination model, and provide this semantic similarity to downstream modules, such as search result ranking models, so that search results semantically related to the search terms are ultimately displayed to users to meet their information search needs. This specification provides a method for training a semantic similarity determination model, which can be executed by a server used for training the model.
[0092] In practical applications, information search platforms can return multiple search results semantically related to the user's input search terms. In the embodiments of this specification, to train the semantic similarity determination model, historical search terms input by the user, one or more historical search results returned by the information search platform based on those historical search terms, and the actual semantic similarity between the historical search terms and the historical search results can be obtained. The actual semantic similarity between historical search terms and historical search results can be determined through manual annotation, user historical behavior (clicks, favorites), etc., and this specification does not limit this determination.
[0093] Furthermore, depending on the application scenario, historical search terms can also be historical search statements. This manual does not restrict historical search terms to words; they can also be single characters or sentences containing multiple words. Historical search results can take different forms depending on the specific type of information provided by the information search platform. For example, on a shopping platform, historical search results can be merchants and products; on a service provider platform, historical search results can be services and the merchants providing those services.
[0094] S102: Input the historical search terms into the first sub-model of the semantic similarity determination model to be trained, and obtain the semantic vector of the historical search terms through the first sub-model.
[0095] The semantic vector of historical search terms can be understood as a character vector or word vector used to represent the semantics of historical search terms. In practical applications, there may be cases where historical search terms are long. In this case, the semantic vector of historical search terms can be generated by combining the syntactic relationships between the words contained in the historical search terms. There may also be cases where historical search terms are short. If there are no surrounding words, the semantic vector of historical search terms can be generated by combining the meaning of each character contained in the historical search terms with the meaning of adjacent characters.
[0096] Specifically, the first sub-model contains at least a first semantic vector layer, which can be used to extract the semantic vectors of historical search terms. By inputting historical search terms into the first sub-model, the semantic vectors of the historical searches output by the first semantic vector layer in the first sub-model can be obtained.
[0097] The first semantic vector layer can effectively represent a word as a vector based on a given corpus. The generated semantic vector of the historical search term has the semantic features of the search term, and the syntactic and semantic features of the historical search term are distributed in each dimension of the semantic vector of the historical search term. The first semantic vector layer used in the embodiments of this specification can be a structure such as Continuous Bag-Of-Words Model (CBOW), Continuous Skip-gram Model (Skip-gram), or Bidirectional Encoder Representation from Transformers (BERT).
[0098] S104: The historical search results are used as input to the second sub-model of the semantic similarity determination model to be trained, and the first feature vector of the historical search results is determined by the feature extraction layer in the second sub-model.
[0099] Specifically, the semantic similarity determination model to be trained may include a second sub-model, and the feature extraction layer in the second sub-model can be used to output a first feature vector of the historical search results based on the historical search results input to the second sub-model. The first feature vector of the historical search results can characterize the semantics of the historical search results.
[0100] Optionally, in order to enable the first feature vector to not only represent the semantics of historical search results, but also features such as the search intent reflected in the semantics of historical search results, other feature vectors can be introduced when generating the first feature vector to make fuller use of the fine-grained matching signals of search results, so as to make the accuracy of the trained semantic similarity determination model higher.
[0101] S106: Input the semantic vector of the historical search terms into the second sub-model, and perform feature fusion on the semantic vector of the historical search terms and the first feature vector of the historical search results through the second sub-model. Use the fused feature vector as the second feature vector of the historical search results.
[0102] Because representation-based similarity calculation methods lack interaction between search terms and search results during model learning, the accuracy of similarity calculation is low. To address this issue, the model training method provided in this specification introduces the semantic vector of historical search terms into the generation process of the second feature vector of historical search results, enabling interaction between historical search terms and historical search results, and making full use of the fine-grained matching signals between historical search terms and historical search results.
[0103] Specifically, the feature fusion layer in the second sub-model can perform feature fusion on the semantic vector of the historical search terms and the specified feature vector of the historical search results using existing vector fusion schemes, such as adding point by point or concatenating vectors, etc. This specification does not limit this.
[0104] It's important to explain that the interaction between historical search terms and historical search results can be considered as the intersection of the semantic features of historical search terms and the semantic features of historical search results, resulting in cross-features between them. This interaction can be applied more fully to scenarios where a single historical search term corresponds to multiple historical search results or vice versa, allowing for more accurate similarity assessments.
[0105] S108: Determine the predicted semantic similarity between the historical search results and the historical search terms based on the similarity between the semantic vector of the historical search terms and the second feature vector of the historical search results.
[0106] Generally, the similarity calculation between the semantic vector of a historical search term and the second feature vector of a historical search result can be considered as calculating the distance between them: the smaller the distance, the higher the similarity; conversely, the larger the distance, the lower the similarity. In the embodiments of this specification, the calculation method used to calculate the similarity between the semantic vector of a historical search term and the second feature vector of a historical search result can be any existing method, such as Pearson correlation coefficient, Euclidean distance, etc., and this specification does not limit this method.
[0107] As can be seen, the first sub-model in the semantic similarity determination model to be trained in the embodiments of this specification includes at least a first semantic vector layer for determining the semantic vector of historical search terms, and the output of the first sub-model is actually the predicted search intent of the historical search terms. The second sub-model in the semantic similarity determination model includes at least a second semantic vector layer for outputting the semantic vector of historical search results, a feature extraction layer for outputting the first feature vector of historical search results, and a feature fusion layer for outputting the second feature vector of historical search results. In fact, the output of the second sub-model is the predicted semantic similarity between historical search results and historical search terms, such as... Figure 2 As shown.
[0108] S110: Train the semantic similarity determination model to be trained based on the predicted semantic similarity and the actual semantic similarity.
[0109] In practical applications, after determining the historical search terms, historical search results, and the actual semantic similarity between historical search results of historical search terms, the server used to train the semantic similarity determination model will adjust the model parameters of the semantic similarity determination model based on the difference between the predicted semantic similarity output by the second sub-model of the semantic similarity model to be trained and the actual semantic similarity between historical search results of historical search terms.
[0110] Additionally, it should be explained that since the semantic similarity determination model trained by the model training method provided in the embodiments of this specification includes at least a first sub-model and a second sub-model, when adjusting the model parameters of the semantic similarity determination model based on the difference between the actual semantic similarity of historical search results between historical search terms and the predicted semantic similarity output by the second sub-model, the model parameters of at least one of the first and second sub-models can be adjusted.
[0111] It is evident that by introducing the semantic vector of historical search terms during the generation of feature vectors for historical search results, the interaction between the semantic features of historical search terms and the semantic features of historical search results is considered. This allows historical search terms, historical search results, and search intent to fully interact during model training, thereby improving the accuracy of semantic similarity determination.
[0112] In an optional embodiment of this specification, in such a way... Figure 1 Before determining the first feature vector of the historical search results by the feature extraction layer in the second sub-model as shown in step S104, multiple preset search intents can be obtained. These preset search intents are then introduced into the feature extraction layer of the second sub-model used to determine the first feature vector of the historical search results. This ensures that the feature vector of the historical search results contains cross-features between the historical search results and the preset search intents of various types, further improving the accuracy of semantic similarity calculation. The specific steps are as follows:
[0113] First, obtain the preset search intents for each type.
[0114] In the embodiments described in this specification, the types of search intent may include the following three:
[0115] The first type is precise search intent. Precise search intent can be characterized by the literal relevance between historical search results and historical search terms. Literal relevance can be seen as the titles or content of historical search results containing words contained in historical search terms. For example, in a service search scenario, the historical search term is "Gaoxing Hot Pot Restaurant". From the historical search term, we can analyze the complete name of the restaurant, that is, the user precisely searched for the restaurant name. In this case, the search intent of the historical search term can be classified as precise search intent. We search for historical search results in the search results database that are semantically related to "Gaoxing Hot Pot Restaurant". This can be historical search results whose titles contain the historical search term "Gaoxing Hot Pot Restaurant", such as "Gaoxing Hot Pot Restaurant (Wudaokou Branch)" and "Gaoxing Hot Pot Restaurant (Xueyuan Road Branch)".
[0116] The second type is broad search intent. Broad search intent can be characterized by search results matching the service category corresponding to the search term. Matching the service category means that the service type corresponding to the historical search results is the same as or similar to the service category corresponding to the historical search term. For example, in a service search scenario, the historical search term is "hot pot". From the historical search term, we can analyze that the user did not explicitly enter the specific name of the service provider (merchant) nor did it involve the location where the user wanted to receive the service. Therefore, the search intent of the historical search term can be classified as broad search intent. Analyzing the historical search term "hot pot", the service category corresponding to the historical search term is "catering". In fine-grained service category classification, "hot pot" can also be directly used as a service category, and historical search results related to "hot pot" or "catering" can be found in the search results database.
[0117] The third type is location-based search intent. Location-based search intent can be characterized by historical search results matching the geographical region indicated by the historical search terms. Matching the geographical region means that the location of the historical search results is the same as or similar to the geographical region indicated by the historical search terms. For example, in a service search scenario, the historical search term is "Wudaokou." Analysis of the historical search term shows that "Wudaokou" corresponds to a region, but the historical search term does not explicitly indicate the service the user needs. Therefore, historical search results containing "Wudaokou" in their titles or content can be returned to the user. Alternatively, the location information corresponding to each historical search result in the search results database can be retrieved, and historical search results with the location information "Wudaokou" can also be returned to the user.
[0118] Secondly, the preset search intents of each type are input into the third sub-model of the semantic similarity determination model to be trained, so as to obtain the feature vectors of the preset search intents of each type output by the third sub-model.
[0119] Specifically, the third sub-model in the semantic similarity determination model extracts the intent features of various preset search intent types, resulting in feature vectors for each type of search intent. These feature vectors characterize the search intent features of each type of search intent, such as precise search intent characterized by literal meaning, broad search intent characterized by service category, and address search intent characterized by geographic location.
[0120] Then, the historical search results are input into the second sub-model, and the semantic vector of the historical search results is output by the second semantic vector layer in the second sub-model.
[0121] To obtain the semantics representing historical search results under various search intents, the semantic features of historical search results can first be extracted using the second semantic vector layer in the second sub-model, resulting in a semantic vector of the historical search results. The structure adopted by the second semantic vector layer in the second sub-model is similar to... Figure 1 The first semantic vector layer in the first sub-model in step S102 is similar, except that the objects for which the generated vectors are targeted are different, which will not be described in detail here.
[0122] Optionally, historical search results can be segmented into characters or words, and the semantic vectors corresponding to each character (or word) in the historical search results can be generated using the second semantic vector layer in the second sub-model. Then, the semantic vectors of the historical search results can be determined according to the specific application scenario.
[0123] Furthermore, for each type of search intent, the semantic vector of the historical search results and the feature vector of the search intent of that type are input into the feature extraction layer corresponding to the search intent of that type in the second sub-model to obtain the first feature vector output by the feature extraction layer corresponding to the search intent of that type; the first feature vector output by the feature extraction layer corresponding to the search intent of that type is used to characterize the semantics of the historical search results under the search intent of that type.
[0124] Since the search intents corresponding to the historical search results obtained during the model training phase are different, especially in cases where one historical search result corresponds to multiple search intents, and the semantic vector of the historical search result only reflects the semantic meaning of the historical search result (title or content) and not the search intent, the feature extraction layer corresponding to each type of search intent in the second sub-model can use an attention mechanism to fuse the preset feature vectors of each type of search intent with the semantic vector of the historical search result to obtain the first feature vector of the historical search result. This allows the first feature vector of the historical search result to reflect not only the semantics of the historical search result but also the semantics of the historical search result under each type of search intent.
[0125] In one optional embodiment of this specification, such as Figure 1In step S106, the semantic vector of the historical search term is input into the second sub-model. Based on the semantic vector of the historical search term, feature fusion is performed on the first feature vector of the historical search result. Before using the fused feature vector as the second feature vector of the historical search result, the actual search intent corresponding to the historical search term can be determined. Then, the actual search intent corresponding to the determined historical search term is introduced into the process of determining the second feature vector of the historical search result. Preset search intents of various types that do not match the actual search intent of the historical search term can be eliminated, so as to prevent the second feature vector of the historical search result from representing the semantics that do not belong to the search intent of the historical search term, thereby affecting the determination of the predicted semantic similarity between the historical search terms and the historical search results.
[0126] In practical applications, some information search platforms can provide users with timely and regional services that include various business categories. By obtaining the search terms entered by users, analyzing and mining the users' search intent, connecting them to the business categories that the users intend to search for, and then returning the services (goods) or service providers (merchants) that match the users' search intent to the users.
[0127] Typically, information search platforms display search results related to user search terms by analyzing and mining the user's search intent from the semantic level of the search terms they input. The goal is to present search results that match the user's search intent. Therefore, the semantic similarity between search terms and search results directly reflects the display effect of the results page. Existing, relatively accurate interaction-based similarity calculation methods also focus on the interaction between the semantic features of historical search terms and the semantic features of historical search results. However, in certain scenarios, combining similarity calculation methods based on the user's search intent can provide more accurate search results. For example, if a user searches for "apple" on a service provider platform, the literal meaning of "apple" can refer to "fruits and vegetables," leading to service categories like food delivery. However, based on the extended meaning of "apple," it can also refer to Apple mobile phones, leading to the category of electronic products. Analyzing and mining the user's search intent from their input search terms can provide more accurate search results.
[0128] It is evident that analyzing and mining users' search intent from historical search terms and incorporating this intent into the training process of the semantic similarity determination model can enhance the influence of historical search intent on the model. This allows the pre-trained semantic similarity determination model to consider the actual search intent reflected in historical search terms when determining the semantic similarity between historical search terms and historical search results. Consequently, search results with high semantic similarity are more aligned with users' search intent, thereby improving the effectiveness of information retrieval.
[0129] The specific steps for determining the second feature vector of historical search results by incorporating the actual search intent corresponding to the identified historical search terms into the process are as follows:
[0130] First, the historical search terms are input into a pre-trained search intent recognition model to obtain the actual search intent of the historical search terms output by the search intent recognition model.
[0131] The actual search intent of the historical search terms mentioned in the embodiments of this specification can be determined based on deep query understanding (DQU). DQU mainly utilizes basic natural language processing (NLP) capabilities to analyze and mine historical search terms, and combines them with different business logics of the information search platform to determine at least one actual search intent corresponding to the historical search terms.
[0132] The actual search intent reflected by historical search terms can be classified in three ways: precise search intent (search results are literally related to the search term), broad search intent (search results match the service category corresponding to the search term), and location search intent (search results match the region corresponding to the search term). This application specification only uses the aforementioned three classifications as examples to illustrate specific technical solutions, and does not imply that the classification of the actual search intent of historical search terms is limited to these three.
[0133] Secondly, based on the first feature vectors of the historical search results and the actual search intent of the historical search terms, a specified feature vector of the historical search results is determined; the specified feature vector of the historical search results is used to characterize the semantics of the historical search results under the actual search intent.
[0134] Specifically, the first feature vectors of historical search results are determined based on the feature vectors of preset search intents for each type. In other words, the first feature vectors of historical search results represent the semantics of the historical search results under each type of search intent. Since the actual search intent of historical search terms is not entirely the same as the preset search intents for each type, to ensure that the search intent reflected in the final determined second feature vectors of historical search results closely matches the actual search intent of the search terms, a filtering process is performed using the actual search intent of the historical search terms: the first feature vectors of historical search results represent the semantics of historical search results under each type of search intent. Based on the actual search intent of historical search terms, the first feature vectors that can represent the semantics of historical search results under the actual search intent of historical search terms are retained, and these retained first feature vectors of historical search results are used as the designated feature vectors of historical search results. Therefore, the designated feature vectors of historical search results can represent the intersection between the semantics of historical search results and the actual search intent.
[0135] Then, the semantic vectors of historical search terms and the specified feature vectors of historical search results are input into the feature fusion layer of the second sub-model. The feature fusion layer fuses the semantic vectors of the historical search terms and the specified feature vectors of the historical search results, and the fused feature vector is used as the second feature vector of the historical search results. The feature fusion method used by the feature fusion layer to fuse the semantic vectors of the historical search terms and the specified feature vectors of the historical search results is the same as described above. Figure 1 Step S106 is similar and will not be repeated here.
[0136] It is evident that the second feature vector of historical search results not only includes the semantic features of the historical search results themselves, but also the semantic features of historical search terms, the actual search intent features of historical search terms, and the intent features of preset search intent. Compared with existing representation-based similarity calculation methods, the second feature vector of historical search results includes the interaction with historical search terms at both the semantic and actual search intent levels. By fully utilizing the fine-grained matching signals of historical search terms and historical search results, the second feature vector of historical search results becomes more accurate, thereby making the trained semantic similarity determination model more accurate in calculating the similarity between search terms and search results.
[0137] Optionally, the search intent recognition model can output at least one actual search intent corresponding to historical search terms. The representation of the actual search intent of historical search terms can be a label of the actual search intent corresponding to the historical search terms, or a probability distribution of the actual search intent. The actual search intent label can be in the form of a one-hot code. Based on preset candidate search intents and the search intents corresponding to the search terms output by the search intent recognition model, the search intent label is determined: when the search intent output by the model matches a candidate search intent, the label of the candidate search intent is 1; when the search intent output by the model does not match a candidate search intent, the label of the candidate search intent is 0. For example, if the search intents determined by the search intent recognition model are address search intent and general search intent, and the preset search intents are precise search intent, address search intent, and general search intent, then the search intent labels corresponding to the user-input search terms can be determined as precise search intent (0), address search intent (1), and general search intent (1).
[0138] Of course, in the embodiments of this specification, the actual search intent of historical search terms can also be determined by manual annotation and obtaining the user's historical behavior (clicks, favorites) for historical search results related to historical search terms. This specification does not limit this.
[0139] In the embodiments described in this specification, such as Figure 1 Step S110, which trains the semantic similarity determination model based on the predicted semantic similarity and the actual semantic similarity, is specifically implemented through the following steps.
[0140] First, the feature vectors of each type of preset search intent are input into the first sub-model. Based on the feature vectors of each type of preset search intent and the semantic vectors of the historical search terms output by the first semantic vector layer in the first sub-model, the predicted search intent of the historical search terms output by the first sub-model is obtained.
[0141] The predicted search intent of historical search terms can be considered as the probability of the search intent represented by the semantics of historical search terms, based on the semantic features of historical search terms and the features of various types of search intent.
[0142] Secondly, the first loss is determined based on the difference between the actual search intent of the historical search terms and the predicted search intent of the historical search terms.
[0143] Furthermore, the weight of the first loss is determined based on the pre-set importance of the search intent, wherein the importance of the search intent is used to characterize the degree of importance of each type of pre-set search intent to the determination of semantic similarity.
[0144] The weight of the first loss is determined by the importance of the pre-set search intent to the semantic similarity calculation. Generally, during the training process of the semantic similarity determination model, the weight of the first loss can be adjusted in real time according to the training effect of the model. Usually, the weight of the first loss is relatively small to avoid the search intent having too much influence on the semantic similarity calculation.
[0145] The loss function used to determine the loss can be the cross-entropy function or any other existing loss function. This specification does not limit this. Taking the cross-entropy function as an example, the specific function representation of the first loss can be as follows:
[0146] Loss1=∑ i -(α i *log(β i )+(1-α i )*log(1-β i ))
[0147] Where, α i β is used to characterize the actual search intent of historical search terms determined by the search intent recognition model. i Predicted search intent used to characterize historical search terms.
[0148] Then, a second loss is determined based on the difference between the predicted semantic similarity and the actual semantic similarity. The functional representation of the second loss can be shown below:
[0149] Loss2=-(A r *log(B r )+(1-A r )*log(1-B r ))
[0150] Among them, A r Used to characterize the actual semantic similarity between search terms and search results, B r Used to characterize predicted semantic similarity.
[0151] Finally, based on the first loss and the second loss, the semantic similarity of the model to be trained is used to determine at least one of the first sub-model, the second sub-model, and the third sub-model in the model.
[0152] The total loss is determined based on the first loss, the weight corresponding to the first loss, and the second loss. The functional representation of the total loss can be shown below:
[0153] Loss = λLoss1 + Loss2
[0154] Where Loss is the total loss, λ is the weight of the first loss, Loss1 is the first loss, and Loss2 is the second loss.
[0155] With the goal of minimizing the total loss, one sub-model from the first and second sub-models of the semantic similarity determination model is trained. The total loss includes a second loss characterizing the difference between the predicted semantic similarity and the actual semantic similarity, and a first loss characterizing the difference between the predicted search intent of historical search terms and the actual search intent of historical search terms.
[0156] The training objective of the semantic similarity determination model provided in the embodiments of this specification is that the predicted semantic similarity between the search results and the search terms is closer to the actual semantic similarity, and the predicted search intent of historical search terms is closer to the actual search intent of historical search terms determined by the search intent recognition model.
[0157] Optionally, to ensure the accuracy of the feature vectors of various types of search intents used when generating the second feature vector of historical search results, the model parameters of the third sub-model can be determined by adjusting the semantic similarity of the feature vectors of various types of search intents generated based on the difference between the actual search intent of the historical search terms and the predicted search intent of the historical search terms determined by the feature vectors of various types of search intents (i.e., the first loss), until the first loss is a preset loss value.
[0158] In an optional embodiment of this specification, historical search terms and at least one historical search result can be used as training sample pairs. The actual semantic similarity between the historical search terms and historical search results is the labeling of the training sample pairs. The constructed training sample pairs can consist of one historical search term and at least one historical search result, or a combination of multiple historical search terms and at least one historical search result. Since historical search results (positive samples) related to historical search terms constitute only a small portion of the overall search result database in information search tasks—that is, most historical search results have a significant semantic distance from the historical search terms—if there are many irrelevant historical search results (negative sample pairs) corresponding to historical search terms during the training of the semantic relevance determination model, the trained model may perform poorly. Therefore, the ratio of negative sample pairs to positive sample pairs can be reasonably determined to construct the training sample pairs.
[0159] It is understood that one or more optional embodiments provided in this specification can be arbitrarily combined to obtain a scheme having at least one of the following: introducing the semantic vector of historical search terms into the second feature vector of historical search results; introducing the actual search intent of historical search terms into the second feature vector of historical search results; and introducing the feature vectors of preset search intents of various types into the first feature vector of historical search results.
[0160] This specification also provides an embodiment of a method based on Figure 1 The model training method trains a semantic similarity determination model, which determines the semantic similarity between the user's input search terms and each search result. This semantic similarity is then passed to downstream modules (such as the information ranking module) so that the information search platform displays search results with high semantic similarity to the search terms and matching the user's search intent. Figure 3 As shown, the specific steps are as follows:
[0161] S200: Obtain the user-inputted search terms and search results.
[0162] In practical applications, users can input their search terms either by typing them into a designated location on the information search page (e.g., typing search terms into the input field) or by clicking on a designated display area on the information search page (e.g., clicking on historical search terms in the historical search term display area). This manual does not specify any particular method for users to input search terms.
[0163] S202: Input the search term into the button. Figure 1 The method trains a first sub-model of the semantic similarity determination model, and obtains the semantic vector of the search term through the first sub-model.
[0164] It's important to understand that during model training, the output of the first sub-model in the semantic similarity determination model is the predicted search intent of historical search terms. This predicted search intent is used to determine the first loss against the actual search intent of the historical search terms, and then the first loss is used to train the third sub-model. Therefore, the output of the first sub-model is only used to train the third sub-model. However, when using the trained semantic similarity determination model to determine the semantic similarity between search terms and search results, the semantic vector of the search term can be output solely from the first semantic vector in the first sub-model, without requiring the first sub-model to output anything. Figure 4 As shown.
[0165] S204: Input each of the search results as input to the... Figure 1 The second sub-model of the semantic similarity determination model trained by the method determines the first feature vector of each search result by the feature extraction layer in the second sub-model.
[0166] Specifically, for each obtained search result, the search result is input into the second sub-model. The second semantic vector layer of the second sub-model outputs the semantic vector of the search result. Then, based on the semantic vector of the search result, the feature extraction layer in the second model determines the first feature vector of the search result. Optionally, preset search intents of various types can be introduced to determine the feature extraction layer corresponding to each type of search intent in the second model. The feature vectors of each type of search intent and the semantic vector of the search result are respectively input into the feature extraction layer corresponding to each type of search intent, thereby obtaining the first feature vector of the search result output by the feature extraction layer corresponding to each type of search intent.
[0167] S206: Input the semantic vector of the search term into the second sub-model, and perform feature fusion on the semantic vector of the search term and the first feature vector of each search result through the second sub-model, and use the fused feature vector as the second feature vector of each search result.
[0168] Optionally, preset search intents of various types can also be obtained. The third sub-model of the semantic similarity determination model outputs feature vectors of the preset search intents of various types, and these feature vectors are also input into the feature fusion layer of the second sub-model to obtain the second feature vector of the search results. The specific process of determining the second feature vector of the search results is as follows: Figure 1 Step S106 is similar and will not be repeated here.
[0169] S208: Determine the semantic similarity between each search result and the search term based on the similarity between the semantic vector of the search term and the second feature vector of each search result.
[0170] Understandably, when using a pre-trained semantic similarity determination model to determine the semantic similarity between search terms and search results, the input to the second sub-model can be search results, feature vectors of various types of search intent, semantic vectors of search terms, etc. The second sub-model includes at least a second semantic vector layer, a feature extraction layer, and a feature fusion layer. Finally, the second sub-model outputs the semantic similarity between the search results and the search terms, such as... Figure 4 As shown.
[0171] S210: Based on the semantic similarity between the search term and each of the search results, return each search result to the user.
[0172] Depending on the specific application scenario, you can choose to directly return search results with a semantic similarity to the search term that is higher than a preset threshold to the user, or you can send the semantic similarity between the search term and each search result to a downstream module, such as a search result ranking model. The downstream module can then perform further operations such as ranking the search results based on the semantic similarity between the search term and each search result before returning them to the user, in order to improve the accuracy and personalization of information retrieval.
[0173] Optionally, since information search scenarios require ensuring the real-time return of search results, to improve real-time performance, one can, for example... Figure 3 Before obtaining the semantic vector of the search term through the first sub-model as shown in step S202, the following steps are performed:
[0174] Determine whether the semantic vector of each pre-stored historical search term contains the semantic vector of the search term.
[0175] If the semantic vector of each pre-stored historical search term contains the semantic vector of the search term, the semantic vector of the pre-stored historical search term is used as the semantic vector of the search term; wherein, the semantic vector of the pre-stored historical search term is obtained by inputting the historical search term into the first sub-model and by the first sub-model.
[0176] If the semantic vectors of the pre-stored historical search terms do not contain the semantic vector of the search term, then as follows: Figure 3 As shown in step S202, the semantic vector of the search term is obtained through the first sub-model.
[0177] Furthermore, for a search term whose semantic vector is obtained through the first sub-model, the search frequency of the search term is obtained. If the search frequency of the search term is higher than a preset search frequency threshold, the search term and its corresponding semantic vector are stored.
[0178] It should be noted that the pre-stored search terms can be historical search terms and their semantic vectors as a search term set. Based on the search terms entered by the user, the search term set is searched, and the semantic vector of the search term is determined based on the search results.
[0179] Specifically, determining whether the semantic vectors of each pre-stored historical search term contain the semantic vector of the search term can be further divided into three cases:
[0180] In the first scenario, if the user-inputted search term is found in the search term set, the semantic vector of the search term can be directly obtained from the pre-stored semantic vector of the historical search terms, which greatly saves the time of generating the semantic vector of the search term.
[0181] The second scenario: If the user-inputted search term is not found in the search term set, the semantic vector of the search term is obtained through the first sub-model.
[0182] The third scenario: If some of the user-inputted search terms are found in the search term set, then the semantic vectors of some words of the search terms not found in the pre-stored search term set are determined through the first sub-model. Based on the partial words of the search terms found in the pre-stored search term set, the semantic vectors corresponding to the partial words of the found search terms are determined. Then, based on the semantic vectors of the partial words of the search terms not found in the pre-stored search term set and the semantic vectors corresponding to the partial words of the found search terms, the semantic vector of the user-input search term is determined.
[0183] Optionally, since information search scenarios require ensuring the real-time return of search results, to improve real-time performance, one can, for example... Figure 3 Before the feature extraction layer in the second sub-model determines the first feature vector of each search result as shown in step S204, the following steps are performed:
[0184] Determine whether the first feature vector of each pre-stored historical search result contains the first feature vector of the search result.
[0185] If the first feature vector of each pre-stored historical search result contains the first feature vector of the search result, the first feature vector of the pre-stored historical search result is used as the first feature vector of the search result; wherein, the first feature vector of the pre-stored historical search result is obtained by inputting the historical search result into the second sub-model and by the feature extraction layer in the first sub-model.
[0186] If the first feature vector of each historical search result stored in advance does not contain the first feature vector of the search result, then the search result is used as input to the second sub-model of the pre-trained semantic similarity determination model, and the first feature vector of the search result is determined by the feature extraction layer in the second sub-model.
[0187] Optionally, after the feature extraction layer in the second sub-model determines the first feature vector of each search result, the acquisition frequency of the search result can be obtained, and when the acquisition frequency of the search result is higher than a preset acquisition frequency threshold, the search result and the first feature vector corresponding to the search result are stored.
[0188] As can be seen, the semantic similarity determination model provided in this specification, which uses a pre-trained semantic similarity determination model to calculate the semantic similarity between the user-input search terms and the search results, greatly improves the real-time performance of information retrieval because it pre-stores the first feature vector of the search results and the semantic vector of the search terms. It improves both the accuracy of semantic similarity calculation and the real-time performance of returning search results based on search terms, thus combining the advantages of expression-based semantic similarity calculation methods and interaction-based semantic similarity calculation methods.
[0189] The above describes one or more embodiments of the model training method and semantic similarity determination method provided in this specification. Based on the same idea, this specification also provides corresponding model training devices and semantic similarity determination devices, such as... Figure 5 , Figure 6 As shown.
[0190] Figure 5 A schematic diagram of a model training device provided in this specification specifically includes:
[0191] The first acquisition module 300 is used to acquire historical search results, historical search terms, and the actual semantic similarity between the historical search results and the historical search terms.
[0192] The first semantic vector determination module 302 is used to input the historical search terms into the first sub-model of the semantic similarity determination model to be trained, and obtain the semantic vector of the historical search terms through the first sub-model.
[0193] The first feature vector determination module 304 is used to take the historical search results as input and input them into the second sub-model of the semantic similarity determination model to be trained, and the feature extraction layer in the second sub-model determines the first feature vector of the historical search results.
[0194] The second feature vector determination module 306 is used to input the semantic vector of the historical search term into the second sub-model, and perform feature fusion on the semantic vector of the historical search term and the first feature vector of the historical search result through the second sub-model, and use the fused feature vector as the second feature vector of the historical search result;
[0195] The predicted semantic similarity determination module 308 is used to determine the predicted semantic similarity between the historical search results and the historical search terms based on the similarity between the semantic vector of the historical search terms and the second feature vector of the historical search results.
[0196] Training module 310 is used to train the semantic similarity determination model to be trained based on the predicted semantic similarity and the actual semantic similarity.
[0197] Optionally, the first sub-model includes a first semantic vector layer;
[0198] Optionally, the first semantic vector determination module 302 is specifically used to input the historical search terms into the first sub-model to obtain the semantic vector of the historical search terms output by the first semantic vector layer in the first sub-model.
[0199] Optionally, the first feature vector determination module 304 is further configured to: obtain preset search intents of various types before the first feature vector determination module 304 inputs the historical search results as input to the second sub-model of the semantic similarity determination model to be trained, and the feature extraction layer in the second sub-model determines the first feature vector of the historical search results; and input the preset search intents of various types into the third sub-model of the semantic similarity determination model to be trained, respectively, to obtain the feature vectors of the preset search intents of various types output by the third sub-model.
[0200] Optionally, the first feature vector determination module 304 is specifically used to output the first feature vector of the historical search results by the feature extraction layer in the second sub-model based on the preset feature vectors of various types of search intents and the historical search results.
[0201] Optionally, the first feature vector determination module 304 is specifically configured to: input the historical search results into the second sub-model to obtain the semantic vector of the historical search results output by the second semantic vector layer in the second sub-model; for each type of search intent, input the semantic vector of the historical search results and the feature vector of the search intent of that type as inputs into the feature extraction layer corresponding to the search intent of that type in the second sub-model to obtain the first feature vector output by the feature extraction layer corresponding to the search intent of that type; the first feature vector output by the feature extraction layer corresponding to the search intent of that type is used to characterize the semantics of the historical search results under the search intent of that type.
[0202] Optionally, the second feature vector determination module 306 is further configured to: before the second feature vector determination module 306 inputs the semantic vector of the historical search term into the second sub-model, performs feature fusion on the semantic vector of the historical search term and the first feature vector of the historical search result through the second sub-model, and uses the fused feature vector as the second feature vector of the historical search result, input the historical search term into a pre-trained search intent recognition model to obtain the actual search intent of the historical search term output by the search intent recognition model; determine a specified feature vector of the historical search result based on each first feature vector of the historical search result and the actual search intent of the historical search term; the specified feature vector of the historical search result is used to characterize the semantics of the historical search result under the actual search intent;
[0203] Optionally, the second feature vector determination module 306 is specifically used to input the semantic vector of the historical search term and the specified feature vector of the historical search result as input into the feature fusion layer in the second sub-model, and perform feature fusion on the semantic vector of the historical search term and the specified feature vector of the historical search result through the feature fusion layer, and use the fused feature vector as the second feature vector of the historical search result.
[0204] Optionally, the training module 310 is further configured to, before training the semantic similarity determination model to be trained based on the predicted semantic similarity and the actual semantic similarity, input the feature vectors of the preset search intentions of each type into the first sub-model, and obtain the predicted search intention of the historical search term output by the first sub-model based on the feature vectors of the preset search intentions of each type and the semantic vector of the historical search term output by the first semantic vector layer in the first sub-model;
[0205] Optionally, the training module 310 is specifically used to train at least one of the first sub-model, the second sub-model, and the third sub-model in the semantic similarity determination model to be trained, based on the difference between the actual search intent of the historical search terms and the predicted search intent of the historical search terms, as well as the difference between the predicted semantic similarity and the actual semantic similarity.
[0206] Figure 6 This specification provides a schematic diagram of a semantic similarity determination device, which specifically includes:
[0207] The second acquisition module 400 is used to acquire the search terms entered by the user and each search result;
[0208] The second semantic vector determination module 402 is used to input the search term into the first sub-model of the semantic similarity determination model trained according to the above model training method, and obtain the semantic vector of the search term through the first sub-model.
[0209] The third feature vector determination module 404 is used to take each search result as input and input it into the second sub-model of the semantic similarity determination model trained according to the above model training method, and the feature extraction layer in the second sub-model determines the first feature vector of each search result respectively.
[0210] The fourth feature vector determination module 406 is used to input the semantic vector of the search term into the second sub-model, and to perform feature fusion on the semantic vector of the search term and the first feature vector of each search result through the second sub-model, and to use the fused feature vector as the second feature vector of each search result;
[0211] The semantic similarity determination module 408 is used to determine the semantic similarity between each search result and the search term based on the similarity between the semantic vector of the search term and the second feature vector of each search result;
[0212] The search result return module 410 is used to return each search result to the user based on the semantic similarity between the search term and each search result.
[0213] Optionally, the second semantic vector determination module 402 is further configured to determine, before the second semantic vector determination module 402 obtains the semantic vector of the search term through the first sub-model, that the semantic vector of each historical search term stored in advance does not contain the semantic vector of the search term.
[0214] Optionally, if the semantic vectors of the pre-stored historical search terms include the semantic vector of the search term, the semantic vectors of the pre-stored historical search terms are used as the semantic vector of the search term; wherein, the semantic vectors of the pre-stored historical search terms are obtained by inputting the historical search terms into the first sub-model and by the first sub-model.
[0215] Optionally, the second semantic vector determination module 402 is further configured to, after obtaining the semantic vector of the search term through the first sub-model, obtain the search frequency of the search term; and when the search frequency of the search term is higher than a preset search frequency threshold, store the search term and the semantic vector corresponding to the search term.
[0216] Optionally, the third feature vector determination module 404 is further configured to determine, before the third feature vector determination module 404 determines the first feature vector of each search result from the feature extraction layer in the second sub-model, that the first feature vector of each historical search result stored in advance does not contain the first feature vector of the search result.
[0217] Optionally, if the first feature vector of each pre-stored historical search result contains the first feature vector of the search result, the first feature vector of the pre-stored historical search result is used as the first feature vector of the search result; wherein, the first feature vector of the pre-stored historical search result is obtained by inputting the historical search result into the second sub-model and by the feature extraction layer in the first sub-model.
[0218] Optionally, the third feature vector determination module 404 is further configured to, after the third feature vector determination module 404 determines the first feature vector of each search result by the feature extraction layer in the second sub-model, obtain the acquisition frequency of the search result; when the acquisition frequency of the search result is higher than a preset acquisition frequency threshold, store the search result and the first feature vector corresponding to the search result.
[0219] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 Provided model training methods or Figure 3 The semantic similarity determination method is shown.
[0220] This instruction manual also provides Figure 7 The diagram shows a schematic structural representation of the electronic device. Figure 7 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1 The model training method or Figure 3 The semantic similarity determination method is shown. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0221] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0222] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0223] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0224] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0225] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0226] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0227] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0228] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0229] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0230] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0231] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0232] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0233] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0234] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0235] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0236] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A model training method, characterized in that, include: Obtain historical search results, historical search terms, and the actual semantic similarity between the historical search results and the historical search terms; The historical search terms are input into the first sub-model of the semantic similarity determination model to be trained, and the semantic vector of the historical search terms is obtained through the first sub-model. The historical search results are used as input to the second sub-model of the semantic similarity determination model to be trained, and the first feature vector of the historical search results is determined by the feature extraction layer in the second sub-model. The semantic vector of the historical search term is input into the second sub-model. The second sub-model performs feature fusion on the semantic vector of the historical search term and the first feature vector of the historical search result. The fused feature vector is used as the second feature vector of the historical search result. Based on the similarity between the semantic vector of the historical search term and the second feature vector of the historical search result, the predicted semantic similarity between the historical search result and the historical search term is determined; The semantic similarity determination model to be trained is trained based on the predicted semantic similarity and the actual semantic similarity.
2. The method as described in claim 1, characterized in that, The first sub-model includes a first semantic vector layer; The historical search terms are input into the first sub-model of the semantic similarity determination model to be trained, and the semantic vector of the historical search terms is obtained through the first sub-model, specifically including: The historical search terms are input into the first sub-model to obtain the semantic vector of the historical search terms output by the first semantic vector layer in the first sub-model.
3. The method as described in claim 2, characterized in that, The method further includes, before the historical search results are input into the second sub-model of the semantic similarity determination model to be trained, and before the feature extraction layer in the second sub-model determines the first feature vector of the historical search results, the method further includes: Obtain preset search intents of various types; The preset search intents of each type are input into the third sub-model of the semantic similarity determination model to be trained, and feature vectors of the preset search intents of each type are output by the third sub-model. The historical search results are used as input to the second sub-model of the semantic similarity determination model to be trained. The feature extraction layer in the second sub-model determines the first feature vector of the historical search results, specifically including: Based on the feature vectors of the preset search intents of each type and the historical search results, the feature extraction layer in the second sub-model outputs the first feature vector of the historical search results.
4. The method as described in claim 3, characterized in that, Based on the feature vectors of the preset search intents for each type and the historical search results, the feature extraction layer in the second sub-model outputs the first feature vector of the historical search results, specifically including: The historical search results are input into the second sub-model to obtain the semantic vector of the historical search results output by the second semantic vector layer in the second sub-model; For each type of search intent, the semantic vector of the historical search results and the feature vector of the search intent of that type are input into the feature extraction layer corresponding to the search intent of that type in the second sub-model to obtain the first feature vector output by the feature extraction layer corresponding to the search intent of that type; the first feature vector output by the feature extraction layer corresponding to the search intent of that type is used to characterize the semantics of the historical search results under the search intent of that type.
5. The method as described in claim 3, characterized in that, The method further includes inputting the semantic vector of the historical search terms into the second sub-model, fusing the semantic vector of the historical search terms and the first feature vector of the historical search results through the second sub-model, and using the fused feature vector as the second feature vector of the historical search results. The historical search terms are input into a pre-trained search intent recognition model to obtain the actual search intent of the historical search terms output by the search intent recognition model. Based on the first feature vectors of the historical search results and the actual search intent of the historical search terms, a specified feature vector of the historical search results is determined; the specified feature vector of the historical search results is used to characterize the semantics of the historical search results under the actual search intent. The semantic vector of the historical search terms is input into the second sub-model. The second sub-model performs feature fusion on the semantic vector of the historical search terms and the first feature vector of the historical search results. The fused feature vector is used as the second feature vector of the historical search results. Specifically, this includes: The semantic vector of the historical search terms and the specified feature vector of the historical search results are input into the feature fusion layer in the second sub-model. The feature fusion layer performs feature fusion on the semantic vector of the historical search terms and the specified feature vector of the historical search results, and the fused feature vector is used as the second feature vector of the historical search results.
6. The method as described in claim 5, characterized in that, Before training the semantic similarity determination model based on the predicted semantic similarity and the actual semantic similarity, the method further includes: The feature vectors of the preset search intents of each type are input into the first sub-model, and the predicted search intent of the historical search terms output by the first sub-model is obtained based on the feature vectors of the preset search intents of each type and the semantic vectors of the historical search terms output by the first semantic vector layer in the first sub-model. Based on the predicted semantic similarity and the actual semantic similarity, the semantic similarity determination model to be trained is trained, specifically including: Based on the difference between the actual search intent of the historical search terms and the predicted search intent of the historical search terms, as well as the difference between the predicted semantic similarity and the actual semantic similarity, at least one of the first sub-model, the second sub-model, and the third sub-model in the semantic similarity determination model to be trained is trained.
7. A method for determining semantic similarity, characterized in that, include: Get the user's input search terms and search results; The search term is input into the first sub-model of the semantic similarity determination model trained according to any one of claims 1 to 6, and the semantic vector of the search term is obtained through the first sub-model. Each search result is used as input to the second sub-model of the semantic similarity determination model trained according to any one of claims 1 to 6 above. The first feature vector of each search result is determined by the feature extraction layer in the second sub-model. The semantic vector of the search term is input into the second sub-model. The second sub-model performs feature fusion on the semantic vector of the search term and the first feature vector of each search result. The fused feature vector is used as the second feature vector of each search result. The semantic similarity between each search result and the search term is determined based on the similarity between the semantic vector of the search term and the second feature vector of each search result. Based on the semantic similarity between the search term and each search result, the search results are returned to the user.
8. The method as described in claim 7, characterized in that, Before obtaining the semantic vector of the search term through the first sub-model, the method further includes: Determine that the semantic vectors of each pre-stored historical search term do not contain the semantic vector of the search term; The method further includes: If the semantic vector of each pre-stored historical search term contains the semantic vector of the search term, the semantic vector of the pre-stored historical search term is used as the semantic vector of the search term; wherein, the semantic vector of the pre-stored historical search term is obtained by inputting the historical search term into the first sub-model and by the first sub-model.
9. The method as described in claim 8, characterized in that, After obtaining the semantic vector of the search term through the first sub-model, the method further includes: Obtain the search frequency of the search term; When the search frequency of the search term is higher than a preset search frequency threshold, the search term and its corresponding semantic vector are stored.
10. The method as described in claim 7, characterized in that, Before determining the first feature vector of each search result by the feature extraction layer in the second sub-model, the method further includes: Determine that the first feature vector of each pre-stored historical search result does not contain the first feature vector of the search result; The method further includes: If the first feature vector of each pre-stored historical search result contains the first feature vector of the search result, the first feature vector of the pre-stored historical search result is used as the first feature vector of the search result; wherein, the first feature vector of the pre-stored historical search result is obtained by inputting the historical search result into the second sub-model and by the feature extraction layer in the second sub-model.
11. The method as described in claim 10, characterized in that, After determining the first feature vector of each search result from the feature extraction layer in the second sub-model, the method further includes: Obtain the frequency of the search results; When the frequency of obtaining the search result is higher than a preset frequency threshold, the search result and the first feature vector corresponding to the search result are stored.
12. A model training device, characterized in that, include: The first acquisition module is used to acquire historical search results, historical search terms, and the actual semantic similarity between the historical search results and the historical search terms. The first semantic vector determination module is used to input the historical search terms into the first sub-model of the semantic similarity determination model to be trained, and obtain the semantic vector of the historical search terms through the first sub-model. The first feature vector determination module is used to input the historical search results as input to the second sub-model of the semantic similarity determination model to be trained, and the feature extraction layer in the second sub-model determines the first feature vector of the historical search results. The second feature vector determination module is used to input the semantic vector of the historical search term into the second sub-model, and to perform feature fusion on the semantic vector of the historical search term and the first feature vector of the historical search result through the second sub-model, and use the fused feature vector as the second feature vector of the historical search result; The predicted semantic similarity determination module is used to determine the predicted semantic similarity between the historical search results and the historical search terms based on the similarity between the semantic vector of the historical search terms and the second feature vector of the historical search results. The training module is used to train the semantic similarity determination model to be trained based on the predicted semantic similarity and the actual semantic similarity.
13. A semantic similarity determination device, characterized in that, include: The second acquisition module is used to acquire the user's input search terms and search results; The second semantic vector determination module is used to input the search term into the first sub-model of the semantic similarity determination model trained according to any one of claims 1 to 6, and obtain the semantic vector of the search term through the first sub-model; The third feature vector determination module is used to input each of the search results as input to the second sub-model of the semantic similarity determination model trained according to the method of any one of claims 1 to 6, and to determine the first feature vector of each of the search results by the feature extraction layer in the second sub-model. The fourth feature vector determination module is used to input the semantic vector of the search term into the second sub-model, and to perform feature fusion on the semantic vector of the search term and the first feature vector of each search result through the second sub-model, and use the fused feature vector as the second feature vector of each search result; A semantic similarity determination module is used to determine the semantic similarity between each search result and the search term based on the similarity between the semantic vector of the search term and the second feature vector of each search result; The search results return module is used to return each search result to the user based on the semantic similarity between the search term and each search result.
14. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 6 or 7 to 11.
15. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 6 or 7 to 11.
Citation Information
Patent Citations
Method and device for showing search result
CN103942279A
Training method and device for semantic similarity matching model
CN111460264A