Semantic correlation prediction model, method and device, storage medium and computer equipment
By introducing the student model and the online prompt model into the semantic relevance model and using the thought chain prompt information for feature extraction and fusion, the problem of insufficient accuracy of the existing model in complex semantic environments is solved, and efficient and accurate semantic relevance prediction is achieved.
Patent Information
- Application Number
- CN202511285148.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing semantic relevance models cannot effectively capture high-order interaction information and deep contextual associations between keywords in complex semantic environments, resulting in insufficient accuracy in semantic relevance prediction.
The student model obtained by distilling the teacher model is adopted, combined with a parameter-sharing dual-tower network, an independent thinking chain tower network and an expert hybrid network. The thinking chain prompt information is generated through the online prompt model, and feature extraction and fusion are performed to predict semantic relevance.
It improves the accuracy and efficiency of semantic relevance prediction, reduces computational overhead, enhances the ability to capture high-order contextual information, and improves the accuracy and robustness of relevance prediction.
Smart Images

Figure CN120804671A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet, and in particular to a semantic correlation prediction model, method, device, storage medium and computer equipment. BACKGROUND
[0002] In the field of information retrieval and natural language processing, semantic correlation models are widely used to measure the degree of semantic association between words, phrases or text segments. Such models play an important role in improving the performance of search engines, recommendation systems and intelligent question-answering systems. With the development of deep learning technology, neural network-based semantic matching methods have gradually become mainstream. This method can automatically learn semantic representations from large-scale data and provide more effective correlation judgment basis for downstream tasks.
[0003] Currently, common semantic correlation models generally use a double-tower structure or a double-encoder structure. This structure uses two parameter-shared or independent encoders to independently encode two input keywords, then maps them to fixed-length vector representations, and finally calculates the correlation score between the two keywords through the similarity between the vectors. However, when applying the above structure to complex semantic environments, the model often fails to capture high-order interaction information or deep contextual associations between keywords during feature extraction, resulting in insufficient expression of complex semantic relationships and severely affecting the accuracy of semantic correlation prediction. SUMMARY
[0004] Therefore, the embodiments of the present application provide a semantic correlation prediction model, method, device, storage medium and computer equipment, which mainly aims to solve the technical problem of low prediction accuracy of the semantic correlation prediction model.
[0005] According to a first aspect of the present application, a semantic correlation prediction model is provided, which includes a student model distilled by a teacher model and a pre-trained online prompt model, and the student model includes: a parameter-shared double-tower network for feature extraction of a first keyword and a second keyword to be predicted to obtain first keyword features and second keyword features; a thought chain tower network for feature extraction of thought chain prompt information generated by the online prompt model to obtain thought chain features, the thought chain tower network not sharing parameters with the double-tower network; an expert mixing network for feature fusion of the first keyword features, the second keyword features and the thought chain features to obtain keyword fusion features; The deep neural network is configured to perform semantic correlation prediction based on the keyword fusion feature to obtain a semantic correlation prediction result of the first keyword and the second keyword.
[0006] According to a second aspect of the present application, a semantic correlation prediction method is provided, which comprises: obtaining a first keyword and a second keyword to be predicted, and inputting the first keyword and the second keyword into the semantic correlation prediction model; performing correlation estimation on the first keyword and the second keyword by an online prompt model in the semantic correlation prediction model to obtain a correlation estimation reason and a correlation estimation result; generating thought chain prompt information based on the correlation estimation reason and the correlation estimation result, and inputting the thought chain prompt information into a student model in the semantic correlation prediction model; The student model performs correlation analysis on the first keyword and the second keyword based on the thought chain prompt information to obtain a correlation prediction result of the first keyword and the second keyword.
[0007] According to a third aspect of the present application, a semantic correlation prediction device is provided, which comprises: a keyword obtaining module configured to obtain a first keyword and a second keyword to be predicted, and input the first keyword and the second keyword into the semantic correlation prediction model; a correlation estimation module configured to perform correlation estimation on the first keyword and the second keyword by an online prompt model in the semantic correlation prediction model to obtain a correlation estimation reason and a correlation estimation result; a thought chain reasoning module configured to generate thought chain prompt information based on the correlation estimation reason and the correlation estimation result, and input the thought chain prompt information into a student model in the semantic correlation prediction model; a correlation prediction module configured to perform correlation analysis on the first keyword and the second keyword by the student model based on the thought chain prompt information to obtain a correlation prediction result of the first keyword and the second keyword.
[0008] According to a fourth aspect of the present application, a storage medium having a computer program stored thereon is provided, and the program is executed by a processor to implement the above-mentioned semantic correlation prediction method.
[0009] According to a fifth aspect of the present application, a computer device is provided, which comprises a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, and the processor implements the above-mentioned semantic correlation prediction method when executing the program.
[0010] By means of the technical solutions, the semantic correlation prediction model, method, device, storage medium and computer equipment provided by the embodiment of the application can reduce the performance overhead of the model when predicting the semantic correlation online by setting the student model distilled by the teacher model in the semantic correlation prediction model, thereby improving the prediction efficiency of the correlation. Moreover, by setting the online prompt model in the semantic correlation prediction model and generating the thought chain prompt information by using the online prompt model, the high-order context information between the keyword pairs can be extracted, thereby helping to improve the accuracy of the correlation prediction. In addition, by setting the parameter-shared double-tower structure and the parameter-independent thought chain tower network in the student model and setting the expert mixed network for feature fusion, the multi-modal features in the keyword pairs to be predicted and the thought chain prompt information can be accurately extracted, thereby significantly improving the capturing ability of the student model for the high-order context information, and thereby the accuracy of the correlation prediction can be enhanced while keeping low computational overhead.
[0011] The above description is only a summary of the technical solutions of the application. In order to enable the technical means of the application to be more clearly understood, the application can be implemented in accordance with the content of the description, and in order to enable the above and other purposes, features and advantages of the application to be more apparent and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings described herein are used to provide further understanding of the application, and form a part of the application. The schematic embodiments of the application and the description thereof are used to explain the application, and do not constitute an improper limitation on the application. In the drawings: Figure 1 A model architecture schematic diagram of a student model in a semantic correlation prediction model provided by an embodiment of the application is shown; Figure 2 A model architecture schematic diagram of a student model in another semantic correlation prediction model provided by an embodiment of the application is shown; Figure 3 A model architecture schematic diagram of a semantic correlation prediction model provided by an embodiment of the application during training is shown; Figure 4 A flow schematic diagram of a semantic correlation prediction method provided by an embodiment of the application is shown; Figure 5 A flow schematic diagram of another semantic correlation prediction method provided by an embodiment of the application is shown; Figure 6 A structure schematic diagram of a semantic correlation prediction device provided by an embodiment of the application is shown. DETAILED DESCRIPTION
[0013] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.
[0014] Currently, common semantic relevance models generally adopt a dual-tower or dual-encoder architecture. These structures use two parameter-sharing or independent encoders to independently encode two input keywords, then map them into fixed-length vector representations. Finally, the relevance score between the two keywords is calculated by calculating the similarity between the vectors. However, when these structures are applied to complex semantic environments, the models are often unable to capture high-level interactions or deep contextual associations between keywords during the feature extraction phase. This results in insufficient representation of complex semantic relationships, which seriously affects the accuracy of semantic relevance prediction.
[0015] In order to solve the above problems, in one embodiment, Figure 1 As shown, a semantic relevance prediction model is provided, which includes a student model obtained by distillation of a teacher model and a pre-trained online prompt model 20, wherein the student model includes a parameter-sharing dual-tower network 11, a thinking chain tower network 12 that does not share parameters with the dual-tower network 11, an expert hybrid network 13, and a deep neural network 14. The parameter-sharing dual-tower network 11 can be used to perform feature extraction on the first keyword and the second keyword to be predicted, thereby obtaining first keyword features and second keyword features; the thinking chain tower network 12 can be used to perform feature extraction on the thinking chain prompt information generated by the online prompt model, thereby obtaining thinking chain features; the expert hybrid network 13 can be used to perform feature fusion on the first keyword features, the second keyword features, and the thinking chain features, thereby obtaining keyword fusion features; and the deep neural network 14 can be used to perform semantic relevance prediction based on the keyword fusion features, thereby obtaining semantic relevance prediction results for the first keyword and the second keyword.
[0016] The semantic correlation prediction model refers to a model for predicting the correlation between two keywords; the online prompt model refers to a pre-trained language model that can automatically generate structured reasoning prompts according to the input keywords, and is run online in real time, which can guide the student model to perform deeper semantic analysis; the student model refers to a lightweight model learned from a teacher model (such as a large language model) with stronger performance but higher computational complexity through knowledge distillation technology, which can significantly reduce reasoning delay and resource consumption while maintaining high prediction accuracy, and is suitable for online real-time search scenarios. In this embodiment, the student model can learn the output of the teacher model by minimizing the output distribution of the teacher model, thereby obtaining reasoning ability close to the teacher model. Through the above method, the understanding ability of the student model for complex semantics can be retained, and the requirements of low delay and high concurrency for data processing can be met.
[0017] Specifically, referring to Figure 1 , the student model can perform feature extraction on the first keyword and the second keyword to be predicted through the parameter-shared double-tower network 11, and at the same time, the thought chain prompt information generated by the online prompt model 20 can be extracted through the parameter-independent thought chain tower network 12 to generate thought chain features reflecting the logical differences in word types. Subsequently, the student model can dynamically fuse the above three types of feature vectors through the expert mixed network 13 (MoE, Mixture of Experts), wherein the expert mixed network can select different expert models (such as specification matching experts, semantic similarity experts, and category matching experts) according to the input feature difference, and perform weighted calculation and fusion, thereby obtaining keyword fusion features. Finally, the student model can perform semantic correlation prediction based on the fused features through the deep neural network 14, and output the semantic correlation prediction result of the first keyword and the second keyword.
[0018] In this embodiment, the thought chain prompt information generated by the online prompt model can provide higher-order context information (such as specification difference and semantic association) compared to only inputting the first keyword and the second keyword to be predicted, therefore, the student model needs to introduce additional parameters to effectively extract and learn such knowledge. Based on this, referring to Figure 1The network structure of the post-interaction type ReprBERT of the student model can be improved, that is, an independent COT tower (i.e., a thinking chain tower network 12) is added to the original double-tower network 11 (a Query tower and an Item tower for processing original first keywords and second keywords) to model the thinking chain prompt information and form a three-tower structure. The output feature vectors of the three towers can be fused through an expert mixing network 13, wherein the expert mixing network 13 includes multiple expert models with independent parameters, and a suitable expert combination can be dynamically selected to weight and fuse the three types of features, so as to obtain a more accurate representation. In this embodiment, the newly added COT tower does not share parameters with the double tower, which can ensure that the network can focus on the extraction of high-order logical features, and the expert mixing network can strengthen the collaborative expression of multi-modal features through expert-level dynamic interaction, thereby improving the processing capability of the student model for complex semantic relationships.
[0019] The semantic correlation prediction model provided in the above embodiment can reduce the performance overhead of the model when predicting semantic correlation online by setting a student model distilled by a teacher model, thereby improving the prediction efficiency of the correlation. Moreover, by setting an online prompt model in the model and generating thinking chain prompt information by using the online prompt model, high-order context information between keyword pairs can be extracted, thereby helping to improve the accuracy of correlation prediction. In addition, the above model can accurately extract multi-modal features in the keyword pairs to be predicted and the thinking chain prompt information generated by the online prompt model by setting a double-tower structure with parameter sharing and a thinking chain tower network with parameter independence in the student model, and setting an expert mixing network for feature fusion, thereby significantly improving the ability of the student model to capture high-order context information, and thereby enhancing the accuracy of correlation prediction while maintaining low computational overhead.
[0020] In one embodiment, as shown in Figure 2 In the student model, the deep neural network 14 can perform the following operations: obtaining explicit cross features and keyword fusion features output by the expert mixing network 13, wherein the explicit cross features can be constructed based on multiple features in the text features, category features and entity features in the first keywords and the second keywords, and then performing semantic correlation prediction according to the explicit cross features and the keyword fusion features to obtain the semantic correlation prediction result of the first keywords and the second keywords.
[0021] Specifically, referring to Figure 2The deep neural network 14 can realize the prediction of semantic relevance by joint features of explicit cross features and keyword fusion features. The explicit cross features can be generated by cross combination of various features in the text features, category features and entity features in the first keyword and the second keyword, and are used to capture the explicit association between the keywords. Then, the explicit cross features and the keyword fusion features (containing implicit high-order context information) output by the expert hybrid network 13 are jointly input into the deep neural network 14, and the cross-modal feature interaction is extracted through multi-layer nonlinear transformation in the deep neural network 14, and finally the semantic relevance prediction result of the first keyword and the second keyword is output. In this embodiment, the explicit cross features can make up for the insufficient modeling of local association of the double tower structure, and at the same time, by means of the fusion ability of the expert hybrid network, the modeling depth of global semantic association can be enhanced.
[0022] The above embodiment can significantly improve the accuracy of semantic relevance prediction by cooperative modeling of explicit cross features and keyword fusion features. Especially in the processing of complex context and long-tail association scenarios, compared with the traditional double tower structure model, the accuracy and robustness of semantic relevance prediction can be effectively improved. At the same time, by introducing explicit cross features, the dependence of the model on the training data distribution can be effectively reduced.
[0023] In one embodiment, as shown in FIG. 1, the semantic relevance prediction model can include an online prompt model 20, a student model 21 and a teacher model 22. The online prompt model 20 can be used to assist the student model 21 to perform semantic relevance prediction by dynamically generating a thought chain prompt information. The teacher model 22 can be used to supervise the student model 21 by providing a thought chain prompt information to the student model 21. Figure 1 Figure 2 As shown in FIG. 1, the online prompt model 20 can perform the following operations: when the student model performs semantic relevance prediction, the first keyword and the second keyword are obtained, and the relevance of the first keyword and the second keyword is estimated to obtain the relevance estimation reason and the relevance estimation result, then the thought chain prompt information is generated based on the relevance estimation reason and the relevance estimation result, and the thought chain prompt information is input into the student model.
[0024] Specifically, the online prompt model can assist the student model to perform semantic relevance prediction by dynamically generating a thought chain prompt information. In this embodiment, when the student model receives the first keyword and the second keyword to be predicted, the online prompt model can synchronously obtain the two keywords, estimate the relevance of the two keywords based on the preset rules set in the model, and then output the text description containing the association basis (the relevance estimation reason) and the preliminary prediction result (the relevance estimation result). Subsequently, the online prompt model can construct a structured thought chain prompt information according to the estimation reason and the estimation result, such as outputting “the first keyword and the second keyword have high relevance because the entity words and the modifier words are matched”, and inputting the thought chain prompt information as an additional input into the student model, so that the student model can combine the explicit reasoning logic and the original features to perform more accurate semantic relevance calculation during decoding.
[0025] The above embodiment can effectively enhance the understanding ability of the student model for the keyword association logic by dynamically generating the thinking chain prompt information by using the online prompt model and injecting it into the student model for relevance prediction, thereby improving the explainability and accuracy of the relevance prediction result while maintaining the lightweight of the model.
[0026] In one embodiment, referring to Figure 3 , the semantic relevance prediction model further includes a pre-trained teacher model and an offline prompt model, wherein the offline prompt model is configured to generate sample thinking chain prompt information based on sample keyword pairs and relevance labels of the sample keyword pairs in a sample data set, the teacher model is configured to receive the sample keyword pairs and the sample thinking chain prompt information as input, and generate a semantic relevance prediction result, and the student model is trained based on the semantic relevance prediction result output by the teacher model and the relevance labels of the sample keyword pairs in the process of knowledge distillation to obtain a trained student model.
[0027] Specifically, referring to Figure 3 , the semantic relevance prediction model can include an offline teacher model (left side) and an online student model (right side). The teacher model, as a large model with a large number of parameters, can include four inputs: [CLS], Query (first keyword), Item (second keyword), and COT (thinking chain prompt information), wherein the COT is generated by the offline prompt model and processed by the BERT module. Subsequently, the Teacher Logits output by the teacher model can be transmitted to the student model by distillation. The loss function of the model includes Logits distillation loss (used to measure the gap between the outputs of the student model and the teacher model) and real loss (used to measure the gap between the output of the student model and the real label). Further, the online student model can be designed to be lightweight and include three inputs: Query (first keyword), Item (second keyword), and COT (thinking chain prompt information), wherein the COT is generated by the online prompt model and processed by the ReprBERT module of the student model, which can support online inference.
[0028] In this embodiment, if the student model lacks COT input (i.e., the thinking chain prompt information generated by the online prompt model) when performing online reasoning, the relevance prediction effect of the student model will decrease. Based on this, the online prompt model is used to dynamically complete the COT information (thinking chain prompt information) online, which can improve the prediction effect. In the knowledge distillation process, the student model can continuously optimize the output Student Logits through the mixed loss function, so as to continuously approach the Teacher Logits. In this embodiment, the soft target of the teacher model can be used to train the student model, so that the student model can imitate the output distribution of the teacher model, rather than just the classification label. Through the above method, the student model can not only learn the correct category, but also learn the confidence of the teacher model for each category. In addition, in order to improve the online efficiency, the student model can also reduce the calculation cost through parameter sharing, and at the same time, a smaller model size can be used to realize fast reasoning. When training the student model, a mixed loss function can be used, which can consider both the soft target of the teacher model and the actual hard label, so as to ensure that the student model can not only imitate the decision boundary of the teacher model, but also can be correctly classified.
[0029] In one embodiment, the student model can be trained by the following method: first, based on the sample keyword pairs in the sample data set and the sample thinking chain prompt information generated by the offline prompt model, the teacher model is pre-trained, and the parameters of the teacher model are frozen after the training is completed, then the teacher model with frozen parameters is used to predict the semantic relevance of the sample keyword pairs, and the prediction result output by the teacher model is temperature scaled to obtain a soft target. The soft target can be generated by a temperature scaled softmax function, and the temperature T is an adjustable hyperparameter. When T>1, the output distribution is smoother and contains more information about the relative probability of each category. Further, the student model can be iteratively trained based on the same sample keyword pairs and sample thinking chain prompt information generated by the pre-trained online prompt model, wherein the loss function of the student model can include the distillation loss between the temperature scaled output result of the student model and the soft target generated by the teacher model, and the cross entropy loss between the output result of the student model and the relevance label. Finally, when the training reaches a preset number of rounds or the performance of the student model on the validation set reaches a preset convergence target, the model training can be stopped, and the trained student model can be obtained.
[0030] Specifically, refer to Figure 3In training the student model, first, a sample dataset containing a large number of sample keyword pairs can be read, and then a teacher model is trained offline based on the sample dataset, and the parameters of the teacher model are frozen after the training is completed. For example, in a search scenario, the input of the teacher model can be a search keyword such as "five batteries", a commodity keyword such as "battery 7 / 4 section", and a thought chain prompt information generated by the offline prompt model such as "the modifier of the search keyword 【5】 is different from the modifier of the commodity keyword 【7】, which belongs to the weak correlation caused by specification difference". Subsequently, the teacher model can process the above features through the BERT module and output a correlation prediction result, and then the output prediction result is temperature scaled to obtain a soft target. Further, the student model can be iteratively trained based on the same sample keyword pairs and sample thought chain prompt information generated by the pre-trained online prompt model as input under the constraint of a hybrid loss function. The loss function can include a distillation loss between the correlation prediction result (soft target) output by the teacher model and the correlation prediction result (temperature scaled) output by the student model, and a cross-entropy loss between the prediction result of the student model and the labeled result. When the hybrid loss converges to a preset range, the training can be stopped to form an online deployable student model. The parameters of the student model are independent of the parameters of the teacher model, which can support real-time inference to reduce the computational cost.
[0031] The above embodiment can make the student model inherit the discriminative ability of the teacher model for complex semantic relationships while keeping low computational overhead, thereby improving the accuracy and robustness of search correlation prediction in the online inference stage, by setting an offline teacher model and an online student model in the semantic correlation prediction model, and cooperatively training the two models, while modeling the multi-modal features of the thought chain prompt information, and adopting a hybrid optimization strategy of distillation loss and cross-entropy loss.
[0032] In one embodiment, the online prompt model can be trained by the following method: loading a pre-trained offline prompt model, mapping the weight parameters of the offline prompt model to an integer interval based on preset quantization parameters, and rounding or truncating the mapped weight parameters, then converting the activation values of neurons from floating-point numbers to integers during each forward propagation of the offline prompt model, and finally performing quantization-aware training on the offline prompt model with adjusted weight parameters and activation values to obtain a trained online prompt model.
[0033] Specifically, referring to Figure 3The online prompt model can be realized by quantization compression of the offline prompt model, and the core target is to greatly reduce the storage requirement and calculation complexity of the model without significant loss of prediction performance. The main process of quantization processing includes: first, loading the pre-trained offline prompt model, and mapping the model weight parameters from floating point numbers to the integer interval based on the preset quantization parameters (such as the integer interval [-128, 127]), and then eliminating the floating point precision redundancy by rounding or truncation. Subsequently, in each forward propagation process, the activation value of the neuron is dynamically converted from a floating point number to an integer representation (such as an 8-bit integer representation) in real time to simulate the low-precision calculation environment in the actual inference scenario. Through the above method, the floating point number weight and activation value in the model can be converted to an integer representation that occupies fewer bits. On this basis, in order to alleviate the performance degradation caused by quantization, quantization aware training (QAT, Quantization Aware Training) can be used to optimize the model end-to-end, that is, inserting a pseudo-quantization node in the training stage, so that the model learns the quantization error compensation mechanism in the back propagation, and finally outputs an online prompt model that can adapt to low-precision operation and maintain semantic integrity. Through the above method, the quantization effect can be simulated during model training, so that the model learns to optimize the performance under the quantization condition.
[0034] The above embodiments can compress the storage space of the online prompt model and improve the inference speed while ensuring high-precision semantic expression capability by introducing weight quantization and dynamic activation quantization in the quantization process of the online prompt model, and combining the error compensation mechanism of quantization aware training. In addition, through the above method, the online prompt model can be compatible with the hardware acceleration instruction set of the offline prompt model without additional calibration data when deployed, thereby significantly reducing the resource consumption of the online inference of the model.
[0035] In one embodiment, a semantic correlation prediction method is provided, which is applied to a computer device such as a server, for example, as shown in Figure 4 The method includes the following steps: Step 101, obtaining a first keyword and a second keyword to be predicted, and inputting the first keyword and the second keyword into a pre-trained semantic correlation prediction model.
[0036] The pre-trained semantic correlation prediction model refers to a joint architecture including an online prompt model and a student model, and the model architecture can refer to any of the above embodiments, which will not be described here.
[0037] Specifically, the server can obtain the first keyword and the second keyword to be predicted through an API interface and / or user input information, and after encapsulating the original texts of the two keywords as structured input vectors, synchronously transmit the input vectors to the input layer of the semantic relevance prediction model to ensure that the model can start the prediction process based on complete context features.
[0038] In step 102, the relevance of the first keyword and the second keyword is estimated by an online hint model in the semantic relevance prediction model, and an estimation reason and an estimation result are obtained.
[0039] Specifically, after receiving the first keyword and the second keyword, the online hint model can splice the input keyword pair in a specific format and take the spliced features as the input of the online hint model. Then, the online hint model can generate a natural language hint containing the reasoning process and the reasoning result based on the input keyword pair, i.e., output the estimation reason and the estimation result of the keyword pair.
[0040] In step 103, based on the estimation reason and the estimation result, a thought chain hint information is generated and input into a student model in the semantic relevance prediction model.
[0041] The thought chain hint information refers to the intermediate text containing reasoning logic generated by the online hint model, which can simulate the thinking process of human beings when judging whether two keywords are relevant.
[0042] Specifically, the online hint model can generate the thought chain hint information based on the estimation reason and the estimation result of the first keyword and the second keyword, and input the thought chain hint as additional input together with the original keyword pair into the student model. Through the above method, the abstract relevance judgment can be converted into a traceable reasoning process, thereby significantly enhancing the transparency and logic of the decision-making of the semantic relevance prediction model, and helping to improve the accuracy of the relevance judgment under complex or fuzzy queries.
[0043] In step 104, the student model analyzes the relevance of the first keyword and the second keyword based on the thought chain hint information, and obtains a relevance prediction result of the first keyword and the second keyword.
[0044] The relevance prediction result refers to the quantitative output of the student model on the relevance between the first keyword and the second keyword under the guidance of the thought chain hint information, which can be a continuous value (such as a relevance score between 0 and 1) or a discrete category (such as “strong correlation”, “weak correlation” and “no correlation”), which can reflect the final judgment of the student model after comprehensively considering the semantics, attributes and reasoning logic.
[0045] Specifically, after receiving the first keyword, the second keyword and the thinking chain prompt information, the student model can perform feature extraction and fusion on the above three texts, and then perform relevance judgment based on the fused features. In this embodiment, the structure of the student model can adopt a double-tower neural network and an independent neural network, and combine a hybrid expert network, wherein the double-tower neural network can be used to encode the first keyword and the second keyword respectively, the independent neural network can be used to encode the thinking chain prompt information, and the hybrid expert network can fuse the above three features for subsequent relevance prediction. In the reasoning process, the student model can not only pay attention to the matching of the surface meaning of the keywords, but also adjust the attention weight according to the logical clues in the thinking chain prompt information, so as to suppress the misjudgment caused by the matching of part of the word items. Through the above-mentioned manner, high-quality and interpretable relevance prediction can be completed on a lightweight student model, so as to balance the relationship between model performance and computational efficiency, and thus improve the accuracy of the search results.
[0046] The semantic relevance prediction method provided by the above embodiment can extract high-order context information between the keyword pair by generating the thinking chain prompt information of the keyword pair to be predicted by the online prompt model, thereby helping to improve the accuracy of relevance prediction. Moreover, the above method can reduce the performance overhead of the model when predicting semantic relevance online by using the student model distilled by the teacher model, thereby improving the prediction efficiency of relevance. In addition, the above method can significantly improve the ability of the student model to capture high-order context information by accurately extracting multi-modal features in the keyword pair to be predicted and the thinking chain prompt information generated by the online prompt model, thereby improving the accuracy of relevance prediction while maintaining low computational overhead.
[0047] In one embodiment, a semantic relevance prediction method in a search scenario is provided. In the search scenario, the first keyword and the second keyword can be a search keyword and a commodity keyword respectively. The above method is applied to a computer device such as a server, for example, as shown in Figure 5 The method comprises the following steps: Step 201, obtaining a search keyword, and based on the search keyword, matching to obtain at least one commodity keyword, and inputting the search keyword and the commodity keyword into a semantic relevance prediction model.
[0048] The search keyword refers to a word or phrase input by a user in a search box of a search engine or e-commerce platform to express information needs or purchase intentions. The commodity keyword refers to a word or phrase closely associated with a commodity on the platform, used to describe the attributes, categories, specifications or uses of the commodity, which can be derived from the commodity title, label, attribute field or category information, and is the basic unit of information matching and retrieval.
[0049] Specifically, the computer device can obtain the text content input by the user in real time through the front-end user interface, that is, obtain the search keyword input by the user, and perform standardization processing on the obtained search keyword, such as removing spaces, correcting spelling errors, synonym normalization, etc., to improve the accuracy of keyword matching. Then, information search can be performed based on the processed search keyword, and corresponding commodity keywords can be found in the commodity keyword library based on the searched commodity information and merchant information. After obtaining the search keyword and the commodity keyword, they can be input into the semantic correlation prediction model. Through the above method, candidate commodity keywords potentially related to the user query information can be preliminarily screened out from a large amount of commodity information.
[0050] Step 202, the relevance of the search keyword and the commodity keyword is estimated by the online prompt model in the semantic correlation prediction model, and the relevance estimation reason and the relevance estimation result are obtained.
[0051] Specifically, after inputting the above keyword pair into the semantic correlation prediction model, the model can splice the input keyword pair in a specific format and use it as the input of the online prompt model. Further, the online prompt model can generate a natural language prompt containing the reasoning process and the reasoning result based on the input keyword pair, that is, output the relevance estimation reason and the relevance estimation result of the search keyword and the commodity keyword.
[0052] Step 203, based on the relevance estimation reason and the relevance estimation result, generate the thought chain prompt information, and input the thought chain prompt information into the student model in the semantic correlation prediction model.
[0053] The thought chain prompt information refers to the intermediate text containing reasoning logic generated by the online prompt model, which can simulate the thinking process of humans when judging whether two keywords are related. For example, first identify the parts of speech of the two keywords, then compare the differences between the two keywords from multiple angles, and finally draw a conclusion. This structured prompt information can guide the student model to perform more in-depth and interpretable relevance analysis.
[0054] Specifically, after receiving the search keyword and the product keyword, the online prompting model can generate the thought chain prompt information by using a preset inference step. For example, the online prompting model can make a relevance judgment based on a preset rule template, can make a relevance judgment based on the free inference logic of the generative model, or can dynamically generate a relevance judgment result in combination with an external knowledge base. Then, the online prompting model can generate the thought chain prompt information based on the relevance inference basis and the relevance inference result, and input the thought chain prompt as additional input together with the original word pair into the student model. Through the above method, the abstract relevance judgment can be converted into a traceable inference process, thereby significantly enhancing the transparency and logic of the semantic relevance prediction model decision, and helping to improve the accuracy of the relevance judgment under complex or fuzzy queries.
[0055] In step 204, the student model performs relevance analysis on the search keyword and the product keyword based on the thought chain prompt information, and obtains a relevance prediction result of the search keyword and the product keyword.
[0056] The relevance prediction result refers to the quantitative output of the relevance between the search keyword and the product keyword by the student model under the guidance of the thought chain prompt information, which can be a continuous value (such as a relevance score between 0 and 1) or a discrete category (such as “strong correlation”, “weak correlation”, “no correlation”), and the result can reflect the final judgment of the student model after considering the semantics, attributes and inference logic.
[0057] Specifically, after receiving the search keyword, the product keyword and the thought chain prompt information, the student model can perform feature extraction and fusion on the above three kinds of text, and then make a relevance judgment based on the fused features. In this embodiment, the structure of the student model can adopt a double-tower neural network and an independent neural network, and combine a hybrid expert network, wherein the double-tower neural network can be used to encode the search keyword and the product keyword respectively, the independent neural network can be used to encode the thought chain prompt information, and the hybrid expert network can fuse the above three features for subsequent relevance prediction. In the inference process, the student model can not only focus on the matching of the surface meaning of the keywords, but also adjust the attention weight according to the logical clues in the thought chain prompt information, so as to suppress the misjudgment caused by the matching of part of the word items. Through the above method, high-quality and interpretable relevance prediction can be completed on a lightweight student model, so as to balance the relationship between model performance and computing efficiency, and thereby improve the accuracy of the search result.
[0058] In step 205, based on the relevance prediction result, a search result corresponding to the search keyword is obtained.
[0059] The search result refers to a list of commodities or a list of merchants presented to the user by the search system according to the relevance prediction result, which can be sorted according to the relevance score of all candidate commodity keywords corresponding to the commodities, and then the top N commodities are selected as the final display result according to the score.
[0060] Specifically, after obtaining the relevance prediction result between the search keyword and the commodity keyword, the relevance prediction result can be used as one of the sorting features, and other sorting features such as commodity sales, user rating, price, and user personalized preference are input into the sorting model for comprehensive scoring. In this embodiment, if the relevance prediction score of a certain commodity is lower than a preset threshold, the commodity is directly filtered out to avoid displaying obviously irrelevant results. For example, for the search keyword "five batteries", the commodity keyword "battery 7 / 4 section" is initially recalled because it contains "battery", but its relevance score is determined to be weakly related because of the reasoning that "five and seven are different attributes", so it is down-weighted or excluded in the final result. Through the above method, the search result presented to the user can be highly relevant to the search keyword and have good sorting quality, thereby improving search efficiency and user satisfaction.
[0061] In a specific application scenario, taking the search keyword "five batteries" and the commodity keyword "battery 7 / 4 section" as an example, the above semantic correlation prediction method is described. First, the search system can obtain the search keyword "five batteries" input by the user, and then match the commodity keywords containing "battery" or "five" based on the inverted index, and preliminarily recall the commodity keyword "battery 7 / 4 section". Then, "five batteries" and "battery 7 / 4 section" can be input into the semantic correlation prediction model, and the correlation between the two keywords can be analyzed by the online prompt model of the semantic correlation prediction model, and the thinking chain prompt information "The modifier "five" in the search keyword is different from the modifier "7" in the commodity keyword, and the two belong to two different specifications of batteries, which does not meet the semantic correlation of compound words, so they are not related" is generated. Subsequently, the student model can receive the above thinking chain prompt information and combine the original word pair for correlation analysis to identify the conflict of key attributes, and finally output a low correlation prediction result (such as a score of 0.2). Finally, the search system can reduce the ranking of the "battery 7 / 4 section" corresponding commodity according to the low score result, or even filter it, to ensure that the user mainly sees "five batteries" or "AA batteries" and other highly relevant commodities, thereby providing accurate search results. The above method can significantly improve the accuracy of the search system in handling complex semantic scenarios such as specification differences, synonyms, and different shapes through commodity keyword matching, deep semantic reasoning, and efficient prediction of lightweight models, while ensuring the real-time and stability of online services, and optimizing the search experience of users.
[0062] The above embodiment matches the search keyword to at least one commodity keyword based on the search keyword, and obtains the thinking chain prompt information by the online prompt model in the semantic correlation prediction model for information reasoning of the search keyword and the commodity keyword, and then obtains the correlation prediction result between the search keyword and the commodity keyword by the student model based on the thinking chain prompt information, and finally obtains the search result by the correlation prediction result. The above method can enhance the explainability and logicality of the correlation prediction by using the real-time generated correlation reasoning process and reasoning result as the prompt information of the correlation prediction, thereby significantly improving the accuracy of the search system in complex semantic scenarios. In addition, the above method can balance the performance and prediction efficiency of the student model after knowledge distillation for correlation prediction, thereby reducing the computing resources required for correlation prediction, ensuring the real-time and stability of online services, and optimizing the search experience of users.
[0063] In an embodiment, in step 201, the search keyword and the commodity keyword can be obtained in the following manner: first, obtaining the search keyword input by the user as a first keyword, and performing information search on the search keyword by using a pre-trained retrieval-augmented generation model to obtain preliminary search results, and then, based on the commodity information and / or the merchant information in the preliminary search results, matching at least one commodity keyword in a commodity keyword database as a second keyword.
[0064] Specifically, after obtaining the search keyword input by the user, first, the retrieval-augmented generation model (RAG, Retrieval-Augmented Generation) is used to perform information search on the search keyword, so as to search for multiple commodity information and merchant information related to the search keyword by using the semantic understanding ability and the text generation ability of the retrieval-augmented generation model, and then, the searched commodity information (such as name, specification, category, etc.) and merchant information (such as store label, etc.) are matched in the commodity keyword database to match at least one commodity keyword related to the search keyword.
[0065] For example, taking the search keyword "five batteries" as an example, after obtaining the search keyword, the search system can first search the keyword by using the retrieval-augmented generation model to search for commodity information containing labels such as "five" or "battery", and then, based on the search results, keyword matching is performed in the commodity keyword database to match multiple search keywords related to the search keyword, such as "five batteries / 4 pieces", "AA type batteries / 10 pieces", "battery 7 / 4 pieces", etc.
[0066] The above embodiment can batch recall commodity keywords related to the search keyword by using the semantic search ability and the text generation ability of the retrieval-augmented generation model, including keywords with specification differences or synonymous expressions, so as to reduce the missed search problem caused by literal mismatch and expand the search range of commodities.
[0067] In one embodiment, in steps 202 and 203, the thought chain prompt information can be generated in the following manner: first, the part-of-speech of the search keyword is analyzed by the online prompt model, and it is determined whether the search keyword is a non-content word, a brand word, and a generic word; then, it is determined in sequence whether the brand name corresponding to the search keyword and the commodity keyword matches when the search keyword is a brand word, and whether the commodity main body, the modifier, and the commodity category corresponding to the search keyword and the commodity keyword match when the search keyword is a generic word; subsequently, based on the matching results of the search keyword and the commodity keyword, a relevance estimation result and a relevance estimation reason are generated; finally, the thought chain prompt information is generated based on the relevance estimation result and the relevance estimation reason, and the thought chain prompt information is input into the student model to use the student model to predict the relevance of the keywords.
[0068] Specifically, after obtaining the search keyword and the commodity keyword, the online prompt model can first analyze the part-of-speech of the search keyword based on a preset prompt template to determine whether the search keyword is a non-content word, a brand word, and a generic word. After determining the part-of-speech of the search keyword, the search keyword and the commodity keyword can be matched based on the judgment logic corresponding to each type of word, and finally, the thought chain prompt information can be generated based on the relevance judgment reason and judgment process of the search keyword and input into the student model.
[0069] For example, taking the search keyword "five batteries" as an example, after obtaining the search keyword, the online prompt model first analyzes the part-of-speech of the keyword and identifies that "five" is a modifier (indicating battery specifications) and "battery" is a commodity main body word, so it is determined to be a generic word rather than a brand word. Subsequently, the thought chain is used to sequentially determine whether the commodity main body (battery), modifier (7), and commodity category (battery category) corresponding to the search keyword and the commodity keyword "battery 7 / 4 piece" match when the search keyword is a generic word. It is found that "five" and "7" differ in specifications (such as AA and AAA models), but the commodity main body and category are consistent. Based on this, the relevance estimation result is "weakly related", the estimation reason is "inconsistent specifications leading to incomplete generic word semantic relevance", and the final integrated thought chain prompt information is "the modifier in the search keyword 【5】 is different from the modifier in the commodity keyword 【7】, which belongs to two different types of batteries with different specifications, and does not completely satisfy the semantic relevance of the generic word, so it is weakly related". Finally, the prompt information is input into the student model to assist in relevance prediction.
[0070] The above embodiments can perform fine matching analysis on search keywords and commodity keywords in multiple dimensions such as brand, subject, modifier, and category, by introducing part-of-speech analysis and structured thought chain reasoning in the online prompt model. Especially for the specification difference problem in the generic word scenario, the online prompt model can generate explicit and interpretable estimation reasons, which can enhance the identification ability of the semantic relevance prediction model for weakly related relationships, thereby improving the accuracy of relevance judgment in complex semantic scenarios.
[0071] In one embodiment, in step 203, the online prompt model can perform thought chain reasoning on the search keywords and commodity keywords in the following manner: when the search keyword is a non-content word, it is determined that the relevance estimation result is strong correlation; when the search keyword is a brand word, if the brand word matches the brand name corresponding to the commodity keyword, it is determined that the relevance estimation result is strong correlation; if the brand word does not match the brand name corresponding to the commodity keyword, it is determined that the relevance estimation result is irrelevant; when the search keyword is a generic word, if the generic word matches the commodity subject, modifier, and commodity category corresponding to the commodity keyword, it is determined that the relevance estimation result is strong correlation; if the generic word matches the commodity subject but not the modifier, or matches the commodity subject but not the commodity category, it is determined that the relevance estimation result is weak correlation; if the generic word does not match the commodity subject, it is determined that the relevance estimation result is irrelevant; finally, based on the judgment condition of the relevance estimation result, the relevance estimation reason is generated.
[0072] Specifically, when the online prompt model performs thought chain reasoning on the search keyword and the commodity keyword, it can first determine the type of the search keyword. When the search keyword is a non-content word, it can be directly determined that the relevance estimation result is strong correlation. For example, when the user searches for "XX Square", "XX Square" itself has no specific direction, so it is only necessary to have a location attribute to be strongly associated. When the search keyword is a brand word, if it matches the brand name in the commodity keyword, that is, the brand names are consistent, it can be determined as strong correlation. If the brand word does not match the brand name in the commodity keyword, that is, the brand names are inconsistent, it is determined as irrelevant. When the search keyword is a generic word, if the commodity main body, the modifier and the category of the commodity keyword are all matched, it is determined as strong correlation. For example, "five batteries" and "battery five / 8 section pack" are strongly correlated. If the commodity main body is matched but the modifier is not matched, such as "five batteries" and "battery 7 / 4 section pack", or the commodity main body is matched but the category is not matched, such as "five batteries" and "battery charger", it is determined as weak correlation. If the commodity main body is not matched, such as "five batteries" and "power bank", it is determined as irrelevant. Finally, the relevance estimation reason can be generated based on the above determination conditions. For example, "the search keyword
five batteries
battery
battery 7 / 4 section pack
five
[0073] The above embodiment classifies and determines in multiple dimensions by introducing non-content words, brand words and generic words, and matches rules based on commodity main body, modifier and category, which can perform structured reasoning on the relevance of search keywords and commodity keywords. In addition, by explicitly defining the judgment conditions for weak correlation, the semantic relevance prediction model can enhance the recognition ability of fuzzy semantic relationship, so as to ensure the accuracy of strong correlation results, and improve the discrimination accuracy and interpretability in weak correlation and irrelevant scenarios.
[0074] In one embodiment, in step 204, the student model can perform relevance prediction in the following way: first, the student model can extract features of the search keyword and the commodity keyword through the parameter-shared double-tower network to obtain the search keyword features and the commodity keyword features, and at the same time, the thought chain prompt information can be extracted through the thought chain tower network to obtain the thought chain features, wherein the thought chain tower network and the double-tower network are not parameter-shared. Further, the search keyword features, the commodity keyword features and the thought chain features can be fused through the expert mixed network, and the relevance prediction result of the search keyword and the commodity keyword can be obtained based on the fused features.
[0075] Specifically, after receiving the search keywords, product keywords, and the thought chain prompt information generated by the online prompt model, the student model can extract features of the search keywords and product keywords through a parameter-sharing dual-tower network. At the same time, the thought chain prompt information generated by the online prompt model can be independently extracted through the thought chain tower network to generate a feature vector that reflects the logic of part-of-speech differences. Subsequently, the student model can dynamically fuse the above three types of feature vectors through the expert mixture network (MoE, Mixture of Experts), wherein the expert mixture network can select different expert models (such as specification matching experts, semantic similarity experts, category matching experts) according to the differences in input features, and perform weighted calculation fusion to obtain the fused features. Finally, the student model can perform correlation prediction based on the fused features and output the correlation prediction results of the search keywords and product keywords.
[0076] In this embodiment, the thought chain prompt information generated by the online prompt model can provide higher-level context information (such as specification differences, semantic associations, etc.) compared to simply entering search keywords and product keywords. Based on this, Figures 1 to 3 , the network structure of the post-interactive ReprBERT of the student model has been improved, that is, an independent COT tower (thinking chain tower network) is added on the basis of the original dual-tower network to model the thinking chain prompt information and form a three-tower structure. The output feature vectors of the three towers can be fused through an expert mixture network, wherein the expert mixture network contains multiple expert models with independent parameters, which can dynamically select a suitable expert combination to perform weighted fusion of search keyword features, product keyword features and thinking chain features, so as to obtain a more accurate representation. In this embodiment, the newly added thinking chain tower network does not share parameters with the dual-tower network, which can ensure that the network can focus on the extraction of high-order logical features, while the expert mixture network can strengthen the collaborative expression of multimodal features through expert-level dynamic interaction, and ultimately improve the student model's ability to handle complex semantic relationships.
[0077] The above embodiment sets a dual-tower structure and an independent thinking chain tower network in the student model, and sets an expert hybrid network, which can accurately extract the multimodal features of search keywords, product keywords and thinking chain prompt information, thereby significantly improving the student model's ability to capture high-order contextual information, and thus enhancing the accuracy of correlation prediction while maintaining low computational overhead.
[0078] For each of the above embodiments, the technical solutions can be applied to the transaction and delivery services of instant e-commerce platforms, such as Taobao Flash Purchase, Taoxianda, Ele.me takeout and retail, etc.
[0079] Further, as Figure 4 andFigure 5 To achieve the specific implementation of the method, the embodiments of the present application provide a semantic correlation prediction device, as shown in the figure Figure 6 The device comprises: A keyword acquisition module 31 is configured to acquire a first keyword and a second keyword to be predicted, and input the first keyword and the second keyword into a pre-trained semantic correlation prediction model; A correlation estimation module 32 is configured to perform correlation estimation on the first keyword and the second keyword by an online prompting model in the semantic correlation prediction model, to obtain a correlation estimation reason and a correlation estimation result; A thought chain reasoning module 33 is configured to generate thought chain prompting information based on the correlation estimation reason and the correlation estimation result, and input the thought chain prompting information into a student model in the semantic correlation prediction model; A correlation prediction module 34 is configured to perform correlation analysis on the first keyword and the second keyword based on the thought chain prompting information by the student model, to obtain a correlation prediction result of the first keyword and the second keyword.
[0080] In a specific application scenario, the first keyword is a search keyword, and the second keyword is a commodity keyword; the correlation estimation module 32 is specifically configured to perform part-of-speech analysis on the search keyword by the online prompting model, and determine whether the search keyword is a non-content word, a brand word, and a generic word; in a thought chain manner, determine whether the brand name corresponding to the search keyword and the commodity keyword matches when the search keyword is a brand word, and whether the main body of the commodity corresponding to the search keyword and the commodity keyword, the modification word, and the commodity category match when the search keyword is a generic word; based on the matching results of the search keyword and the commodity keyword, obtain the correlation estimation reason and the correlation estimation result of the search keyword and the commodity keyword.
[0081] In a specific application scenario, the relevance estimation module 32 is further configured to: when the search keyword is a non-content keyword, determine that the relevance estimation result is strong correlation; when the search keyword is a brand keyword, if the brand keyword matches the brand name corresponding to the product keyword, determine that the relevance estimation result is strong correlation; if the brand keyword does not match the brand name corresponding to the product keyword, determine that the relevance estimation result is irrelevant; when the search keyword is a generic keyword, if the generic keyword matches the product subject, the modifier, and the product category corresponding to the product keyword, determine that the relevance estimation result is strong correlation; if the generic keyword matches the product subject but does not match the modifier, or matches the product subject but does not match the product category, determine that the relevance estimation result is weak correlation; if the generic keyword does not match the product subject, determine that the relevance estimation result is irrelevant; and generate a relevance estimation reason based on the judgment condition of the relevance estimation result.
[0082] In a specific application scenario, the relevance prediction module 34 can be configured to: the student model extracts features of the search keyword and the product keyword through a parameter-shared double-tower network to obtain search keyword features and product keyword features; the student model extracts features of the thought chain prompt information through a thought chain tower network to obtain thought chain features, where the parameters of the thought chain tower network are not shared with the double-tower network; and the student model performs feature fusion on the search keyword features, the product keyword features, and the thought chain features through an expert hybrid network, and performs relevance prediction based on the fused features to obtain a relevance prediction result of the search keyword and the product keyword.
[0083] In a specific application scenario, the keyword acquisition module 31 is further configured to: acquire a search keyword input by a user as the first keyword, and perform information search on the search keyword through a pre-trained retrieval enhancement generation model to obtain a preliminary search result; based on product information and / or merchant information in the preliminary search result, match at least one product keyword in a product keyword database to obtain the second keyword; and the device further includes a search result determination module, where the search result determination module can be configured to: based on the relevance prediction result of the search keyword and the product keyword, obtain a search result corresponding to the search keyword, and send the search result to a client.
[0084] It should be noted that other corresponding descriptions of the various functional units involved in the semantic relevance prediction device provided by the embodiments of the present application can be referred to the corresponding descriptions in the method, which will not be repeated here. Figures 4 to 5
[0085] The embodiments of the present application also provide a computer device, which can be a personal computer, a server, a network device, etc. The computer device comprises a bus, a processor, a memory and a communication interface, and can further comprise an input / output interface and a display device. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store location information. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the steps in the method embodiments.
[0086] Those skilled in the art can understand that the structure of the computer device described above is only part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can comprise more or fewer components, or combine certain components, or have a different component arrangement.
[0087] In one embodiment, a computer readable storage medium is provided, which can be non-volatile or volatile, and stores a computer program. The computer program is executed by the processor to implement the steps in the method embodiments described above.
[0088] In one embodiment, a computer program product is provided, which comprises a computer program. The computer program is executed by the processor to implement the steps in the method embodiments described above.
[0089] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.
[0090] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0091] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0092] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A semantic relevance prediction model, characterized in that: The semantic relevance prediction model includes a student model obtained by distilling a teacher model and a pre-trained online prompt model, and the student model includes: A parameter-sharing dual-tower network is used to extract features of the first keyword and the second keyword to be predicted, thereby obtaining first keyword features and second keyword features; A thinking chain tower network is used to extract features from the thinking chain prompt information generated by the online prompt model to obtain thinking chain features. The thinking chain tower network and the dual-tower network do not share parameters. A mixture of experts network is used to perform feature fusion on the first keyword feature, the second keyword feature, and the thought chain feature to obtain a keyword fusion feature; A deep neural network is used to perform semantic relevance prediction based on the keyword fusion feature to obtain a semantic relevance prediction result between the first keyword and the second keyword.
2. The semantic relevance prediction model according to claim 1, characterized in that: The deep neural network of the student model is specifically used to: Obtaining the keyword fusion feature and the explicit cross feature, where the explicit cross feature is constructed based on the cross-text features, category features, and entity features in the first keyword and the second keyword; Semantic relevance prediction is performed based on the explicit cross feature and the keyword fusion feature to obtain a semantic relevance prediction result between the first keyword and the second keyword.
3. The semantic relevance prediction model according to claim 1, characterized in that: The online prompt model is specifically used for: When the student model performs semantic relevance prediction, the first keyword and the second keyword are obtained, and a relevance estimation is performed on the first keyword and the second keyword to obtain a relevance estimation reason and a relevance estimation result; Based on the correlation estimation reason and the correlation estimation result, thought chain prompt information is generated, and the thought chain prompt information is input into the student model.
4. The semantic relevance prediction model according to any one of claims 1 to 3, characterized in that: The semantic relevance prediction model also includes a pre-trained teacher model and an offline prompt model, wherein, The offline prompt model is used to generate sample thought chain prompt information based on sample keyword pairs in the sample data set and correlation labels of the sample keyword pairs; The teacher model is used to receive the sample keyword pairs and the sample thought chain prompt information as input, generate semantic relevance prediction results, and train the student model based on the semantic relevance prediction results output by the teacher model and the relevance labels of the sample keyword pairs during the knowledge distillation process to obtain a trained student model.
5. The semantic relevance prediction model according to claim 4, characterized in that: The training method of the student model comprises: Pre-training the teacher model based on sample keyword pairs in the sample data set and sample thought chain prompt information generated by the offline prompt model, and freezing the parameters of the teacher model; Predicting semantic relevance of the sample keyword pairs using a teacher model with frozen parameters, and obtaining a soft target by temperature scaling the prediction result output by the teacher model; Iteratively training the student model based on the sample keyword pairs and sample thought chain prompt information generated by a pre-trained online prompt model, wherein the loss function of the student model includes a distillation loss between the temperature-scaled output of the student model and the soft target generated by the teacher model, and a cross-entropy loss between the output of the student model and a relevance label; When the training reaches a preset number of rounds or the performance of the student model on the validation set reaches a preset convergence target, the model training is stopped to obtain a trained student model.
6. The semantic relevance prediction model according to claim 4, characterized in that: The training method of the online prompt model includes: Loading a pre-trained offline prompting model, mapping the weight parameters of the offline prompting model to an integer interval based on a preset quantization parameter, and rounding or truncating the mapped weight parameters; During each forward propagation of the offline prompt model, converting the activation value of the neuron from a floating point number to an integer; The offline prompt model with adjusted weight parameters and activation values is trained with quantization perception to obtain a trained online prompt model.
7. A semantic relevance prediction method, characterized in that: The method comprises: Obtaining a first keyword and a second keyword to be predicted, and inputting the first keyword and the second keyword into the semantic relevance prediction model according to any one of claims 1 to 6; Predicting the relevance of the first keyword and the second keyword using an online prompt model in the semantic relevance prediction model to obtain a relevance prediction reason and a relevance prediction result; generating thought chain prompt information based on the relevance estimation reason and the relevance estimation result, and inputting the thought chain prompt information into the student model in the semantic relevance prediction model; The student model performs a correlation analysis on the first keyword and the second keyword based on the thought chain prompt information to obtain a correlation prediction result of the first keyword and the second keyword.
8. The semantic relevance prediction method according to claim 7, characterized in that: The first keyword is a search keyword, and the second keyword is a product keyword; then, the online prompt model in the semantic relevance prediction model is used to perform relevance estimation on the first keyword and the second keyword, and obtain a relevance estimation reason and a relevance estimation result, including: Performing part-of-speech analysis on the search keyword using the online prompt model, and determining whether the search keyword is a non-content word, a brand word, or a general word; Use a chain of thought approach to determine whether the brand name corresponding to the product keyword matches the search keyword when the search keyword is a brand word, and whether the product subject, modifier, and product category corresponding to the product keyword match the search keyword when the search keyword is a general word; Based on the matching result between the search keyword and the product keyword, a correlation estimation reason and a correlation estimation result between the search keyword and the product keyword are obtained.
9. The semantic relevance prediction method according to claim 8, characterized in that: The obtaining of the correlation estimation reason and correlation estimation result between the search keyword and the product keyword based on the matching result between the search keyword and the product keyword includes: When the search keyword is a non-content word, determining that the relevance estimation result is a strong correlation; When the search keyword is a brand word, if the brand word matches the brand name corresponding to the product keyword, the correlation estimation result is determined to be strongly correlated; if the brand word does not match the brand name corresponding to the product keyword, the correlation estimation result is determined to be unrelated; When the search keyword is a general word, if the general word matches the product subject, modifier, and product category corresponding to the product keyword, the correlation estimation result is determined to be strongly correlated; if the general word matches the product subject corresponding to the product keyword but the modifier does not match, or if the product subject matches but the product category does not match, the correlation estimation result is determined to be weakly correlated; if the general word does not match the product subject corresponding to the product keyword, the correlation estimation result is determined to be irrelevant; Based on the judgment condition of the correlation estimation result, a correlation estimation reason is generated.
10. The semantic relevance prediction method according to claim 8 or 9, characterized in that: The student model performs a correlation analysis on the first keyword and the second keyword based on the thought chain prompt information to obtain a correlation prediction result between the first keyword and the second keyword, including: The student model extracts features of the search keyword and the product keyword through a parameter-sharing dual-tower network to obtain search keyword features and product keyword features; The student model extracts features from the thought chain prompt information through a thought chain tower network to obtain thought chain features, wherein the thought chain tower network and the double tower network do not share parameters; The student model performs feature fusion on the search keyword features, the product keyword features and the thought chain features through an expert hybrid network, and performs correlation prediction based on the fused features to obtain a correlation prediction result between the search keyword and the product keyword.
11. The semantic relevance prediction method according to claim 8 or 9, characterized in that: The obtaining of the first keyword and the second keyword to be predicted includes: Obtaining a search keyword input by a user as the first keyword, and performing an information search on the search keyword using a pre-trained retrieval enhancement generation model to obtain preliminary search results; Based on the product information and / or merchant information in the preliminary search results, at least one product keyword is matched in a product keyword database as the second keyword; After obtaining the correlation prediction result of the first keyword and the second keyword, the method further includes: obtaining search results corresponding to the search keyword based on the correlation prediction result of the search keyword and the product keyword, and sending the search results to the client.
12. A semantic relevance prediction device, characterized in that: The device comprises: a keyword acquisition module, configured to acquire a first keyword and a second keyword to be predicted, and input the first keyword and the second keyword into the semantic relevance prediction model according to any one of claims 1 to 6; a relevance estimation module, configured to estimate the relevance of the first keyword and the second keyword using an online prompt model in the semantic relevance prediction model, and obtain a relevance estimation reason and a relevance estimation result; a thought chain reasoning module, configured to generate thought chain prompt information based on the relevance estimation reason and the relevance estimation result, and input the thought chain prompt information into a student model in the semantic relevance prediction model; A relevance prediction module is used for the student model to perform relevance analysis on the first keyword and the second keyword based on the thought chain prompt information to obtain a relevance prediction result of the first keyword and the second keyword.
13. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 7 to 11 is implemented.
14. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 7 to 11 is implemented.
Citation Information
Patent Citations
Correlation judgment method and LLM-based correlation judgment model construction method
CN117743950A
Knowledge distillation method, device, equipment, storage medium and computer program product
CN119005176A
Financial field training data construction method based on dual feedback mechanism
CN120123471A
Geological domain named entity recognition and classification method based on thinking chain and hybrid experts
CN120387454A
Method and apparatus for artificial neural network-based search term dictionary generation and search
WO2024185948A1