Content Completion Method, Apparatus, Device, and Medium Based on an Auto-Completion Model
By dynamically updating a language model with user search logs and retrieval library data, the method improves the accuracy and user experience of query suggestions in natural language processing systems.
Patent Information
- Application Number
- CN202311329042.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-10-13
AI Technical Summary
The existing automatic completion method has low retrieval efficiency when the lexicon content increases, the recommended algorithm has poor effect when cold starts, and the data source is high, resulting in poor user experience.
By obtaining user search text to generate search logs, building content completion data sets, and using preset language models for online updates, combining search library text for vectorization and model loss calculation, the model is realized continuously optimized.
It improves the accuracy and user experience of query prompts, solves the problems of low retrieval efficiency and limitations of data source, and ensures that the model can be continuously optimized after cold start.
Smart Images

Figure CN117235210B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and in particular to a content completion method, device, equipment and medium based on an automatic completion model. Background Art
[0002] With the development of deep learning and neural networks, natural language technology has ushered in an opportunity for development. In the field of search engine applications, natural language processing technology can be used to help users focus on the content they are interested in and realize semantic analysis of search content. Automatic completion of search content means that when using a search engine, the user only enters individual keywords, and the algorithm intelligently prompts the user's complete query sentence, helping the user to quickly locate the search content from the massive search library.
[0003] There are three commonly used methods for automatic completion in the prior art. The first is the character matching method. This method requires manual maintenance of the dictionary library, regular expansion and reduction of the dictionary library content, and matching the user query words with the characters in the dictionary library to lock the recommendation list. However, the retrieval efficiency of this method is too slow. When the content of the dictionary increases, the prompt return time often exceeds the user's expectations, which reduces the user experience; the second method is the recommendation system method. This method collects data features of user search behavior, uses the current more mature recommendation algorithm, analyzes and models the user's search behavior, and can achieve personalized prompts. However, this method has a cold start process when the recommendation algorithm is launched. User behavior data may be missing during this process, resulting in poor recommendation results; the third method is an intelligent prompt method based on log mining, which uses search log information to mine user search information and make corresponding recommendations, but the data of this method only comes from search logs, and the data source has high limitations. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a content completion method, device, equipment and medium based on an automatic completion model. The model can be updated online through user search data and retrieval library text data to improve the accuracy of query prompts and user experience. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a content completion method based on an automatic completion model, comprising:
[0006] Acquire a search text input by a user terminal, and process the search text based on a preset log generation rule to generate a search log;
[0007] Count the number of logs in the search log, and determine whether the number of logs meets a preset condition based on a preset quantity threshold. If so, add start and end symbols to the log text in the search log and the retrieval library text in the preset retrieval library, and intercept the added text to construct a content completion dataset based on the obtained several text segments;
[0008] Perform vectorization processing on the several text segments in the content completion dataset, and input the obtained several text tensors into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors;
[0009] Update the preset language model based on the overall model loss, use the obtained updated model to provide an auto-completed text for the client, and then re-jump to the step of obtaining the search text input by the client, and process the search text based on a preset log generation rule to generate a search log for the next round of model update.
[0010] Optionally, the step of obtaining the search text input by the client and processing the search text based on a preset log generation rule to generate a search log includes:
[0011] Obtain the search text input by the client, and determine the query time and user identifier corresponding to the search text;
[0012] Determine whether the search text is the same as the actual query text. If not, determine the target ranking of the actual query text among several recommended texts, and determine the actual query text as the log text to generate a first search log based on the query time, the user identifier, the search text, the log text, and the target ranking; the several recommended texts are recommended texts related to the search text generated by the preset language model; the actual query text is the search text determined by the client from the several recommended texts;
[0013] If so, determine the search text as the log text to generate a second search log based on the query time, the user identifier, and the log text.
[0014] Optionally, the step of counting the number of logs in the search log and determining whether the number of logs meets a preset condition based on a preset quantity threshold. If so, add start and end symbols to the log text in the search log and the retrieval library text in the preset retrieval library, and intercept the added text to construct a content completion dataset includes:
[0015] Count the number of logs of all the search logs locally, and determine whether the number of logs is not less than several integer multiples of a preset number threshold;
[0016] If so, determine the log text in all the search logs and the text in the preset retrieval library, and add a start symbol and an end symbol to the log text and the text in the retrieval library to obtain the added text;
[0017] Perform truncation processing on the added text respectively based on each text truncation length in the preset truncation range, so as to construct a content completion data set based on the obtained several text segments.
[0018] Optionally, after counting the number of logs of all the search logs locally and determining whether the number of logs is not less than several integer multiples of a preset number threshold, it further includes:
[0019] If not, jump to the step of obtaining the search text input by the user end and processing the search text based on the preset log generation rule to generate search logs, until the number of logs of all the search logs locally is not less than several integer multiples of the preset number threshold, so as to construct the content completion data set based on the search logs.
[0020] Optionally, the vectorization processing of the several text segments in the content completion data set and inputting the obtained several text tensors into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors includes:
[0021] Perform vectorization processing on the several text segments in the content completion data set, and obtain a first text tensor corresponding to the several text segments, a second text tensor corresponding to the corresponding labels of the several text segments, and a third text tensor corresponding to the prediction direction of the several text segments based on a preset dictionary index correspondence table;
[0022] Perform dimensionality reduction and feature extraction on the first text tensor to obtain a processed first text tensor;
[0023] Perform a max pooling operation on the processed first text tensor, so as to calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor.
[0024] Optionally, the performing a max pooling operation on the processed first text tensor to calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor includes:
[0025] Input the processed first text tensor into the Maxpooling layer to perform max pooling operation on the processed first text tensor, so as to obtain the processed first text tensor;
[0026] Input the processed first text tensor into the first linear classification layer and the second linear classification layer respectively, so as to obtain the corresponding first output tensor and second output tensor;
[0027] Calculate the cross-entropy loss of the preset language model based on the first output tensor and the second text tensor to obtain the target cross-entropy loss; calculate the mean square error of the preset language model based on the second output tensor and the third text tensor, and use the obtained mean square error as the prediction loss of the preset language model;
[0028] Use the sum of the target cross-entropy loss and the prediction loss as the overall model loss of the preset language model.
[0029] Optionally, the step of updating the preset language model based on the overall model loss, using the obtained updated model to provide auto-completed text for the client, and then re-jumping to the step of obtaining the search text input by the client and processing the search text based on the preset log generation rule to generate a search log for the next round of model update includes:
[0030] Update the preset language model based on the overall model loss to obtain an updated language model;
[0031] Judge whether the current search text input by the client is received. If so, generate several auto-completed texts corresponding to the current search text based on the updated language model, and then re-jump to the step of obtaining the search text input by the client and processing the search text based on the preset log generation rule to generate a search log for the next round of model update.
[0032] In a second aspect, the present application discloses a content completion device based on an auto-completion model, including:
[0033] A log generation module, configured to obtain the search text input by the client and process the search text based on a preset log generation rule to generate a search log;
[0034] A dataset construction module, configured to count the number of logs in the search log, and judge whether the number of logs meets a preset condition based on a preset number threshold. If so, add start and end symbols to the log text in the search log and the retrieval library text in the preset retrieval library, and intercept the obtained added text, so as to construct a content completion dataset based on the obtained several text segments;
[0035] A loss calculation module, configured to perform vectorization processing on the several text segments in the content completion dataset, and input the obtained several text tensors into a preset language model, so as to calculate the overall model loss of the preset language model based on the several text tensors;
[0036] A model update module, configured to update the preset language model based on the overall model loss, use the obtained updated model to provide automatically completed text for the client, and then re-jump to the step of obtaining the search text input by the client, and process the search text based on a preset log generation rule to generate a search log, so as to perform the next round of model update.
[0037] In a third aspect, the present application discloses an electronic device, including:
[0038] A memory, configured to store a computer program;
[0039] A processor, configured to execute the computer program to implement the content completion method based on the auto-completion model as described above.
[0040] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program, and when the computer program is executed by a processor, the content completion method based on the auto-completion model as described above is implemented.
[0041] In this application, first, the search text input by the client is obtained, and the search text is processed based on a preset log generation rule to generate a search log. The log quantity of the search log is counted, and it is determined whether the log quantity meets the preset condition based on a preset quantity threshold. If so, start and end symbols are added to the log text in the search log and the retrieval text in the preset retrieval library, and the obtained text after addition is intercepted to construct a content completion data set based on the obtained several text segments. Then, the several text segments in the content completion data set are vectorized, and the obtained several text tensors are input into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors. Finally, the preset language model is updated based on the overall model loss, and the updated model is used to provide automatic completion text for the client. Then, it jumps back to the step of obtaining the search text input by the client and processing the search text based on the preset log generation rule to generate a search log to perform the next round of model update. It can be seen that through the method of this application, a search log can be generated based on the user's search text. After the search log reaches the preset quantity threshold, the retrieval text in the preset retrieval library and the search text in the search log are processed, and the processed data is used to construct a content completion data set. Then, the text segments in the content completion data set are vectorized to train the preset language model through the obtained several text tensors, and the model loss of the preset language model is calculated. The language model is updated according to the obtained model loss, and when using the updated model to provide completion text, the search text is collected simultaneously to perform the next round of update. In this way, the model can be updated online through the user search data and the retrieval library text data, improving the accuracy of query prompts and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0043] Figure 1 Flowchart of a content completion method based on an auto-completion model provided by this application;
[0044] Figure 2 Specific flowchart of a content completion method based on an auto-completion model provided by this application;
[0045] Figure 3 Schematic diagram of an auto-completion model framework provided by this application;
[0046] Figure 4 Structural schematic diagram of a content completion device based on an auto - completion model provided for this application;
[0047] Figure 5 Structural diagram of an electronic device provided for this application. Specific embodiments
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0049] There are three commonly used methods for auto - completion in the prior art. The first is the character matching method. This method has too slow retrieval efficiency. When the content of the thesaurus increases day by day, the prompt return time often exceeds the user's expectation, reducing the user experience. The second method is the recommendation system method. There is a cold - start process when the recommendation algorithm is launched. During this process, user behavior data may be missing, resulting in poor recommendation effects. The third method is that the data of this method only comes from search logs, and the data source has high limitations.
[0050] In order to overcome the above - mentioned technical problems, the present invention discloses a content completion method, device, equipment and medium based on an auto - completion model. The model can be updated online through user search data and retrieved library text data, improving the accuracy of query prompts and the user experience.
[0051] See Figure 1 As shown, the embodiments of the present invention disclose a content completion method based on an auto - completion model, including:
[0052] Step S11: Obtain the search text input by the user terminal, and process the search text based on a preset log generation rule to generate a search log.
[0053] In this embodiment, in order to update the preset language model, that is, to update the auto-complete model, it is necessary to process the text in the preset retrieval library and the search text input by the client collected in real time to obtain training data for training the model. First, it is necessary to obtain the search text input by the client to generate a search log. The specific process is as follows: Obtain the search text input by the client, and determine the query time and user identifier corresponding to the search text; Determine whether the search text is the same as the actual query text. If not, then determine the target ranking of the actual query text among several recommended texts, and determine the actual query text as the log text, so as to generate a first search log based on the query time, the user identifier, the search text, the log text, and the target ranking; The several recommended texts are recommended texts related to the search text generated by the preset language model; The actual query text is the search text determined by the client from the several recommended texts; If so, determine the search text as the log text, so as to generate a second search log based on the query time, the user identifier, and the log text. That is, it is necessary to obtain the search text input by the client, and it is necessary to determine the query time and user identifier when the search text is input, that is, the user ID (Identity document), and it is necessary to determine whether the actual query text finally searched by the user is the same as the search text. If not, it means that the query term used by the user for searching is the text content recommended based on the preset language model, that is, based on the auto-complete model. At this time, it is necessary to determine the actual query text and the ranking of the actual query text in the returned results, and then use the actual query text as the log text for generating the search log, and generate a search log according to the query time, user ID, search text, log text, and target ranking; In another case, if the search text is consistent with the actual query text, the search text can be determined as the log text for generating the search log, and then a search log is generated based on the log text, query time, and user ID.
[0054] Step S12: Count the number of logs in the search log, and judge whether the number of logs meets the preset conditions based on a preset number threshold. If so, add start and end symbols to the log text in the search log and the retrieval library text in the preset retrieval library, and intercept the obtained added text, so as to construct a content completion data set based on the obtained several text fragments.
[0055] In this embodiment, in order to achieve automatic update of the model, conditions for model update need to be set, and the specific process is as follows: count the number of logs of all the local search logs, and determine whether the number of logs is not less than several integer multiples of a preset number threshold; if so, determine the log texts in all the search logs and the retrieval library texts in the preset retrieval library, and add a start symbol and an end symbol to the log texts and the retrieval library texts to obtain the added texts; respectively perform truncation processing on the added texts based on each text truncation length in the preset truncation range, so as to construct a content completion data set based on the obtained several text segments. That is to say, the automatic update rule of the model can be set as follows: count all the local search logs. If the data volume of the search logs is greater than an integer multiple of the preset threshold, the text can be processed. It should be noted that the preset threshold is not limited in this application and can be set according to user needs. For example, the preset threshold can be set to 5000. Then, when the number of local search logs is not less than 5000, the model can be updated. When the number of local search logs is not less than 10000, the model can be updated in the next round. Further, when it is determined to update the model, the retrieval library texts in the preset retrieval library can be extracted, and start and end symbols are added to the retrieval library texts and the log texts. <start>”and" <end>”, and intercept a text segment of length L, and construct a content completion dataset based on the obtained text segment, and the value of L can be {2, 3, 4, 5}. And the scope of the retrieval library text includes parts such as the title, abstract, and detailed description. Taking the title as an example, "Withdraw housing provident fund for self-occupied housing purchase" is the title of the retrieval library text. After adding the start and end symbols, it becomes " <start>Withdrawal of housing provident fund for purchasing self-occupied housing <end>", where L takes the value of 3, and the formed text segment is <start>Purchase, purchase from, buy for self-occupation, self-occupation residence, residence housing, housing withdrawal, withdrawal housing, housing provident, provident fund, provident fund, accumulation fund <end>After obtaining all text segments, a content completion dataset can be created based on the text segments. Moreover, the sample format in the content completion dataset is {text segment, label, prediction direction}, where label is the label corresponding to the text segment, and the prediction direction indicates whether the position where the text segment appears in the sentence is in the first half or the second half of the sentence. In this way, the completion dataset for model training can be generated by combining the data in the retrieval library and the data of the user's real-time search, making the training of the auto-completion model more comprehensive and avoiding the situation where the model recommendation is inaccurate due to training the model with a single data source.
[0056] It should be noted that after counting the number of logs of all the search logs locally and determining whether the number of logs is not less than several integer multiples of a preset number threshold, it further includes: if not, then jump to the step of obtaining the search text input by the user terminal and processing the search text based on a preset log generation rule to generate search logs until the number of logs of all the search logs locally is not less than several integer multiples of the preset number threshold, so as to construct the content completion dataset based on the search logs. That is to say, if the current number of search logs does not reach an integer multiple of the preset number threshold, it is necessary to continue collecting the input search text and generate new search logs according to the search text until the number of search logs is not less than an integer multiple of the preset number threshold.
[0057] Step S13: Perform vectorization processing on the several text segments in the content completion dataset, and input the obtained several text tensors into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors.
[0058] In this embodiment, the text segments in the completion dataset can be vectorized to improve the data for model training, reduce the time for model training, and after obtaining text tensors by vectorizing the text, the text tensors can be input into the model to calculate the overall model loss of the model. The specific process is as follows: Perform vectorization processing on the several text segments in the content completion dataset, and obtain a first text tensor corresponding to the several text segments, a second text tensor corresponding to the labels of the several text segments, and a third text tensor corresponding to the prediction directions of the several text segments based on a preset dictionary index correspondence table; perform dimensionality reduction and feature extraction on the first text tensor to obtain a processed first text tensor; perform a max pooling operation on the processed first text tensor to calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor. That is to say, perform vectorization processing on the samples in the dataset, where the maximum character length in the text is set to L, and a tensor T corresponding to the text segment can be generated using the dictionary index correspondence table, with a dimension of R L The number of characters is v, the tensor corresponding to label is Y, and the dimension is R 1 , and both v and L are positive integers. The tensor M represents the prediction direction of the text segment in the sample. M being 1 indicates that the prediction direction is forward, and M being 0 indicates that the prediction direction is backward. Here, R represents the real number space, and both v and L are positive integers. Then, the tensor T can be input into the embedding layer of the model to obtain the dimension-reduced tensor X e , and perform feature extraction on X e to obtain the tensor X T . Perform max pooling operation on, and obtain X o , and calculate the overall model loss of the model using X o , M, and Y
[0059] Step S14: Update the preset language model based on the overall model loss, use the updated model obtained to provide the client with automatically completed text, and then jump back to obtaining the search text input by the client, and process the search text based on the preset log generation rule to generate a search log, so as to perform the next round of model update
[0060] In this embodiment, updating the preset language model based on the overall model loss, using the updated model obtained to provide the client with automatically completed text, and then jumping back to obtaining the search text input by the client, and processing the search text based on the preset log generation rule to generate a search log, so as to perform the next round of model update, includes: updating the preset language model based on the overall model loss to obtain an updated language model; determining whether the current search text input by the client is received. If so, generate several automatically completed texts corresponding to the current search text based on the updated language model, and then jump back to obtaining the search text input by the client, and process the search text based on the preset log generation rule to generate a search log, so as to perform the next round of model update. That is to say, the model can be updated based on the obtained overall model loss, and after the user inputs the search text, the user's search text can be input into the updated model, so as to use the updated model to recommend search terms for the user, and regenerate the search log based on the user's search text to perform the next round of model update. In this way, after the initial model is launched, the persistent update of the model can be realized to ensure the accuracy of content completion
[0061] It can be seen that in this embodiment, first, the search text input by the client is obtained, and the search text is processed based on a preset log generation rule to generate a search log. Then, the number of logs in the search log is counted, and it is judged whether the number of logs meets a preset condition based on a preset quantity threshold. If so, start and end symbols are added to the log text in the search log and the retrieval library text in the preset retrieval library, and the obtained text after addition is intercepted to construct a content completion data set based on the obtained several text fragments. Then, the several text fragments in the content completion data set are vectorized, and the obtained several text tensors are input into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors. Finally, the preset language model is updated based on the overall model loss, and the updated model is used to provide auto-completed text for the client, and it jumps back to the step of obtaining the search text input by the client and processing the search text based on the preset log generation rule to generate a search log for the next round of model update. It can be seen that through the method of this application, a search log can be generated based on the user's search text. After the search log reaches the preset quantity threshold, the retrieval text in the preset retrieval library and the search text in the search log are processed, and the processed data is used to construct a content completion data set. Then, the text fragments in the content completion data set are vectorized to train the preset language model through the obtained several text tensors, and the model loss of the preset language model is calculated, so as to update the language model according to the obtained model loss, and when using the updated model to provide completion text, the search text is collected at the same time for the next round of update. In this way, on the one hand, the model can be updated online through the user search data and the retrieval library text data; on the other hand, the text in the retrieval library and the text in the real-time generated search log can be used to train the model to improve the accuracy of query prompts.
[0062] Based on the foregoing embodiments, it can be known that through the method of this application, it is necessary to implement the update of the model. Therefore, this embodiment describes in detail how to update the model. Refer to Figure 2 As shown, an embodiment of the present invention discloses a content completion method based on an auto-completion model, including:
[0063] Step S21: Obtain the search text input by the client, and process the search text based on a preset log generation rule to generate a search log.
[0064] Step S22: Count the number of logs in the search log, and determine whether the number of logs meets the preset conditions based on a preset quantity threshold. If so, add start and end symbols to the log texts in the search log and the retrieval library texts in the preset retrieval library, and intercept the added texts to construct a content completion dataset based on the obtained several text segments.
[0065] Step S23: Vectorize the several text segments in the content completion dataset, and obtain a first text tensor corresponding to the several text segments, a second text tensor corresponding to the labels of the several text segments, and a third text tensor corresponding to the prediction directions of the several text segments based on a preset dictionary index correspondence table.
[0066] In this embodiment, as Figure 3 shown, it is necessary to vectorize several text segments in the content completion dataset, and then it is necessary to use the dictionary index correspondence table to determine the tensor T corresponding to the text segment, the tensor Y corresponding to the text label label, and the tensor M corresponding to the prediction direction of the text segment, where the dimension of the tensor T is R L , the dimension of the tensor Y is R 1 , and the value of the tensor M is 1 or 0.
[0067] Step S24: Perform dimensionality reduction and feature extraction on the first text tensor to obtain a processed first text tensor.
[0068] In this embodiment, it is necessary to input the tensor T into the Embedding layer for dimensionality reduction to obtain the tensor X e , and X e = Embedding(T). Further, it is necessary to input the obtained tensor X e into the Transformer layer for feature extraction to obtain the tensor X T , and X T = Transformer(X e ), and the dimension of X T is R L×E .
[0069] Step S25: Perform a max pooling operation on the processed first text tensor to calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor.
[0070] In this embodiment, it is necessary to input the processed first text tensor, that is, X T into the Maxpooling layer to perform a max pooling operation on X T to obtain more obvious semantic information in the sentence length dimension and obtain the tensor X O , and X O = Maxpooling(X T ), where the dimension of X O is R E . Then X O needs to be input into two linear classification layers MLP respectively. The output of the first linear classification layer is X M , with the dimension of R V , and X M = MLP(X O ). In the second linear classification layer, the sigmoid function is selected as the activation function to determine the forward probability and backward probability of the text, and the output is X S , and X S = sigmoid(MLP2(X O )). The probability of each text segment appearing can be obtained through the sigmoid function. When X S > 0.5, the prediction direction is forward; when X S < 0.5, the prediction direction is backward. For example, there is a text { <start>, w1, w2, …, w n , <end>}, then the overall probability of the text is P{ <start>, w1, w2, …, w n , <end>}=P( <start>)*P(w1│ <start> )*…*P( <end>|w n , …, w1, <start>), taking L as 3 for example, to simplify the calculation, it is assumed that the above probabilities follow the Markov assumption, and the forward probability calculation formula is as follows: P{ <start>, w1, w2, …, w n , <end>}≈P( <start>)*P(w1│ <start>)*P(w2│w1 <start>)*P(w3│w2,w1, <start> )*…*P( <end>|w n , w n-1 , w n-2 ); The formula for calculating the backward probability is as follows: P{ <start>, w1, w2, …, w n , <end>}=P( <end>)*P(w n │ <end>)*P(w n-1 │w n , <end>)*P(w n-2 │w n-1 ,w n , <end> )*…*P( <start>|(w1, w2, w3).
[0071] Further, it should be noted that, based on the obtained X M and X S To calculate the overall loss of the model, first, based on the output X of the linear layer M and the tensor Y corresponding to the label label, calculate the cross-entropy loss, and the cross-entropy loss loss1 = CrossEntropyLoss(X M , Y). It is necessary to calculate the mean squared error loss as the prediction loss based on X S and the tensor M corresponding to the prediction direction of the text segment to correct the prediction direction of the text segment, and the prediction loss loss2 = MSE(X S , M). The overall model loss loss = loss1 + loss2.
[0072] Step S26: Update the preset language model based on the overall model loss, use the obtained updated model to provide automatically completed text for the client, and then redirect to obtaining the search text input by the client, and process the search text based on the preset log generation rule to generate a search log, so as to perform the next round of model update.
[0073] Therefore, in this embodiment, in order to calculate the model loss, first, vectorize the several text segments in the content completion dataset, and obtain the first text tensor corresponding to the several text segments, the second text tensor corresponding to the labels of the several text segments, and the third text tensor corresponding to the prediction directions of the several text segments based on the preset dictionary index correspondence table. Then, perform dimensionality reduction and feature extraction on the first text tensor to obtain the processed first text tensor. Finally, perform a max-pooling operation on the processed first text tensor to calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor. In this way, the model can be updated through the model loss to make the automatic completion model more accurate, and by using Transformer to process the semantic connection between contexts, the cold start problem in the recommendation algorithm is solved.
[0074] See Figure 4 As shown, an embodiment of the present invention discloses a content completion device based on an automatic completion model, including:
[0075] A log generation module 11, configured to obtain the search text input by the client, and process the search text based on a preset log generation rule to generate a search log;
[0076] The dataset construction module 12 is used to count the number of logs in the search log, and determine whether the number of logs meets a preset condition based on a preset quantity threshold. If so, start and end symbols are added to the log text in the search log and the retrieval library text in the preset retrieval library, and the obtained added text is intercepted to construct a content completion dataset based on a number of obtained text segments;
[0077] The loss calculation module 13 is used to perform vectorization processing on the number of text segments in the content completion dataset, and input the obtained number of text tensors into a preset language model to calculate the overall model loss of the preset language model based on the number of text tensors;
[0078] The model update module 14 is used to update the preset language model based on the overall model loss, use the obtained updated model to provide an automatically completed text for the client, and then redirect to the step of obtaining the search text input by the client, and process the search text based on a preset log generation rule to generate a search log, so as to perform the next round of model update.
[0079] It can be seen that in this application, the search text input by the user end is first obtained, and the search text is processed based on a preset log generation rule to generate a search log, and the log quantity of the search log is counted, and it is judged whether the log quantity meets the preset condition based on a preset quantity threshold. If so, start and end symbols are added to the log text in the search log and the retrieval library text in the preset retrieval library, and the obtained text after addition is intercepted, so as to construct a content completion data set based on the obtained several text segments. Then, the several text segments in the content completion data set are vectorized, and the obtained several text tensors are input into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors. Finally, the preset language model is updated based on the overall model loss, and the updated model is used to provide an auto-completed text for the user end, and it jumps back to the step of obtaining the search text input by the user end and processing the search text based on the preset log generation rule to generate a search log to perform the next round of model update. It can be seen that through the method of this application, a search log can be generated based on the user's search text, and after the search log reaches the preset quantity threshold, the retrieval text in the preset retrieval library and the search text in the search log are processed, and a content completion data set is constructed by using the processed data. Then, the text segments in the content completion data set are vectorized to train the preset language model through the obtained several text tensors, and the model loss of the preset language model is calculated, so as to update the language model according to the obtained model loss, and when using the updated model to provide a completion text, the search text is collected at the same time to perform the next round of update. In this way, the model can be updated online through the user search data and the retrieval library text data, improving the accuracy of query prompts and the user experience.
[0080] In some embodiments, the log generation module 11 may specifically include:
[0081] A data acquisition unit, configured to acquire the search text input by the user end and determine the query time and user identifier corresponding to the search text;
[0082] A first log generation unit, configured to determine whether the search text is the same as the actual query text. If not, determine the target ranking of the actual query text among several recommended texts, and determine the actual query text as the log text, so as to generate a first search log based on the query time, the user identifier, the search text, the log text, and the target ranking; the several recommended texts are recommended texts related to the search text generated by the preset language model; the actual query text is the search text determined by the user end from the several recommended texts;
[0083] A second log generation unit, which is configured to, if yes, determine the search text as the log text, and generate a second search log based on the query time, the user identifier, and the log text.
[0084] In some embodiments, the dataset construction module 12 may specifically include:
[0085] A quantity judgment unit, which is configured to count the log quantity of all the search logs locally, and judge whether the log quantity is not less than several integer multiples of a preset quantity threshold;
[0086] A data processing unit, which is configured to, if yes, determine the log text in all the search logs and the retrieval library text in the preset retrieval library, and add a start symbol and an end symbol to the log text and the retrieval library text to obtain the added text;
[0087] A first dataset construction unit, which is configured to perform an interception process on the added text respectively based on each text interception length in a preset interception range, and construct a content completion dataset based on the obtained several text segments.
[0088] In some embodiments, the content completion device based on the auto-completion model may further include:
[0089] A first dataset construction unit, which is configured to, if not, jump to the step of obtaining the search text input by the user terminal, and process the search text based on a preset log generation rule to generate a search log until the log quantity of all the search logs locally is not less than several integer multiples of the preset quantity threshold, and construct the content completion dataset based on the search log.
[0090] In some embodiments, the loss calculation module 13 may specifically include:
[0091] A tensor determination unit, which is configured to perform vectorization processing on the several text segments in the content completion dataset, and obtain a first text tensor corresponding to the several text segments, a second text tensor corresponding to the corresponding labels of the several text segments, and a third text tensor corresponding to the prediction directions of the several text segments based on a preset dictionary index correspondence table;
[0092] A tensor dimensionality reduction unit, which is configured to perform dimensionality reduction and feature extraction on the first text tensor to obtain a processed first text tensor;
[0093] A loss calculation sub-module, which is configured to perform a max pooling operation on the processed first text tensor, and calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor.
[0094] In some embodiments, the loss calculation sub-module may specifically include:
[0095] The first tensor processing unit is configured to input the processed first text tensor into the Maxpooling layer to perform a max pooling operation on the processed first text tensor to obtain a processed first text tensor;
[0096] The second tensor processing unit is configured to input the processed first text tensor into the first linear classification layer and the second linear classification layer respectively to obtain corresponding first and second output tensors;
[0097] The loss calculation unit is configured to calculate the cross-entropy loss of the preset language model based on the first output tensor and the second text tensor to obtain a target cross-entropy loss; calculate the mean square error of the preset language model based on the second output tensor and the third text tensor, and use the obtained mean square error as the prediction loss of the preset language model;
[0098] The loss determination unit is configured to use the sum of the target cross-entropy loss and the prediction loss as the overall model loss of the preset language model.
[0099] In some embodiments, the model update module 14 may specifically include:
[0100] The first model update unit is configured to update the preset language model based on the overall model loss to obtain an updated language model;
[0101] The text completion unit is configured to determine whether the current search text input by the user terminal is received. If so, generate a plurality of automatically completed text corresponding to the current search text based on the updated language model;
[0102] The second model update unit is configured to re-jump to the step of obtaining the search text input by the user terminal and processing the search text based on a preset log generation rule to generate a search log for the next round of model update.
[0103] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of the present application.
[0104] Figure 5 Schematic diagram of the structure of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the content completion method based on the auto-completion model disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0105] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.
[0106] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0107] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the content completion method based on the auto-completion model executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks.
[0108] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the content completion method based on the auto-completion model disclosed above. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0109] In this specification, the various embodiments are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0110] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0111] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0112] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0113] The technical solutions provided in this application have been introduced in detail above. Specific examples have been used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.< / start> < / end> < / end> < / end> < / end> < / end> < / start> < / end> < / start> < / start> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / end> < / start> < / end> < / start> < / end> < / start> < / end> < / start>
Claims
1. A content completion method based on an auto-completion model, characterized in that, Including: Obtain the search text input by the client, and process the search text based on a preset log generation rule to generate a search log; Count the number of logs in the search log, and judge whether the number of logs meets a preset condition based on a preset quantity threshold. If so, add start and end symbols to the log text in the search log and the retrieval library text in the preset retrieval library, and intercept the obtained added text to construct a content completion data set based on the obtained several text segments; Perform vectorization processing on the several text segments in the content completion data set, and input the obtained several text tensors into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors; Update the preset language model based on the overall model loss, use the obtained updated model to provide automatic completion text for the client, and re-jump to the step of obtaining the search text input by the client and processing the search text based on a preset log generation rule to generate a search log, so as to perform the next round of model update.
2. The content completion method based on the auto-completion model according to claim 1, characterized in that, The step of obtaining the search text input by the client and processing the search text based on a preset log generation rule to generate a search log includes: Obtain the search text input by the client, and determine the query time and user identifier corresponding to the search text; Determine whether the search text is the same as the actual query text. If not, determine the target ranking of the actual query text in several recommended texts, and determine the actual query text as the log text to generate a first search log based on the query time, the user identifier, the search text, the log text, and the target ranking; the several recommended texts are recommended texts related to the search text generated by the preset language model; the actual query text is the search text determined by the client from the several recommended texts; If so, determine the search text as the log text to generate a second search log based on the query time, the user identifier, and the log text.
3. The content completion method based on an auto - completion model according to claim 1, wherein, The step of counting the number of logs in the search log and judging whether the number of logs meets a preset condition based on a preset quantity threshold. If so, add start and end symbols to the log text in the search log and the retrieval library text in the preset retrieval library, and intercept the obtained added text to construct a content completion data set based on the obtained several text segments includes: Count the number of logs of all the search logs locally, and judge whether the number of logs is not less than several integer multiples of the preset quantity threshold; If so, determine the log text in all the search logs and the retrieval library text in the preset retrieval library, and add a start symbol and an end symbol to the log text and the retrieval library text to obtain the added text; Perform interception processing on the added text respectively based on each text interception length in the preset interception range to construct a content completion data set based on the obtained several text segments.
4. The content completion method based on the auto-completion model according to claim 3, wherein After counting the number of logs of all the search logs locally and determining whether the number of logs is not less than several integer multiples of a preset number threshold, it further includes: If not, jump to the step of obtaining the search text input by the client and processing the search text based on a preset log generation rule to generate search logs until the number of logs of all the search logs locally is not less than several integer multiples of the preset number threshold, so as to construct the content completion data set based on the search logs.
5. The content completion method based on the auto - completion model according to claim 1, characterized in that, The vectorization processing of the several text fragments in the content completion data set and inputting the obtained several text tensors into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors includes: Performing vectorization processing on the several text fragments in the content completion data set, and obtaining a first text tensor corresponding to the several text fragments, a second text tensor corresponding to the corresponding labels of the several text fragments, and a third text tensor corresponding to the prediction directions of the several text fragments based on a preset dictionary index correspondence table; Performing dimensionality reduction and feature extraction on the first text tensor to obtain a processed first text tensor; Performing a max pooling operation on the processed first text tensor to calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor.
6. The content completion method based on the auto - completion model according to claim 5, wherein, The performing a max pooling operation on the processed first text tensor to calculate the overall model loss of the preset language model based on the obtained output tensor, the second text tensor, and the third text tensor includes: Inputting the processed first text tensor into a Maxpooling layer to perform a max pooling operation on the processed first text tensor to obtain a processed first text tensor; Inputting the processed first text tensor into a first linear classification layer and a second linear classification layer respectively to obtain corresponding first output tensor and second output tensor; Calculating the cross-entropy loss of the preset language model based on the first output tensor and the second text tensor to obtain a target cross-entropy loss; calculating the mean square error of the preset language model based on the second output tensor and the third text tensor, and using the obtained mean square error as the prediction loss of the preset language model; Using the sum value of the target cross-entropy loss and the prediction loss as the overall model loss of the preset language model.
7. The content completion method based on an auto-complete model according to any one of claims 1 to 6, characterized in that Updating the preset language model based on the overall model loss, using the obtained updated model to provide automatic completion text for the client, and re-jumping to the step of obtaining the search text input by the client and processing the search text based on a preset log generation rule to generate search logs to perform the next round of model update, including: Updating the preset language model based on the overall model loss to obtain an updated language model; Determine whether the current search text input by the client is received. If so, generate several auto-completed texts corresponding to the current search text based on the updated language model, and then re-jump to the step of obtaining the search text input by the client, and process the search text based on a preset log generation rule to generate a search log for the next round of model update.
8. A content completion device based on an autocomplete model, characterized in that, It includes: A log generation module, configured to obtain the search text input by the client and process the search text based on a preset log generation rule to generate a search log; A dataset construction module, configured to count the number of logs in the search log and determine whether the number of logs meets a preset condition based on a preset quantity threshold. If so, add start and end symbols to the log text in the search log and the retrieval library text in a preset retrieval library, and intercept the obtained added text to construct a content completion dataset based on the obtained several text segments; A loss calculation module, configured to perform vectorization processing on the several text segments in the content completion dataset, and input the obtained several text tensors into a preset language model to calculate the overall model loss of the preset language model based on the several text tensors; A model update module, configured to update the preset language model based on the overall model loss, use the obtained updated model to provide auto-completed texts for the client, and then re-jump to the step of obtaining the search text input by the client, and process the search text based on a preset log generation rule to generate a search log for the next round of model update.
9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the content completion method based on an auto-completion model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program, when executed by a processor, implements the content completion method based on an auto-completion model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Language model training system, a voice identification system and corresponding method
CN103871402A
Query automatic completion method, device and equipment and computer storage medium
CN111222058A