Multi-dimensional category enhanced word vector fusion method and system suitable for digital human
By establishing a standard problem database and performing positive and negative sample annotation, combining gradient descent algorithm and multi-dimensional category weight optimization word vectors, the problem of low accuracy of question-answer matching in the existing technology is solved, and more efficient and accurate question-and-answer matching is achieved.
Patent Information
- Application Number
- CN202510063710.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-15
AI Technical Summary
When building user interest models, it is difficult for the prior art to effectively reflect the distribution of characteristic words and is easily affected by the skew of the data set, resulting in low accuracy of question-and-answer matching.
Improve the quality of model training data by establishing a database containing standard questions and their corresponding answers, and performing positive and negative sample annotations. Combined with the gradient descent algorithm, optimize word vectors, incorporate multi-dimensional category weights, and generate extended vectors to improve matching accuracy.
It significantly improves the accuracy and response speed of question-and-answer matching, enhances the expression ability and distinction of word vectors, can better adapt to different categories of problems, and improves the flexibility and adaptability of overall representation.
Smart Images

Figure CN120011502A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of user interest model construction, and in particular to a multi-dimensional category enhanced word vector fusion method and system suitable for digital humans. Background Art
[0002] At present, due to the rapid development of science and technology and the Internet, online information resources are growing rapidly, and the resources available to netizens are becoming more and more abundant. Any questions that they do not understand can be found on the Internet. As the information injected into the Internet continues to increase, there is also a lot of duplicate information. Duplicate web page information is of little significance to search engines. Duplicate information not only occupies excessive storage resources, but also makes it difficult for people to obtain the information they really need, affecting the user's browsing experience. Therefore, in the massive amount of data, it is particularly important to efficiently and accurately check the similarity of sentences.
[0003] The Chinese patent with publication number CN114997181A discloses an intelligent question-answering method and system based on user feedback correction, step 1, the user inputs the question, performs text segmentation and feature word extraction; step 2, vectorizes the feature word extraction result and the question-answering library text, and calculates the cosine similarity of the text to generate a preliminary matching result; step 3, inputs the user input text vector and the user feedback database vector into the deep learning network model, and performs the user feedback correction process; step 4, outputs the final answer text to the user based on the results of the comprehensive similarity matching and user feedback correction model. However, the above scheme cannot effectively reflect the distribution of feature words and is easily affected by the skewness of the data set, thereby affecting the calculation of question similarity. Therefore, it is very necessary to provide a multi-dimensional category enhanced word vector fusion method and system suitable for digital people to improve the accuracy of question-answer matching. Summary of the invention
[0004] In view of this, the present invention proposes a multi-dimensional category enhanced word vector fusion method and system suitable for digital humans. By establishing a database containing standard questions and their corresponding answers and annotating positive and negative samples, the quality of model training data is improved, thereby enhancing the accuracy of question-answer matching.
[0005] The present invention provides a multi-dimensional category enhanced word vector fusion method applicable to digital humans, the method comprising:
[0006] Constructing a question-and-answer database and collecting user input questions, wherein the question-and-answer database includes a set of standard questions and an answer corresponding to each standard question in the set of standard questions;
[0007] Constructing a question sample pair according to the standard question corresponding to the user input question and the user input question, and marking each question sample pair in turn to obtain a positive and negative sample set;
[0008] Performing similarity calculation on the positive and negative sample sets based on a gradient descent algorithm to obtain word vector scaling parameters and category weight parameters corresponding to the positive and negative sample sets, and forming a first extended vector corresponding to the user input question according to the word vector scaling parameters, the category weight parameters and the positive and negative sample sets, wherein the similarity includes cosine similarity and dot product similarity;
[0009] The cosine similarity between the first extended vector and the extended vectors corresponding to all standard questions is calculated, the standard question with the highest cosine similarity is selected, and the answer corresponding to the standard question is output to the user.
[0010] Based on the above technical solution, preferably, after constructing the question-answer database and collecting user input questions, the method further includes:
[0011] All professional terms in the standard question set are extracted, and all professional terms are classified in multiple dimensions to obtain the one-hot encoding vectors corresponding to the standard questions in the standard question set under each dimension.
[0012] On the basis of the above technical solution, preferably, the obtaining of the word vector scaling parameters and the category weight parameters corresponding to the positive and negative sample sets specifically includes:
[0013] Loading the positive and negative sample sets, initial word vector scaling parameters, and initial category weight parameters into the machine learning model, and using the BERT model to respectively calculate the word vectors corresponding to the user input questions and standard questions in the positive and negative sample sets;
[0014] Calculate the one-hot encoding vectors corresponding to the user input question and the standard question in the positive and negative sample sets, and obtain the extended vector corresponding to the sampled input question according to the one-hot encoding vector and the sampled input question;
[0015] Calculate the similarity between the user input question and the standard question in each question sample pair in the positive and negative sample sets using cosine similarity, and calculate the total loss function according to the transition word vector scaling parameter and the transition category weight parameter;
[0016] When the total loss function is greater than the loss threshold or reaches the maximum number of iterations, the machine learning model outputs the word vector scaling parameter and the category weight parameter.
[0017] More preferably, the obtaining the extended vector corresponding to the sample input problem specifically includes:
[0018] Selecting a sampled input question from any question sample pair in the positive and negative sample sets, wherein the sampled input question includes a first user input question and a first standard question corresponding to the first user input question;
[0019] Calculate the sample input question according to the BERT model to obtain a first word vector corresponding to the sample input question;
[0020] An extended vector corresponding to the sampled input problem is obtained according to the transition word vector scaling parameter, the transition category weight parameter, the first word vector, and the one-hot encoding vector corresponding to the sampled input problem.
[0021] More preferably, the expression of the total loss function is:
[0022]
[0023] L total =L+L reg
[0024] in, represents the cosine similarity between the user input question u and the standard question w in the positive and negative sample sets, represents the expansion vector of the user input question u, represents the expanded vector of the standard problem w, · represents the dot product, |||| represents the Euclidean norm of the vector, L represents the contrastive loss function, N represents the total number of samples, y i represents the sample label, represents the expansion vector of the i-th user input question u, represents the expansion vector of the i-th standard problem w, represents the cosine similarity between the ith user input question u and the ith standard question w, γ represents the boundary parameter that controls the upper limit of the similarity of negative samples, max() represents the maximum value function, L reg represents the regularization function, λ represents the regularization coefficient, q i represents the weight parameter of the i-th category, m represents the dimension corresponding to the weight parameter of the i-th category, k represents the word vector scaling parameter, L total Represents the total loss function.
[0025] More preferably, the selecting the standard question with the highest cosine similarity and outputting the answer corresponding to the standard question to the user specifically includes:
[0026] calculating a first cosine similarity between the first extended vector and an extended vector corresponding to the currently selected standard question;
[0027] Determine whether the currently stored maximum similarity value is less than the first cosine similarity, and if the currently stored maximum similarity value is less than the first cosine similarity, assign the first cosine similarity to the currently stored maximum similarity value;
[0028] It is determined whether all the extended vectors of the standard question have been traversed. If all the extended vectors have been traversed, the answer to the standard question corresponding to the first cosine similarity is output.
[0029] More preferably, the expressions of the word vector scaling parameter and the category weight parameter are respectively:
[0030]
[0031] Among them, k (t+1) represents the word vector scaling parameter after the t+1th iteration, represents the weight parameter of the i-th category after the t+1th iteration, q i represents the weight parameter of the i-th category, k represents the word vector scaling parameter, and L total represents the total loss function, t represents the number of steps in the current iteration, and η represents the learning rate.
[0032] In a second aspect of the present application, a multi-dimensional category enhanced word vector fusion system applicable to digital humans is provided, wherein the multi-dimensional category enhanced word vector fusion system comprises a data acquisition module, a sample expansion module and an answer output module, wherein:
[0033] The data collection module is used to build a question-answer database and collect user input questions, wherein the question-answer database includes a set of standard questions and an answer corresponding to each standard question in the set of standard questions;
[0034] The sample expansion module is used to construct a question sample pair according to the standard question corresponding to the user input question and the user input question, mark each question sample pair in turn to obtain a positive and negative sample set, perform similarity calculation on the positive and negative sample sets based on a gradient descent algorithm to obtain a word vector scaling parameter and a category weight parameter corresponding to the positive and negative sample sets, and form a first extended vector corresponding to the user input question according to the word vector scaling parameter, the category weight parameter and the positive and negative sample sets, wherein the similarity includes cosine similarity and dot product similarity;
[0035] The answer output module is used to calculate the cosine similarity between the first extended vector and the extended vectors corresponding to all standard questions, select the standard question with the highest cosine similarity, and output the answer corresponding to the standard question to the user.
[0036] In a third aspect of the present application, an electronic device is provided, comprising a processor, a memory, a user interface and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory.
[0037] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of a multi-dimensional category enhanced word vector fusion method applicable to digital humans.
[0038] The multi-dimensional category enhanced word vector fusion method and system applicable to digital humans provided by the present invention have the following beneficial effects compared with the prior art:
[0039] (1) By establishing a database containing standard questions and their corresponding answers, the integrity and accuracy of the data are ensured, the possibility of matching errors is reduced, and question sample pairs are constructed and positive and negative samples are labeled, which improves the quality of the model training data and enhances the accuracy of question-answer matching. The word vector is optimized by combining the gradient descent algorithm, and the multi-dimensional category weight is incorporated to enhance the expressive power of the word vector, so that it not only contains semantic information but also category information, thereby improving the discrimination and recognition ability of the vector. At the same time, by optimizing the scaling parameters and category weight parameters, the word vector can better adapt to different categories of questions, improving the flexibility and adaptability of the overall representation. The cosine similarity is used to quickly calculate the matching degree between the user question and the standard question. Combined with the optimized extended vector, an efficient matching process is achieved, which significantly improves the response speed of the question-answering process.
[0040] (2) By combining the multidimensional category vector with the original word vector to generate an extended vector, rich semantic information is retained and category information in professional fields is introduced, making the vector representation more comprehensive and detailed. The word vector scaling parameters are optimized through the gradient descent algorithm to ensure that the expressive power of the word vector in different categories is enhanced, and the discrimination and recognition of the vector in a specific context are improved. The cosine similarity is used to calculate the similarity between the user input question and the standard question. Combined with the optimized extended vector, the subtle differences in semantics and categories are effectively captured, which improves the matching accuracy. By optimizing the category weight parameters, the system can be compatible with and effectively handle multi-classification problems, ensuring that a high level of matching performance is maintained in a multi-category context. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 A schematic diagram of the flow of the multi-dimensional category enhanced word vector fusion method applicable to digital humans provided by the present invention;
[0043] Figure 2 A schematic diagram of the structure of the multi-dimensional category enhanced word vector fusion system provided by the present invention;
[0044] Figure 3 This is a schematic structural diagram of an electronic device provided by the present invention.
[0045] Explanation of the accompanying drawings: 1. Multi-dimensional category enhanced word vector fusion system; 11. Data acquisition module; 12. Sample expansion module; 13. Sample expansion module; 2. Electronic device; 21. Processor; 22. Communication bus; 23. User interface; 24. Network interface; 25. Memory. DETAILED DESCRIPTION
[0046] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] The present invention discloses a multi-dimensional category enhanced word vector fusion method suitable for digital human, referring to Figure 1 The steps of the method include steps S1 to S4.
[0048] Step S1, constructing a question-answering database and collecting user input questions, wherein the question-answering database includes a set of standard questions and an answer corresponding to each standard question in the set of standard questions.
[0049] In this step, prepare all standard questions and their corresponding answers, and establish a mapping relationship between the two. Extract professional terms from the standard question set, classify these professional terms in multiple dimensions, determine the number of classification dimensions, and determine the order of the elements of the classification set in each dimension, and then construct a unique hot vector for each dimension. Finally, obtain positive and negative samples in the original system, with the number of samples in each category exceeding m*100, and make the number of positive and negative samples equal.
[0050] In this embodiment, after constructing the question-answer database and collecting user input questions, it also includes extracting all professional terms in the standard question set and classifying all professional terms in multiple dimensions to obtain the unique hot encoding vector corresponding to the standard question in the standard question set under each dimension.
[0051] Furthermore, all standard questions and their corresponding answers are collected from existing resources, databases or documents. A clear correspondence is established between each standard question and the answer to ensure that they are correctly matched. Duplicates are removed and the collated data is stored in the database.
[0052] Manually identify and extract all professional terms from the set of standard questions. Determine the number of dimensions for classification, with each dimension representing a classification standard, such as subject, variety, application field, etc. Under each dimension, classify the professional terms, and determine the order of the elements of the classification set in each dimension to establish a clear classification hierarchy. For example, under the variety dimension, establish a classification of rice varieties, and the professional terms are in the order of: "Nanjing 9108", "Nanjing 5718", "Jingliangyou 534", and "Jingliangyou Huazhan". If a standard question contains "Jingliangyou 534", the one-hot encoding vector of the question on this dimension is [0,0,1,0].
[0053] Step S2, constructing question sample pairs according to the standard questions corresponding to the user input questions and the user input questions, and marking each question sample pair in turn to obtain a set of positive and negative samples.
[0054] In this step, we collect the questions raised by users and the standard questions answered by the digital human in the past, and establish the corresponding relationship between them. Through manual judgment, we determine whether each pair of user questions and standard questions matches, and mark them as positive samples or negative samples. Among the collected sample pairs, we select m*100 positive sample pairs and m*100 negative sample pairs for each classification dimension (where m is the number of dimensions of multidimensional classification), ensuring that the number of positive and negative samples is equal and that the samples can fully cover all elements in each classification.
[0055] Step S3, calculate the similarity of the positive and negative sample sets based on the gradient descent algorithm to obtain the word vector scaling parameters and category weight parameters corresponding to the positive and negative sample sets, and form a first extended vector corresponding to the user input question based on the word vector scaling parameters, category weight parameters and positive and negative sample sets, wherein the similarity includes cosine similarity and dot product similarity.
[0056] In this step, set the parameters required for the machine learning model, including the negative sample similarity upper limit parameter γ, regularization coefficient λ, learning rate η, loss threshold, and maximum number of iterations. Once set, these parameters will remain unchanged during the calculation process unless learning is restarted. Load all positive and negative sample data, mainly covering the following three parts: user input questions, screened standard questions; judgment labels (for positive samples, the label is 1; for negative samples, the label is 0). If there is an empirical formula, initialize the parameter value to the square root of the corresponding empirical formula; if there is no empirical formula, initialize the parameters k and q i Initialize to 1. Initialize the iteration counter t to 1.
[0057] This step also includes steps S31 to S34.
[0058] Step S31, load the positive and negative sample sets, initial word vector scaling parameters and initial category weight parameters into the machine learning model, and use the BERT model to calculate the word vectors corresponding to the user input questions and standard questions in the positive and negative sample sets respectively.
[0059] In this step, both standard questions and user input questions are converted into 768-dimensional word vectors through a unified model (such as the BERT model). Let each standard question be w, then the word vector obtained by the BERT model for standard question w is V w .
[0060] Assume that there are m categories, each of which contains multiple subcategories. For example, the place classification includes "Beijing", "Shanghai", etc.; the variety classification includes "Nanjing 9108", "Nanjing 5718", etc.; the purpose classification includes "commercial", "home use", etc. Each category can be represented as a one-hot vector C wi .
[0061] If a question belongs to multiple categories, it corresponds to multiple category vectors. By combining the original word vector with these category vectors, a new extended vector is formed.
[0062]
[0063] Among them, k represents the scaling parameter of the word vector, q i Represents the weight parameter of the i-th category, the one-hot encoding vector C of the standard problem w wi When the standard question w contains a term of a certain category, the unique hot vector element corresponding to the category is 1, otherwise it is 0, and m is the dimension of the category. The plus sign here indicates expansion in the vector dimension, that is, adding a new dimension to construct the expanded vector
[0064] After the user enters the question, there are two main methods to filter the standard questions by similarity. One is to use the dot product as the similarity metric and select the standard question with the greatest similarity; the other is to use cosine similarity for independent evaluation.
[0065] The dot product is a direct similarity measurement method with obvious linear characteristics, which is suitable for linear combination indicators. Set T as the measurement indicator of dot product similarity.
[0066] If the user input question is u, the expanded vector is The expansion vector of the standard problem w is The calculation formula of dot product similarity is as follows:
[0067]
[0068] A common method is to set the weights of word vectors and the weights of each category through experience. The calculation formula is usually:
[0069]
[0070] Among them, a represents the semantic weight, b i Represents the weights corresponding to each category and satisfies:
[0071]
[0072] Since all parameters are non-negative, it can be deduced that:
[0073]
[0074] Therefore, when using the dot product as a similarity measure to screen the standard problem, there is the above relationship between it and the empirical formula. If the empirical formula is reliable and the dot product is used for screening, k and q can be easily determined i , thereby quickly constructing a new expansion vector
[0075] In one example, since dot product similarity is very sensitive to the length of the vector, it cannot directly measure the directional similarity between vectors, and its result range is not limited, so cosine similarity can also be used for measurement. In the screening process, cosine similarity is used for measurement. However, the calculation parameters k and q i The method is also applicable to the calculation of dot product similarity.
[0076] If the user inputs question u and standard question w, then the cosine similarity The calculation formula is:
[0077]
[0078] in, They represent the expanded vector representations of the user input question u and the standard question w respectively, · represents the dot product, and |||| represents the Euclidean norm of the vector.
[0079] The designed contrast loss function L is as follows:
[0080]
[0081] Where N is the total number of samples. i Represents the sample label. When the i-th sample is a positive sample pair, it is 1; when it is a negative sample pair, it is 0; represents the expansion vector of the i-th user input question, represents the expansion vector of the ith standard problem, γ represents the boundary parameter that controls the upper limit of the similarity of negative samples, and the boundary parameter is set to 0.1. The value can be adjusted between 0.01 and 0.2 as needed, but cannot be less than 0.01 or more than 0.2.
[0082] When i = 1 (positive sample pair), the goal is to make As close to 1 as possible, so the loss term is
[0083] When i = 0 (negative sample pair), the goal is to make Therefore, the loss term is
[0084] In order to prevent overfitting caused by excessive parameters, the L2 regularization term is introduced, and its formula is as follows:
[0085]
[0086] Among them, λ is the regularization coefficient, which is generally set to 0.1, but can be adjusted as needed, and m is the parameter q i The dimension is the number of categories.
[0087] The total loss function is the sum of the contrast loss and the regularization term, that is:
[0088] L total =L+L reg
[0089] The optimization goal is to find the parameters k and q i =[q1,q2,…,q m ], so that the total loss function L total minimize.
[0090] In order to minimize the loss function, using the gradient descent method, it is necessary to calculate the loss function with respect to the parameters k and q i The partial derivative of .
[0091] Partial derivative with respect to parameter k:
[0092]
[0093] in It is the partial derivative of k in the contrast loss function, which is calculated by the chain rule for the positive sample part (similar to the negative sample part):
[0094]
[0095] For parameter q i The partial derivative of :
[0096]
[0097] Likewise, It is calculated by the chain rule of the contrast loss function. For the positive sample part (similar to the negative sample part), we have:
[0098]
[0099] The partial derivative of cosine similarity with respect to the parameters is as follows:
[0100]
[0101] Among them, C i Represents a vector of the same dimension as the extended vector, where the corresponding positions of the i-th category are all 1 and the remaining positions are 0.
[0102] Step S32, calculate the one-hot encoding vectors corresponding to the user input questions and the standard questions in the positive and negative sample sets, and obtain the expansion vector corresponding to the sampled input questions based on the one-hot encoding vectors and the sampled input questions.
[0103] In this step, the BERT model is first used to calculate the word vectors of user questions and standard questions in all samples. Then, according to the constructed multidimensional classification, the corresponding professional terms are retrieved in user questions and standard questions. If they exist, the corresponding one-hot encoding variable is set to 1; if they do not exist, the corresponding one-hot encoding variable is set to 0. In this way, the one-hot encoding variable of the corresponding question is obtained. Next, the parameters k and q are used to calculate the corresponding professional terms. i , generate the expanded vectors of user questions and standard questions in all samples.
[0104] Select a sampled input question from any question sample pair in the positive and negative sample sets, where the sampled input question includes a first user input question and a first standard question corresponding to the first user input question; calculate the sampled input question according to the BERT model to obtain a first word vector corresponding to the sampled input question; obtain an expanded vector corresponding to the sampled input question according to a transition word vector scaling parameter, a transition category weight parameter, the first word vector and a one-hot encoding vector corresponding to the sampled input question.
[0105] Furthermore, in the question sample pair, there are user questions and screened standard questions, both of which can be used as inputs of the expansion vector. Here, one of the questions is randomly selected as an example. In the sample pair, there are user questions and screened standard questions, both of which can be used as inputs of the expansion vector. Here, one of the questions is randomly selected as an example. In the sample pair, there are user questions and screened standard questions, both of which can be used as inputs of the expansion vector. First, all dimensions of the one-hot encoding vector of the question are initialized to 0. Then, using the professional terms in multidimensional classification, search for each professional term in the question: if a professional term is found in the question, the value of the corresponding one-hot encoding vector is set to 1; if not found, the corresponding value remains 0. After completing the search of all professional terms, the one-hot vector of the question is obtained. Using parameters k and q i , multiply them with the corresponding word vectors and one-hot vector elements one by one to generate an expanded vector with a dimension greater than 768.
[0106] Step S33, using cosine similarity to calculate the similarity between the user input question and the standard question in each question sample pair in the positive and negative sample sets, and calculating the total loss function according to the transition word vector scaling parameter and the transition category weight parameter.
[0107] Given the current parameters k and q i In the case of , first calculate the regularization loss L reg Then, the contrast loss L obtained in the previous step is combined with the regularization loss L reg Add together to get the total loss L total The total loss is expressed as i , the overall loss value of the machine learning model.
[0108] In this step, the expression of the total loss function is:
[0109]
[0110] L total =L+L reg
[0111] in, represents the cosine similarity between the user input question u and the standard question w in the positive and negative sample sets, represents the expansion vector of the user input question u, represents the expanded vector of the standard problem w, · represents the dot product, |||| represents the Euclidean norm of the vector, L represents the contrastive loss function, N represents the total number of samples, y i represents the sample label, represents the expansion vector of the i-th user input question u, represents the expansion vector of the i-th standard problem w, represents the cosine similarity between the ith user input question u and the ith standard question w, γ represents the boundary parameter that controls the upper limit of the similarity of negative samples, max() represents the maximum value function, L reg represents the regularization function, λ represents the regularization coefficient, q i represents the weight parameter of the i-th category, m represents the dimension corresponding to the weight parameter of the i-th category, k represents the word vector scaling parameter, L total Represents the total loss function.
[0112] Step S34: When the total loss function is greater than the loss threshold or reaches the maximum number of iterations, the machine learning model outputs the word vector scaling parameters and category weight parameters.
[0113] In this step, the expressions of the word vector scaling parameter and the category weight parameter are:
[0114]
[0115] Among them, k (t+1) represents the word vector scaling parameter after the t+1th iteration, represents the weight parameter of the i-th category after the t+1th iteration, q i represents the weight parameter of the i-th category, k represents the word vector scaling parameter, and L total represents the total loss function, t represents the number of steps in the current iteration, and η represents the learning rate.
[0116] By combining the multidimensional category vector with the original word vector to generate an extended vector, rich semantic information is retained and category information in professional fields is introduced, making the vector representation more comprehensive and detailed. The word vector scaling parameters are optimized through the gradient descent algorithm to ensure that the expressive ability of the word vector in different categories is enhanced, and the distinction and recognition of the vector in a specific context are improved. The cosine similarity is used to calculate the similarity between the user input question and the standard question. Combined with the optimized extended vector, it effectively captures the subtle differences in semantics and categories and improves the matching accuracy. Through the optimization of the category weight parameters, the system is compatible with and effectively handles multi-classification problems, ensuring that a high level of matching performance is maintained in a multi-category context.
[0117] Step S4, calculating the cosine similarity between the first extended vector and the extended vectors corresponding to all standard questions, selecting the standard question with the highest cosine similarity, and outputting the answer corresponding to the standard question to the user.
[0118] This step also includes steps S41 to S43.
[0119] Step S41, calculating a first cosine similarity between the first extended vector and the extended vector corresponding to the currently selected standard question.
[0120] Step S42, determining whether the currently stored maximum similarity value is less than the first cosine similarity, if the currently stored maximum similarity value is less than the first cosine similarity, assigning the first cosine similarity to the currently stored maximum similarity value.
[0121] Step S43, determining whether all the extended vectors of the standard question have been traversed, and if all the extended vectors have been traversed, outputting the answer to the standard question corresponding to the first cosine similarity.
[0122] In an example, the steps to get the answer based on the user input question are as follows:
[0123] Step 1, define two variables tmp and final, and initialize them to 0. tmp is used to store the maximum similarity value currently found. final is used to store the number of the standard question with the maximum similarity to the user question extension vector. After completion, go to step 2.
[0124] Step 2: The user enters the question to be asked and submits it to the system backend. After completion, proceed to step 3.
[0125] Step 3: Generate an extended vector of the question based on the question input by the user using the method described in step 5 in part 2. After completion, proceed to step 4.
[0126] Step 4: Select an extension vector from the set of standard problem extension vectors: if this step is performed for the first time, select the first one; if it is not performed for the first time, select the next one after the previous one. After completion, proceed to step 5.
[0127] Step 5, calculate the cosine similarity between the expansion vector of the user question and the expansion vector of the currently selected standard question, and store the calculation result in the variable Sim. After completion, go to step 6.
[0128] Step 6: If Sim≥tmp, it means that the currently calculated similarity is greater than or equal to the previous maximum value, and go to step 7; otherwise, go to step 9.
[0129] Step 7, assign the value of Sim to tmp, update the maximum similarity, so that tmp always stores the maximum similarity. After completion, proceed to step 8.
[0130] Step 8: Assign the number of the current standard question to final, and update the standard question number so that it is the number with the maximum similarity. After completion, proceed to step 9.
[0131] Step 9, determine whether all the extension vectors of the standard questions have been traversed and processed. If not, return to step 4 to continue processing the next standard question extension vector; otherwise, all the extension vectors have been traversed, proceed to step 10.
[0132] Step 10: Find the corresponding standard answer according to the standard question number stored in the variable final, and output the answer to the user.
[0133] Through the above steps, the system accurately finds the standard question that best matches the question entered by the user, and provides the corresponding standard answer to the user. This process ensures that users can get the most relevant and accurate answers. By establishing a database containing standard questions and their corresponding answers, the integrity and accuracy of the data are ensured, the possibility of matching errors is reduced, and the quality of model training data is improved by constructing question sample pairs and annotating positive and negative samples, thereby enhancing the accuracy of question-answer matching. The word vector is optimized by combining the gradient descent algorithm, and the multi-dimensional category weight is integrated to enhance the expression ability of the word vector, so that it not only contains semantic information but also covers category information, improving the discrimination and recognition ability of the vector. At the same time, by optimizing the scaling parameters and category weight parameters, the word vector can better adapt to different categories of questions, improving the flexibility and adaptability of the overall representation, using cosine similarity to quickly calculate the matching degree between user questions and standard questions, combined with the optimized extended vector, an efficient matching process is achieved, and the response speed of the question-answer process is significantly improved.
[0134] Based on the above method, the embodiment of the present application discloses a multi-dimensional category enhanced word vector fusion system suitable for digital humans, referring to Figure 2 The multi-dimensional category enhanced word vector fusion system 1 includes a data acquisition module 11, a sample expansion module 12 and an answer output module 13, wherein:
[0135] The data collection module 11 is used to construct a question-answering database and collect user input questions, wherein the question-answering database includes a set of standard questions and an answer corresponding to each standard question in the set of standard questions;
[0136] The sample expansion module 12 is used to construct a question sample pair according to the standard question corresponding to the user input question and the user input question, mark each question sample pair in turn to obtain a positive and negative sample set, perform similarity calculation on the positive and negative sample sets based on the gradient descent algorithm to obtain a word vector scaling parameter and a category weight parameter corresponding to the positive and negative sample sets, and form a first expansion vector corresponding to the user input question according to the word vector scaling parameter, the category weight parameter and the positive and negative sample sets;
[0137] The answer output module 13 is used to calculate the cosine similarity between the first extended vector and the extended vectors corresponding to all standard questions, select the standard question with the highest cosine similarity, and output the answer corresponding to the standard question to the user.
[0138] In one example, the data collection module 11 is used to extract all professional terms in the standard question set, and perform multi-dimensional classification on all professional terms to obtain the one-hot encoding vector corresponding to the standard question in the standard question set under each dimension.
[0139] In one example, the sample expansion module 12 is used to load the positive and negative sample sets, the initial word vector scaling parameters and the initial category weight parameters in the machine learning model, and use the BERT model to respectively calculate the word vectors corresponding to the user input questions and the standard questions in the positive and negative sample sets; calculate the one-hot encoding vectors corresponding to the user input questions and the standard questions in the positive and negative sample sets, and obtain the expansion vector corresponding to the sampled input questions according to the one-hot encoding vectors and the sampled input questions; use cosine similarity to calculate the similarity between the user input questions and the standard questions in each question sample pair in the positive and negative sample sets, and calculate the total loss function according to the transition word vector scaling parameters and the transition category weight parameters; when the total loss function is greater than the loss threshold or reaches the maximum number of iterations, the machine learning model is made to output the word vector scaling parameters and the category weight parameters. .
[0140] In one example, the sample expansion module 12 is used to select a sampled input question from any question sample pair in a positive and negative sample set, wherein the sampled input question includes a first user input question and a first standard question corresponding to the first user input question; calculate the sampled input question according to the BERT model to obtain a first word vector corresponding to the sampled input question; obtain an expansion vector corresponding to the sampled input question according to a transition word vector scaling parameter, a transition category weight parameter, the first word vector, and a one-hot encoding vector corresponding to the sampled input question.
[0141] In one example, the expression of the total loss function is:
[0142]
[0143] L total =L+L reg
[0144] in, represents the cosine similarity between the user input question u and the standard question w in the positive and negative sample sets, represents the expansion vector of the user input question u, represents the expanded vector of the standard problem w, · represents the dot product, |||| represents the Euclidean norm of the vector, L represents the contrastive loss function, N represents the total number of samples, y i represents the sample label, represents the expansion vector of the i-th user input question u, represents the expansion vector of the i-th standard problem w, represents the cosine similarity between the ith user input question u and the ith standard question w, γ represents the boundary parameter that controls the upper limit of the similarity of negative samples, max() represents the maximum value function, L reg represents the regularization function, λ represents the regularization coefficient, q i represents the weight parameter of the i-th category, m represents the dimension corresponding to the weight parameter of the i-th category, k represents the word vector scaling parameter, L total Represents the total loss function.
[0145] In one example, the answer output module 13 is used to calculate the first cosine similarity between the first extended vector and the extended vector corresponding to the currently selected standard question; determine whether the currently stored maximum similarity value is less than the first cosine similarity, if the currently stored maximum similarity value is less than the first cosine similarity, assign the first cosine similarity to the currently stored maximum similarity value; determine whether the extended vectors of all standard questions have been traversed, if all extended vectors have been traversed, output the answer to the standard question corresponding to the first cosine similarity.
[0146] In an example, the expressions of the word vector scaling parameter and the category weight parameter are:
[0147]
[0148] Among them, k (t+1) represents the word vector scaling parameter after the t+1th iteration, represents the weight parameter of the i-th category after the t+1th iteration, q i represents the weight parameter of the i-th category, k represents the word vector scaling parameter, and L total represents the total loss function, t represents the number of steps in the current iteration, and η represents the learning rate.
[0149] See also Figure 3 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device 2 may include: at least one processor 21 , at least one network interface 24 , a user interface 23 , a memory 25 , and at least one communication bus 22 .
[0150] The communication bus 22 is used to realize the connection and communication between these components.
[0151] The user interface 23 may include a display screen (Display) and a camera (Camera), and the optional user interface 23 may also include a standard wired interface and a wireless interface.
[0152] The network interface 24 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0153] Among them, the processor 21 may include one or more processing cores. The processor 21 uses various interfaces and lines to connect various parts in the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 25, and calling data stored in the memory 25. Optionally, the processor 21 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 21 can integrate one or more combinations of a central processing unit (Central Processing Unit, CPU), a graphics processor (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 21, and it can be implemented separately through a chip.
[0154] Among them, the memory 25 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 25 includes a non-transitory computer-readable storage medium. The memory 25 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 25 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data involved in the above-mentioned method embodiments, etc. The memory 25 may also be optionally at least one storage device located away from the aforementioned processor 21. As Figure 3 As shown, the memory 25 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a multi-dimensional category enhanced word vector fusion method suitable for digital humans.
[0155] exist Figure 3In the electronic device 2 shown, the user interface 23 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 21 can be used to call a multi-dimensional category enhanced word vector fusion method suitable for digital humans stored in the memory 25. When executed by one or more processors, the electronic device executes one or more methods in the above-mentioned embodiments.
[0156] A computer-readable storage medium stores instructions, which, when executed by one or more processors, enable the computer to execute one or more methods in the above-mentioned embodiments.
[0157] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.
[0158] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0159] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0160] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0162] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: various media that can store program codes, such as USB flash drives, mobile hard drives, magnetic disks or optical disks.
[0163] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereto. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and the truth of practice, it will be easy for those skilled in the art to think of other embodiments of the present disclosure. This application is intended to cover any variation, use or adaptive change of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the technical field that are not recorded in the present disclosure.
[0164] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-dimensional category enhanced word vector fusion method suitable for digital humans, characterized in that: The method comprises: Constructing a question-and-answer database and collecting user input questions, wherein the question-and-answer database includes a set of standard questions and an answer corresponding to each standard question in the set of standard questions; Constructing a question sample pair according to the standard question corresponding to the user input question and the user input question, and marking each question sample pair in turn to obtain a positive and negative sample set; Performing similarity calculation on the positive and negative sample sets based on a gradient descent algorithm to obtain word vector scaling parameters and category weight parameters corresponding to the positive and negative sample sets, and forming a first extended vector corresponding to the user input question according to the word vector scaling parameters, the category weight parameters and the positive and negative sample sets, wherein the similarity includes cosine similarity and dot product similarity; The cosine similarity between the first extended vector and the extended vectors corresponding to all standard questions is calculated, the standard question with the highest cosine similarity is selected, and the answer corresponding to the standard question is output to the user.
2. The multi-dimensional category enhanced word vector fusion method applicable to digital humans as claimed in claim 1, characterized in that: After the question-answer database is constructed and user input questions are collected, the following steps are also included: All professional terms in the standard question set are extracted, and all professional terms are classified in multiple dimensions to obtain the one-hot encoding vectors corresponding to the standard questions in the standard question set under each dimension.
3. The multi-dimensional category enhanced word vector fusion method applicable to digital humans as claimed in claim 1, characterized in that: The obtaining of the word vector scaling parameters and the category weight parameters corresponding to the positive and negative sample sets specifically includes: Loading the positive and negative sample sets, initial word vector scaling parameters, and initial category weight parameters into the machine learning model, and using the BERT model to respectively calculate the word vectors corresponding to the user input questions and standard questions in the positive and negative sample sets; Calculate the one-hot encoding vectors corresponding to the user input question and the standard question in the positive and negative sample sets, and obtain the extended vector corresponding to the sampled input question according to the one-hot encoding vector and the sampled input question; Calculate the similarity between the user input question and the standard question in each question sample pair in the positive and negative sample sets using cosine similarity, and calculate the total loss function according to the transition word vector scaling parameter and the transition category weight parameter; When the total loss function is greater than the loss threshold or reaches the maximum number of iterations, the machine learning model outputs the word vector scaling parameter and the category weight parameter.
4. The multi-dimensional category enhanced word vector fusion method applicable to digital humans as claimed in claim 3, characterized in that: The obtaining of the extended vector corresponding to the sample input problem specifically includes: Selecting a sampled input question from any question sample pair in the positive and negative sample sets, wherein the sampled input question includes a first user input question and a first standard question corresponding to the first user input question; Calculate the sample input question according to the BERT model to obtain a first word vector corresponding to the sample input question; An extended vector corresponding to the sampled input problem is obtained according to the transition word vector scaling parameter, the transition category weight parameter, the first word vector, and the one-hot encoding vector corresponding to the sampled input problem.
5. The multi-dimensional category enhanced word vector fusion method applicable to digital humans as claimed in claim 3, characterized in that: The expression of the total loss function is: L total =L + L reg in, represents the cosine similarity between the user input question u and the standard question w in the positive and negative sample sets, represents the expansion vector of the user input question u, represents the expanded vector of the standard problem w, · represents the dot product, || || represents the Euclidean norm of the vector, L represents the contrastive loss function, N represents the total number of samples, y i represents the sample label, represents the expansion vector of the i-th user input question u, represents the expansion vector of the i-th standard problem w, represents the cosine similarity between the ith user input question u and the ith standard question w, γ represents the boundary parameter that controls the upper limit of the similarity of negative samples, max() represents the maximum value function, L reg represents the regularization function, λ represents the regularization coefficient, q i represents the weight parameter of the i-th category, m represents the dimension corresponding to the weight parameter of the i-th category, k represents the word vector scaling parameter, L total Represents the total loss function.
6. The multi-dimensional category enhanced word vector fusion method applicable to digital humans as claimed in claim 1, characterized in that: The selecting the standard question with the highest cosine similarity and outputting the answer corresponding to the standard question to the user specifically includes: calculating a first cosine similarity between the first extended vector and an extended vector corresponding to the currently selected standard question; Determine whether the currently stored maximum similarity value is less than the first cosine similarity, and if the currently stored maximum similarity value is less than the first cosine similarity, assign the first cosine similarity to the currently stored maximum similarity value; It is determined whether all the extended vectors of the standard question have been traversed. If all the extended vectors have been traversed, the answer to the standard question corresponding to the first cosine similarity is output.
7. The multi-dimensional category enhanced word vector fusion method applicable to digital humans as claimed in claim 1, characterized in that: The expressions of the word vector scaling parameter and the category weight parameter are: Among them, k (t+1) represents the word vector scaling parameter after the t+1th iteration, represents the weight parameter of the i-th category after the t+1th iteration, q i represents the weight parameter of the i-th category, k represents the word vector scaling parameter, and L total represents the total loss function, t represents the number of steps in the current iteration, and η represents the learning rate.
8. A multi-dimensional category enhanced word vector fusion system suitable for digital humans, characterized by: The multi-dimensional category enhanced word vector fusion system (1) comprises a data collection module (11), a sample expansion module (12) and an answer output module (13), wherein: The data collection module (11) is used to construct a question-answer database and collect user input questions, wherein the question-answer database includes a set of standard questions and an answer corresponding to each standard question in the set of standard questions; The sample expansion module (12) is used to construct a question sample pair according to the standard question corresponding to the user input question and the user input question, mark each question sample pair in turn to obtain a positive and negative sample set, perform similarity calculation on the positive and negative sample sets based on a gradient descent algorithm to obtain a word vector scaling parameter and a category weight parameter corresponding to the positive and negative sample sets, and form a first expansion vector corresponding to the user input question according to the word vector scaling parameter, the category weight parameter and the positive and negative sample sets, wherein the similarity includes cosine similarity and dot product similarity; The answer output module (13) is used to calculate the cosine similarity between the first extended vector and the extended vectors corresponding to all standard questions, select the standard question with the highest cosine similarity, and output the answer corresponding to the standard question to the user.
9. An electronic device, characterized in that: The electronic device (2) comprises a processor (21), a memory (25), a user interface (23) and a network interface (24), wherein the memory (25) is used to store instructions, the user interface (23) and the network interface (24) are used to communicate with other devices, and the processor (21) is used to execute the instructions stored in the memory (25) so that the electronic device (2) executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Intelligent question answering method and system based on user feedback correction
CN114997181A
Intelligent text dialogue generation method and device and computer readable storage medium
CN111221942A
Similar question determination method and device in question and answer application and electronic equipment
CN111782762A
Generation method and device of intelligent question and answer model, computing equipment and storage medium
CN114547267A
Intelligent question and answer model construction method and device
CN116450796A