Data processing method, device, electronic device and storage medium
By clustering the semantic vectors of preset problem samples and using central vector comparison to select the most matching problem, the problem of inaccurate responses caused by semantic differences in electronic devices is solved, and more accurate responses and faster result feedback are achieved.
Patent Information
- Application Number
- CN202110062249.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-18
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-01-18
AI Technical Summary
In the prior art, electronic devices have large semantic differences in response to recall problems, resulting in insufficient accuracy of responses.
By clustering the semantic vectors of the preset problem samples, multiple semantic vectors are obtained, and the center vectors of each semantic vector set are calculated. These center vectors are used to compare with the semantic vectors of the target problem, and the first type of problem that is close to the semantics of the target problem is selected, so as to determine the most matching problem and provide a reply.
It improves the accuracy of the response, reduces the amount of calculation, and improves the timeliness of the result feedback.
Smart Images

Figure CN114817483B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language processing, and in particular to a data processing method, device, electronic device, and storage medium. Background Art
[0002] With the development of technology, some electronic devices provide users with channels to obtain search results.
[0003] In the related art, the electronic device recalls questions related to the user question from the database based on the characters contained in the received user question and the characters contained in the questions in the database, and then selects the question closest to the user question from the recalled questions, and uses the answer corresponding to the question as the reply to the user question.
[0004] However, the characters contained in the question can only represent its literal meaning. The questions recalled by the above solution may have a significant difference in semantics from the user's question, resulting in inaccurate responses. Summary of the Invention
[0005] The present disclosure provides a data processing method, device, electronic device and storage medium for solving the problem that the semantics of recalled questions and user questions are quite different, resulting in inaccurate responses.
[0006] In a first aspect, an embodiment of the present disclosure provides a data processing method, the method comprising:
[0007] Obtaining a target question and determining a target semantic vector for the target question;
[0008] Determining the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library, wherein the semantic vector set is obtained by clustering the semantic vectors of a plurality of preset question samples, and the central vector of the semantic vector set represents each semantic vector in the semantic vector set;
[0009] Selecting, from the plurality of preset question samples, a plurality of first-category questions that are semantically close to the target question according to the similarities corresponding to the respective semantic vector sets;
[0010] Based on the multiple first-category questions, a question with the highest matching degree with the target question is determined, and an answer corresponding to the question with the highest matching degree is determined as a reply to the target question.
[0011] The above scheme obtains multiple semantic vector sets by clustering semantic vectors based on multiple preset question samples, so that the semantic vectors in the same semantic vector set have a high similarity. The central vector of the semantic vector set calculated based on the semantic vectors in the semantic vector set represents the overall characteristics of each semantic vector in the semantic vector set. In this way, the target semantic vector representing the semantics of the target question can be compared with the above central vectors respectively. There is no need to compare the target semantic vector with the semantic vectors of each preset question sample one by one. The first type of question with semantics close to the target question can be determined with less calculation amount. The semantic difference between the first type of question and the target question is small. In this way, the question that best matches the user's question is selected from the first type of question, and the answer corresponding to the best matching question is more in line with the expected result, that is, a more accurate reply is obtained.
[0012] In a second aspect, an embodiment of the present disclosure provides a data processing device, including:
[0013] A vector determination module, configured to obtain a target question and determine a target semantic vector for the target question;
[0014] a similarity determination module, configured to respectively determine the similarity between the target semantic vector and a central vector of each semantic vector set in a semantic library, wherein the semantic vector set is obtained by clustering the semantic vectors of a plurality of preset question samples, and the central vector of the semantic vector set represents each semantic vector in the semantic vector set;
[0015] a question selection module, configured to select, from the plurality of preset question samples, a plurality of first-category questions that are semantically close to the target question based on the similarities corresponding to the respective semantic vector sets;
[0016] The reply determination module is used to determine the question with the highest matching degree with the target question based on the multiple first-category questions, and determine the reply corresponding to the question with the highest matching degree as the reply to the target question.
[0017] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising one or more processors and a memory for storing instructions executable by the processors; wherein the processors are configured to execute the instructions to implement the data processing method as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in the first aspect when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A schematic flow chart of a first data processing method provided in an embodiment of the present disclosure;
[0021] Figure 2 A schematic flow chart of a second data processing method provided in an embodiment of the present disclosure;
[0022] Figure 3 A schematic flow chart of a third data processing method provided in an embodiment of the present disclosure;
[0023] Figure 4A A schematic diagram of a target sub-semantic vector provided by an embodiment of the present disclosure;
[0024] Figure 4B Schematic diagram of the sub-semantic vector and target sub-semantic vector of the preset question sample provided in an embodiment of the present disclosure;
[0025] Figure 4C A schematic diagram of the corresponding relationship between the target sub-semantic vector and the central sub-vector provided in an embodiment of the present disclosure;
[0026] Figure 4D A schematic diagram of determining the second target similarity provided by an embodiment of the present disclosure;
[0027] Figure 5 A schematic flow chart of a fourth data processing method provided in an embodiment of the present disclosure;
[0028] Figure 6 A schematic diagram of determining a third target similarity provided in an embodiment of the present disclosure;
[0029] Figure 7 A schematic flow chart of a fifth data processing method provided in an embodiment of the present disclosure;
[0030] Figure 8 A schematic structural diagram of a data processing device provided in an embodiment of the present disclosure;
[0031] Figure 9 A schematic block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] To make the objectives, technical solutions, and advantages of the present disclosure more clear, the present disclosure will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only a portion of the embodiments of the present disclosure, rather than all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without creative effort are intended to fall within the scope of protection of the present disclosure.
[0033] In the embodiments of the present disclosure, the term "and / or" describes the association relationship between associated objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0034] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature identified as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout this disclosure, unless otherwise specified, "plurality" means two or more.
[0035] In the description of this disclosure, it should be noted that, unless otherwise specified or limited, the term "connection" should be understood in a broad sense. For example, it can mean direct connection, indirect connection through an intermediate medium, or internal communication between two devices. Those skilled in the art will understand the specific meaning of the above terms in this disclosure based on the specific circumstances.
[0036] With the development of technology, smart devices (such as robots and mobile terminals) provide users with channels for obtaining search results. In some embodiments, the electronic device matches questions related to the user's question from a database based on the characters contained in the received user's question text. The electronic device then selects the question that is closest to the user's question from the matched questions and uses the answer corresponding to that question as the reply to the user's question.
[0037] For example, the user question received is "How do I know the balance of my bank card account?", and the database contains questions such as "How do I check the balance on my card" and "Withdraw cash from my bank card account balance." Among them, the above user question has fewer characters in common with "How do I check the balance on my card?" and more characters in common with "Withdraw cash from my bank card account balance." When recalling questions related to the user question from the database based on the characters contained in the user question, it is very likely that "Withdraw cash from my bank card account balance" will be recalled, but "How do I check the balance on my card" will not be recalled. However, the semantics of "How do I check the balance on my card" and the user question are basically the same. The above method is difficult to fully recall questions related to the user question, and may even miss questions with similar semantics to the user question. In this way, even if the question closest to the user question is selected from the recalled questions, the semantics of this question may be quite different from the user question, resulting in an inaccurate response that is difficult to meet user expectations.
[0038] In order to solve the problem that the semantic difference between the recalled questions and the user questions is large, resulting in the obtained responses being not accurate enough, the embodiments of the present disclosure provide a data processing method, device, electronic device and storage medium. The technical solution of the present disclosure and how the technical solution of the present disclosure solves the above technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.
[0039] Figure 1 A schematic flow chart of the first data processing method provided in the embodiment of the present disclosure is shown as follows: Figure 1 As shown, the method may include:
[0040] Step 101: Obtain a target question and determine a target semantic vector for the target question.
[0041] Step 102: Determine the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library.
[0042] The semantic vector set is obtained by clustering the semantic vectors of multiple preset question samples, and the central vector of the semantic vector set represents each semantic vector in the semantic vector set.
[0043] For a text, its semantic vector expresses the text's semantic meaning in the form of a vector. Semantic vectors have good semantic properties, and semantically similar texts have high similarity between their semantic vectors. In this embodiment, by comparing semantic vectors, it is possible to identify pre-set question samples that are semantically close to the target question.
[0044] The above semantic library contains multiple preset question samples and also contains text vector indexes for retrieving these preset question samples.
[0045] Generally speaking, the number of preset question samples in the semantic library is huge, and the amount of calculation required to directly determine the similarity between the target semantic vector and the semantic vectors of each preset question sample is large. In this embodiment, multiple semantic vector sets are obtained by clustering the semantic vectors of the question samples, so that the similarity between the semantic vectors in the same semantic vector set is within the set range threshold, and the difference between the semantic vectors in different semantic vector sets is large (the similarity is small); the central vector of the semantic vector set is calculated based on the semantic vectors in the semantic vector set. The central vector of each semantic vector set represents the overall characteristics of the semantic vectors in the semantic vector set, so that during semantic retrieval, only the target semantic vector can be compared with the above-mentioned central vectors, which greatly reduces the amount of calculation.
[0046] This embodiment does not limit the method of clustering to obtain a semantic vector set and calculating a central vector. For example, clustering can be performed by using a K-means clustering algorithm (K-means), a hierarchical agglomerative clustering (HAC), a maximum-minimum distance clustering algorithm, or the like to obtain a semantic vector set, and then a central vector is calculated based on each semantic vector in the semantic vector set. Taking K-means clustering as an example: K initial vectors are randomly selected from all / a certain class of semantic vectors, and the distance from the remaining semantic vectors in all / a certain class of semantic vectors to each initial vector is calculated respectively. The remaining vectors are classified into the class with the closest distance to the initial vector. After classification, K semantic vector sets are obtained, and the average value or weighted average value of all semantic vectors in each semantic vector set is used as the central vector of the semantic vector set.
[0047] The above is merely an example of how to obtain the semantic vector set and center vector in K-means clustering. The specific number of initial vectors and the method used to calculate the center vector of the semantic vector set can be determined based on the actual application scenario. Furthermore, when using other clustering algorithms, the steps may vary, so we will not elaborate on the implementation of each clustering algorithm here.
[0048] Each element in the semantic vector represents a feature with a certain semantic and grammatical interpretation. It is understood that the dimensions of the target semantic vector, the semantic vectors of each preset question sample, and each center vector are all the same, so that the similarity between the vectors can be determined.
[0049] This embodiment does not limit the method for determining the similarity between the target semantic vector and the central vector, such as:
[0050] 1) Calculate the distance between the target semantic vector and the center vector (Euclidean distance, Manhattan distance, Chebyshev distance, Minkowski distance, standardized Euclidean distance, Mahalanobis distance, or Langmuir distance, etc.). The smaller the distance between the vectors, the higher the similarity between the two vectors.
[0051] 2) Calculate the cosine of the angle between the target semantic vector and the center vector. The closer the cosine of the angle between the vectors is to 1, the higher the similarity between the two vectors.
[0052] In addition to using the above two methods to determine the similarity between the target semantic vector and the central vector, this embodiment can also determine the similarity between the target semantic vector and the central vector using other parameters that can represent vector similarity, which is not limited in the embodiments of the present invention.
[0053] In addition, this embodiment does not limit the method of obtaining the above-mentioned target problem, for example:
[0054] 1) The user voice is received by a voice receiving device (such as a microphone), and the user voice is recognized to obtain the user question in text form. The user question is directly used as the target question, and / or the user question is processed to obtain the target question.
[0055] 2) A user question in text form input by the user, directly using the user question as the target question, and / or processing the user question to obtain the target question.
[0056] 3) The user sends an instruction carrying a user question in text form using a terminal device, and directly uses the user question carried in the instruction as the target question, and / or processes the user question to obtain the target question.
[0057] The above examples are just some feasible ways to obtain the target question. This embodiment can also use other methods to obtain the target question, such as receiving an image containing a gesture, determining the target question based on the correspondence between preset text and gestures, etc., and no examples will be given here one by one.
[0058] As mentioned in the above embodiment, the obtained user question can be used as the target question, the question obtained by processing the user question can be used as the target question, or both the user question and the question obtained by processing the user question can be used as the target question.
[0059] In some specific embodiments, the received user question and the expanded question obtained by performing synonymous replacement on at least one content word contained in the user question may be used as the user question, for example:
[0060] The user question is "Where is the toilet?". "Toilet" can be replaced with "restroom," and "where" can be replaced with "which location." This gives us three expanded questions: "Where is the toilet?", "Where is the restroom?", and "Where is the restroom?". This gives us four target questions: "Where is the toilet?", "Where is the restroom?", and "Where is the restroom?"
[0061] If the target question includes a user question and an extended question, the target semantic vector for the target question includes the target semantic vector for each question in the target question. Therefore, it is necessary to determine the similarity between each target semantic vector and the central vector of the semantic vector set. The method for determining the similarity between each target semantic vector and the central vector of the semantic vector set can be referenced to the above embodiment and will not be repeated here.
[0062] This embodiment uses the received user question as the target question and selects questions with semantic similarities to the user question using semantic vectors. It also uses the aforementioned extended question as the target question and selects questions with semantic similarities to the extended question using semantic vectors. These questions are also related to the user question. Compared to related techniques that search only based on user questions, this method can more comprehensively select questions related to the user question.
[0063] Step 103: Select multiple first-category questions with semantics close to the target question from multiple preset question samples based on the similarities corresponding to each semantic vector set.
[0064] Step 104: Based on the multiple first-category questions, determine the question with the highest matching degree with the target question, and determine the answer corresponding to the question with the highest matching degree as the reply to the target question.
[0065] In this embodiment, after selecting multiple first-category questions that are semantically close to the target question, it is also necessary to select one question that best matches the target question based on the multiple first-category questions. The answer corresponding to this question has a high probability of meeting the user's expectations. Based on this, the answer corresponding to the question with the highest matching degree is determined as the answer to the target question.
[0066] This embodiment does not limit the specific method for determining the matching degree. For example, two texts for which the matching degree is to be determined (such as the target question and any of the first-category questions) are input into the matching degree model to obtain the matching degree of the two texts. Of course, this embodiment can also use other methods to determine the matching degree, and these examples will not be given here.
[0067] This embodiment obtains multiple semantic vector sets by clustering semantic vectors based on preset question samples, so that the semantic vectors in the same semantic vector set have a high similarity. The central vector of the semantic vector set calculated based on the semantic vectors in the semantic vector set represents the overall characteristics of each semantic vector in the semantic vector set. In this way, the target semantic vector representing the semantics of the target question can be compared with the above-mentioned central vectors respectively. There is no need to compare the target semantic vector with the semantic vectors of each preset question sample one by one. The first category of questions with semantics close to the target question can be determined with less calculation amount. The semantic difference between the first category of questions and the target question is small. In this way, the question that best matches the user question is selected from the first category of questions, and the reply corresponding to the best matching question is more in line with the expected result, that is, a more accurate reply is obtained.
[0068] In some embodiments, the semantic vector of the preset question sample is a semantic vector that represents the complete semantics of the preset question sample, that is, one preset question corresponds to one semantic vector; correspondingly, the above-mentioned target semantic vector is a semantic vector that represents the complete semantics of the target question. Figure 2 As shown, the embodiment of the present disclosure provides a schematic flow chart of a second data processing method, which may include:
[0069] Step 201: Obtain a target question and determine a target semantic vector for the target question.
[0070] Step 202: Determine the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library.
[0071] The implementation of steps 201-202 may refer to the above-mentioned steps 101-102.
[0072] The target semantic vector in this embodiment is a semantic vector that represents the complete semantics of the target question. The target semantic vector can be obtained by, but not limited to, the following methods: inputting the target question into a target vector model to obtain the target semantic vector of the target question; wherein, the target vector model is trained based on the semantic vectors of two selected sample texts, the difference between the semantic vectors of the two sample texts, and the product of the semantic vectors of the two sample texts.
[0073] Related technologies only use the semantic vectors of two selected sample texts for model training. This embodiment performs model training based on the semantic vectors of two selected sample texts, the difference between the semantic vectors of the two sample texts, and the product of the semantic vectors of the two sample texts. During the model training process, the model can better learn the difference between the two sample texts through the difference between the semantic vectors of the two sample texts; and the model can better learn the similarity between the two sample texts through the product of the semantic vectors of the two sample texts. The similarity between the semantic vectors of texts with similar semantics is enhanced, that is, the semantic vectors obtained by inputting texts with similar semantics into the target vector model will be closer, which is more suitable for the scenario of selecting similar vectors for question recall.
[0074] In some embodiments, the target vector model may be obtained by:
[0075] Two sample texts and their annotation information are used as inputs to a first model, where the annotation information represents the degree of similarity between the two sample texts. The first model includes a vector sub-model and a linear layer, where the inputs to the linear layer include the semantic vectors of the two sample texts output by the vector model, the difference between the semantic vectors of the two sample texts, and the product of the semantic vectors of the two sample texts. The first model is iteratively trained based on the output results and the annotation information of the first model to obtain a second model. The vector sub-model in the second model is determined as a target vector model.
[0076] The above-mentioned vector sub-model can be any model that can obtain the semantic vector of the text, such as the ChineseWord2Vector (Chinese word vector) model, the Global Vectors (GloVe) model or the Bidirectional Encoder Representations from Transformer (BERT) model, etc.
[0077] The specific dimension of the target semantic vector depends on the target vector model used. Different models produce different output dimensions. A higher dimension yields a more comprehensive semantic vector, but also requires more computation. This choice can be tailored to different application scenarios and is not specified here.
[0078] Step 203: According to the similarities corresponding to the semantic vector sets, select N semantic vector sets with the greatest similarity, where N is an integer greater than or equal to 1.
[0079] In this embodiment, the central vector of each semantic vector set represents the semantic vector in the corresponding semantic vector set. Therefore, the similarity between the corresponding semantic vector sets can, to a certain extent, reflect the semantic similarity between the target question and the preset question samples corresponding to the semantic vectors in the set. In other words, if the similarity between the corresponding semantic vector sets is large, the similarity between the target semantic vector and the semantic vectors in the semantic vector set is generally also large, and thus the target question and the preset question corresponding to these semantic vectors are semantically similar.
[0080] Based on this, in this embodiment, it is necessary to select a semantic vector set with greater similarity. Taking 10 semantic vector sets as an example, the semantic vector set can be selected by, but not limited to, the following methods:
[0081] 1) According to the similarities corresponding to the 10 semantic vector sets, a semantic vector set having a similarity greater than a first preset similarity threshold is selected.
[0082] The third preset similarity threshold is denoted as S3. Assuming that the corresponding similarities of the set semantic vectors 2, 3, and 7 are all greater than S3, these three semantic vector sets are selected;
[0083] 2) Sort the similarities corresponding to these 10 semantic vector sets from high to low, and select the N semantic vector sets with the largest similarity.
[0084] In this embodiment, either of the above methods is used to select a semantic vector set with the highest similarity, or both methods are used simultaneously to select a semantic vector set that meets either condition. Furthermore, the number of semantic vector sets and the corresponding similarity levels are provided to better illustrate how to select a semantic vector set and are not intended to limit this embodiment.
[0085] Step 204: Determine first target similarities between the target semantic vector and each semantic vector in the selected semantic vector set.
[0086] Step 205: Based on the first target similarity, select a first type of question from the preset question samples corresponding to the semantic vectors in the selected semantic vector set.
[0087] As mentioned above, if the similarity corresponding to the semantic vector set is large, the similarity between the target semantic vector and each semantic vector in the semantic vector set is generally also large. However, it is not ruled out that due to factors such as poor clustering effect, there are semantic vectors in the semantic vector set that are not very similar to the target semantic vector. There is no need to recall the preset problem samples corresponding to these semantic vectors. If these problems are also recalled, the subsequent calculation amount will increase.
[0088] Based on this, we first need to determine the first target similarity between the semantic vectors included in the selected semantic vector set and the target semantic vector. Since the number of these semantic vectors is significantly smaller than the number of semantic vectors for all pre-set question samples, we can determine the similarity between the target semantic vector and each of these semantic vectors one by one. Only the pre-set question samples corresponding to the semantic vectors with the highest similarity are selected as the first category of questions.
[0089] Step 206: Based on the multiple first-category questions, determine the question with the highest matching degree with the target question, and determine the answer corresponding to the question with the highest matching degree as the reply to the target question.
[0090] The implementation of step 206 is the same as that of step 104, and will not be repeated here.
[0091] The above scheme aims at the scenario where the semantic vector of the preset question sample is the semantic vector that represents the complete semantics of the preset question sample, and the target semantic vector is the semantic vector that represents the complete semantics of the target question. By selecting a set of semantic vectors corresponding to a central vector that has a high similarity with the target semantic vector, and then determining the similarity between the semantic vectors in these selected semantic vector sets and the above target semantic vector, the first type of question that is semantically close to the target question can be selected from the preset question samples corresponding to the semantic vectors in these selected semantic vector sets. In this way, when some semantic vectors in the selected semantic vector set are not very similar to the target semantic vector, there is no need to recall the preset question samples corresponding to these semantic vectors, so as to avoid increasing the subsequent calculation amount.
[0092] When the dimension of the semantic vector representing the complete semantics of the preset problem sample is high, the differences between these semantic vectors are relatively large, and the clustering effect is likely to be poor: the similarity of the vectors in the same semantic vector set may not be very large, and the central vector of each semantic vector set cannot well represent the characteristics of each semantic vector in the semantic vector set. There will be a situation where the similarity corresponding to the semantic vector set is large, but there are many semantic vectors in the semantic vector set that are not close to the semantic vector representing the complete semantics of the target problem, resulting in some semantic vectors in the selected semantic vector set corresponding to the preset problem samples being not close enough to the target problem semantics. Figure 2 In this embodiment, the similarity between the target semantic vector and each of the selected semantic vector sets is determined one by one. Although this method reduces the computational effort compared to determining the target semantic vector and the semantic vectors of all pre-set question samples, in scenarios like intelligent question answering, where timely feedback is crucial, this method still requires a relatively high computational effort, resulting in slower feedback.
[0093] In order to solve this technical problem, this embodiment segments the semantic vector that represents the complete semantics of the preset problem sample to obtain low-dimensional sub-semantic vectors; the above-mentioned semantic vector set is obtained by clustering the sub-semantic vectors at the same segmentation position of multiple preset problem samples, and the clustering effect is good. The sub-semantic vectors in the same semantic vector set are highly similar, and the central vector of each semantic vector set can better represent the characteristics of the corresponding sub-semantic vector. In this way, the target sub-semantic vector has a greater similarity to the central vector of the semantic vector set, and the sub-semantic vectors in the set will also have a greater similarity to the target sub-semantic vector.
[0094] like Figure 3 As shown, a schematic flow chart of a third data processing method is provided for an embodiment of the present disclosure. The method may include:
[0095] Step 301: Obtain a target question, determine a semantic vector representing the complete semantics of the target question, and segment the semantic vector representing the complete semantics of the target question according to a preset segmentation method to obtain multiple target sub-semantic vectors of the target question.
[0096] Step 302: Determine the similarity between the target sub-semantic vector and the central sub-vector of each sub-semantic vector set corresponding to the same segment position in the semantic library.
[0097] Among them, the sub-semantic vector set is obtained by clustering the sub-semantic vectors of the same segmentation position of multiple preset question samples. The central sub-vector of the sub-semantic vector set represents each sub-semantic vector in the sub-semantic vector set. The sub-semantic vector of the preset question sample is obtained by segmenting the semantic vector representing the complete semantics of the preset question sample based on a preset segmentation method.
[0098] As mentioned above, the center vector of each sub-semantic vector set has the same dimension as the sub-semantic vectors in the set. Based on this, the semantic vector representing the complete semantics of the target question needs to be segmented in the same way to obtain multiple target sub-semantic vectors.
[0099] This embodiment does not limit the above-mentioned preset segmentation method. For example, the semantic vector representing the complete semantics of the target question is a 64-dimensional vector. The preset segmentation method is to divide the 64-dimensional vector into 8 segments on average:
[0100] See Figure 4AAs shown in the figure, each 8 dimensions constitute a target sub-semantic vector, and the semantic vector representing the complete semantics of the target question is divided into 8 target sub-semantic vectors. In the order in which they appear in the semantic vector representing the complete semantics of the target question, these 8 target sub-semantic vectors are recorded as: target sub-semantic vector 1 (first segment position), target sub-semantic vector 2 (second segment position), target sub-semantic vector 3 (third segment position), ..., target sub-semantic vector 8 (eighth segment position).
[0101] The dimensions of the semantic vector representing the complete semantics of the target question and the preset segmentation method are only for better illustrating the segmentation process. This embodiment is not limited thereto. For example, the preset segmentation method can segment vectors of any dimension evenly or proportionally.
[0102] As mentioned above, the value of each dimension of the semantic vector represents a feature with a certain semantic and grammatical interpretation. The values of different dimensions in the vector are also incomparable. In other words, whether the target sub-semantic vector at a certain segment position (such as the third segment position) is similar to the central sub-vectors at other segment positions (such as the first segment position, the second segment position, etc.) of the preset question sample is high or low is of practical significance. Therefore, it is necessary to compare the target sub-semantic vector (such as the target sub-semantic vector at the third segment position) with the central vector corresponding to the same segment position (the third segment position).
[0103] Take 100 preset question samples, the semantic vector representing the complete semantics of the preset question samples and the semantic vector representing the complete semantics of the target question are both 64-dimensional vectors. The preset segmentation method is to divide the 64-dimensional vector into 8 segments. For example, see Figure 4B As shown:
[0104] According to the order in the semantic vector representing the complete semantics of the target question, the 8 vectors of the semantic vector representing the complete semantics of the target question are recorded as target sub-semantic vector 1, target sub-semantic vector 2, target sub-semantic vector 3, ..., target sub-semantic vector 8;
[0105] According to the order in the semantic vector representing the complete semantics of the first preset question sample, the 8 vectors of the semantic vector representing the complete semantics of the first preset question sample are recorded as sub-semantic vectors 11 , sub-semantic vector 21 , sub-semantic vector 31 , ..., sub-semantic vector 81 ;
[0106] According to the order in the semantic vector representing the complete semantics of the second preset question sample, the 8 vectors of the semantic vector representing the complete semantics of the second preset question sample are recorded as sub-semantic vectors 12 , sub-semantic vector 22 , sub-semantic vector32 , ..., sub-semantic vector 82 ;
[0107] According to the order in the semantic vector representing the complete semantics of the third preset question sample, the 8 vectors of the semantic vector representing the complete semantics of the third preset question sample are recorded as sub-semantic vectors 13 , sub-semantic vector 23 , sub-semantic vector 33 , ..., sub-semantic vector 83 ;
[0108] …
[0109] According to the order in the semantic vector representing the complete semantics of the 100th preset question sample, the 8 vectors of the semantic vector representing the complete semantics of the 100th preset question sample are recorded as sub-semantic vectors 1100 , sub-semantic vector 2100 , sub-semantic vector 3100 , ..., sub-semantic vector 8100 .
[0110] The sub-semantic vectors of multiple preset question samples at the first segment position have sub-semantic vectors 11 , sub-semantic vector 12 , sub-semantic vector 13 , ..., sub-semantic vector 1100 ;
[0111] The sub-semantic vectors of multiple preset question samples at the second segment position have sub-semantic vectors 21 , sub-semantic vector 22 , sub-semantic vector 23 , ..., sub-semantic vector 2100 ;
[0112] The sub-semantic vectors of multiple preset question samples at the 3rd segment position have sub-semantic vectors 31 , sub-semantic vector 32 , sub-semantic vector 33 , ..., sub-semantic vector 3100 ;
[0113] …
[0114] The sub-semantic vectors of multiple preset question samples at the 8th segment position have sub-semantic vectors 81 , sub-semantic vector 82 , sub-semantic vector 83 , ..., sub-semantic vector 8100 .
[0115] Among them, the 100 sub-semantic vectors of multiple preset question samples at the first segment position are clustered to obtain the sub-semantic vector set11 , sub-semantic vector set 12 and sub-semantic vector set 13 . Sub-semantic vector set 11 Contains sub-semantic vectors 11 , sub-semantic vector 15 , ..., corresponding to the central sub-vector 11 ; Sub-semantic vector set 12 Contains sub-semantic vectors 17 , sub-semantic vector 163 , sub-semantic vector 16 , ..., corresponding to the central sub-vector 12 ; Sub-semantic vector set 13 Contains sub-semantic vectors 18 , sub-semantic vector 154 , ..., corresponding to the central sub-vector 13 The specific implementation of clustering can refer to the above embodiment and will not be described here in detail. Figure 4C As shown, the target sub-semantic vector 1 and the center sub-vector are determined respectively. 11 The similarity between the target sub-semantic vector 1 and the central sub-vector 12 The similarity between the target sub-semantic vector 1 and the central sub-vector 13 similarity.
[0116] Cluster the 100 sub-semantic vectors of multiple preset question samples at the second segment position to obtain a sub-semantic vector set 21 , sub-semantic vector set 22 and sub-semantic vector set 23 . Sub-semantic vector set 21 Contains sub-semantic vectors 279 (Referring to the sub-semantic vectors of the 100 preset question samples in the above embodiment, it can be seen that the "2" in the subscript "279" represents the second segment position, and "79" represents the 79th preset question sample. The subscripts of other sub-semantic vectors have the same meanings). 225 , ..., the central sub-vector is the central sub-vector 21 ; Semantic vector set 22 Contains sub-semantic vectors 26 , sub-semantic vector 253 , sub-semantic vector 214 , ..., the central sub-vector is the central sub-vector 22 ; Sub-semantic vector set 23 Contains sub-semantic vectors 28 , sub-semantic vector 29 , ..., the central sub-vector is the central sub-vector 23 . Determine the target sub-semantic vector 2 and the center sub-vector 21The similarity between the target sub-semantic vector 2 and the central sub-vector 22 The similarity between the target sub-semantic vector 2 and the central sub-vector 23 similarity.
[0117] Similarly, the 100 sub-semantic vectors of multiple preset question samples at other segmented positions are clustered to obtain the sub-semantic vector sets and central sub-vectors corresponding to each segment, and the similarity between the target sub-semantic vector and the central sub-vector of each sub-semantic vector set corresponding to the same segmented position is determined. This will not be repeated here. Figure 4C The 100 sub-semantic vectors of multiple preset question samples in the 3rd segment position are clustered to obtain three sub-semantic vector sets, namely, the sub-semantic vector sets 31 (corresponding to the central sub-vector 31 ), sub-semantic vector set 32 (corresponding to the central sub-vector 32 ) and sub-semantic vector set 33 (corresponding to the central sub-vector 33 ); The 100 sub-semantic vectors at the 8th segment position are clustered to obtain three sub-semantic vector sets, which are the sub-semantic vector sets 81 (corresponding to the central sub-vector 81 ), sub-semantic vector set 82 (corresponding to the central sub-vector 82 ) and sub-semantic vector set 83 (corresponding to the central sub-vector 83 ).
[0118] It is understandable that the number of sub-semantic vector sets corresponding to each segment position is not necessarily the same. Taking K-means clustering as an example, the number of sub-semantic vector sets corresponding to a segment position is related to the number of selected initial vectors.
[0119] Step 303: For any preset question sample, determine a second target similarity corresponding to the preset question sample according to the similarities corresponding to the sub-semantic vectors of the preset question sample.
[0120] The similarity corresponding to each sub-semantic vector of the preset question sample is the similarity corresponding to the sub-semantic vector set to which the sub-semantic vector belongs.
[0121] As described above, after segmentation, we obtain sub-semantic vectors, each of which is a low-dimensional vector. When clustering sub-semantic vectors at the same segment position, the sub-semantic vectors within the same sub-semantic vector set are highly similar. The central sub-vector of each sub-semantic vector set can well characterize the corresponding sub-semantic vector. Therefore, the similarity between the corresponding sub-semantic vector set can be directly used as the similarity between the target sub-semantic vector and the sub-semantic vectors contained in that semantic vector set.
[0122] As described above, a preset question sample corresponds to multiple sub-semantic vectors. The above steps yield the similarity between the target sub-semantic vector and the sub-semantic vector at the same segmented position of the preset question sample. The target question has multiple target sub-semantic vectors, and any preset question sample also has the same number of sub-semantic vectors. Based on the similarity between the target sub-semantic vector and the sub-semantic vector at the same segmented position of the preset question sample, the similarity between the semantic vector representing the complete semantics of the target question, composed of multiple target sub-semantic vectors, and the semantic vector representing the complete semantics of the preset question sample, composed of multiple sub-semantic vectors, can be determined.
[0123] In some embodiments, determining the second target similarity corresponding to the preset question sample based on the similarities corresponding to the sub-semantic vectors of the preset question sample can be achieved by, but not limited to, the following methods:
[0124] Determine the sum of the similarities corresponding to the sub-semantic vectors of the preset question sample as the second target similarity corresponding to the preset question sample; or
[0125] An average value of the similarities corresponding to the sub-semantic vectors of the preset question sample is determined as the second target similarity corresponding to the preset question sample.
[0126] Still taking the above embodiment as an example, refer to Figure 4D As shown:
[0127] According to the order in the semantic vector representing the complete semantics of the target question, the 8 semantic vectors divided into the semantic vector representing the complete semantics of the target question are recorded as target sub-semantic vector 1, target sub-semantic vector 2, target sub-semantic vector 3, ..., target sub-semantic vector 8;
[0128] According to the order in the semantic vector representing the complete semantics of the first preset question sample, the 8 semantic vectors that represent the complete semantics of the first preset question sample are divided into sub-semantic vectors 11 , sub-semantic vector 21 , sub-semantic vector 31 , ..., sub-semantic vector 81 ;
[0129] Determine the target sub-semantic vector 1 and the sub-semantic vector 11 Similarity A , target sub-semantic vector 2 and sub-semantic vector 21 Similarity B , target sub-semantic vector 3 and sub-semantic vector 31 Similarity C , ..., target sub-semantic vector 8 and sub-semantic vector 81 Similarity H, the average value of the above 8 similarities or the sum of the similarities is used as the second target similarity corresponding to the first preset question sample.
[0130] The above embodiment is merely an example of how to determine the second target similarity corresponding to a preset question sample. In this embodiment, any method of determining the second target similarity corresponding to the preset question sample by comprehensively considering the similarities between multiple target sub-semantic vectors and the sub-semantic vectors at the corresponding segmented positions of the preset question sample is feasible. This embodiment does not further elaborate on the implementation method of determining the second target similarity of other preset question samples.
[0131] Step 304: Based on the determined second target similarity of the preset question samples, select a first type of question from a plurality of preset question samples.
[0132] The above steps determine the second target similarity of multiple preset question samples. This second target similarity can accurately reflect the semantic proximity between the target question and the preset question samples. Based on this, this embodiment needs to select the preset question samples with large second target similarity as the first type of questions. The first type of questions can be selected by, but not limited to, the following methods:
[0133] 1) Selecting a preset question sample having a similarity greater than a second preset similarity threshold from a plurality of preset question samples as a first type of question;
[0134] 2) Sort the multiple preset question samples from high to low according to the second target similarity, and select the M preset question samples with the largest similarity as the first category of questions.
[0135] In this embodiment, the preset question samples with a large second target similarity are selected by any of the above methods; or both methods are used at the same time to select the preset question samples that meet any condition.
[0136] Step 305: Based on the multiple first-category questions, determine the question with the highest matching degree with the target question, and determine the answer corresponding to the question with the highest matching degree as the reply to the target question.
[0137] The implementation of step 305 is the same as that of step 104, and will not be repeated here.
[0138] This embodiment segments the semantic vector representing the complete semantics of the preset question sample to obtain low-dimensional sub-semantic vectors, and clusters the sub-semantic vectors at the same segmentation position of multiple preset question samples. The clustering effect is good, and the central sub-vector of each sub-semantic vector set can well represent the corresponding sub-semantic vector. Therefore, the similarity corresponding to the sub-semantic vector set accurately reflects the similarity between the target sub-semantic vector and the sub-semantic vectors contained in the set; the target question has multiple target sub-semantic vectors, and the preset question sample also has the same number of sub-semantic vectors. According to the similarity between the target sub-semantic vector and the sub-semantic vector at the same segmentation position of the preset question sample, a second target similarity can be determined that can accurately reflect the degree of semantic proximity between the target question and the preset question sample. Based on this, selecting preset question samples that are semantically close to the target question can reduce the amount of calculation, improve the timeliness of feedback, and is more suitable for scenarios such as intelligent question answering.
[0139] like Figure 5 As shown, the embodiment of the present disclosure provides a schematic flow chart of a fourth data processing method, which may include:
[0140] Step 501: Obtain a target question, determine a semantic vector representing the complete semantics of the target question, and segment the semantic vector representing the complete semantics of the target question according to a preset segmentation method to obtain multiple target sub-semantic vectors of the target question.
[0141] Step 502: Determine the similarity between the target sub-semantic vector and the central sub-vector of each sub-semantic vector set corresponding to the same segment position in the semantic library.
[0142] The implementation of steps 501-502 is the same as that of steps 301-302, and will not be repeated here.
[0143] Step 503: For any target sub-semantic vector, select the target central sub-vector with the highest similarity to the target sub-semantic vector from the central sub-vectors corresponding to the same segment position.
[0144] In this embodiment, if the preset question sample and the preset segmentation method remain unchanged, the sub-semantic vectors of the preset question sample and the center vectors corresponding to the same segmentation position are fixed, so that the similarity between the sub-semantic vectors of these preset question samples and the center vectors corresponding to the same segmentation position can be determined in advance.
[0145] By selecting the target central subvector with the greatest similarity to the target subsemantic vector from the central subvectors at the same segment position, the similarity between the target central subvector and the subsemantic vector at the corresponding segment position of the preset problem sample can be quickly and efficiently obtained according to the above predetermined similarity. Figure 4B and Figure 4C Take the embodiment shown as an example:
[0146] The first segment position corresponds to the semantic vector set 11 (corresponding to the central sub-vector 11 ), semantic vector set 12 (corresponding to the central sub-vector 12 ) and semantic vector set 13 (corresponding to the central sub-vector 13 ), from the target sub-semantic vector 1 and the center sub-vector 11 The similarity between the target sub-semantic vector 1 and the central sub-vector 12 The similarity between the target sub-semantic vector 1 and the central sub-vector 13 The maximum value is selected among the three similarities, such as the similarity between the target sub-semantic vector 1 and the center sub-vector 11 The similarity is the largest, and the central subvector 11 The target center subvector as the first segment position.
[0147] The target center subvectors of other segment positions are selected in the same way, which will not be repeated here.
[0148] As mentioned above, the center vector of the semantic vector set corresponding to any sub-semantic vector and the same segment position is fixed, so these similarities can be determined in advance, for example:
[0149] Through the use of a graph, the similarity between the central sub-vector of each sub-semantic vector set and the sub-semantic vector of the corresponding segment position is set; by traversing the above graph, the target central sub-vector is found, and the similarity between the target central sub-vector and the sub-semantic vectors of each preset problem sample corresponding to the same segment position is obtained.
[0150] Step 504: For any preset question sample, determine a third target similarity corresponding to the preset question sample based on the similarities between the target center sub-vector and the sub-semantic vectors at the corresponding segmented positions of the preset question sample.
[0151] The similarity between the target center sub-vector and the sub-semantic vector at the same segment position is obtained through the above step 504. There are multiple target center sub-vectors, and there are the same number of sub-semantic vectors for any preset question sample. According to the similarity between each target center sub-vector and the sub-semantic vector at the corresponding segment position of the preset question sample, it can be determined that the semantic closeness between the preset question sample and the target question can be reflected more accurately. Still taking the above embodiment as an example, refer to Figure 6 As shown:
[0152] The target center subvector at the first segment position is the center subvector 11 , recorded as target center subvector 1; the target center subvector at the second segment position is the center subvector23 , recorded as target sub-semantic vector 2; the target center sub-vector at the third segment position is the center sub-vector 32 , recorded as target sub-semantic vector 3; ...; the target central sub-vector at the 8th segment position is the central sub-vector 81 , recorded as the target sub-semantic vector 8; according to the order in the semantic vector representing the complete semantics of the first preset question sample, the 8 vectors are recorded as sub-semantic vectors 11 , sub-semantic vector 21 , sub-semantic vector 31 , ..., sub-semantic vector 81 ;
[0153] According to the above predetermined similarity, the sub-semantic vector can be directly determined 11 With the central subvector 11 The similarity of (target center subvector 1) is recorded as similarity a ; Determine the sub-semantic vector 21 With the central subvector 23 The similarity of (target center subvector 2) is recorded as similarity b ; Determine the sub-semantic vector 31 With the central subvector 32 The similarity of (target center subvector 3) is recorded as similarity c ; ...; Determine the sub-semantic vector 81 With the central subvector 81 The similarity of (target center subvector 8) is recorded as similarity h . The above similarity a - Similarity h The average value of these 8 similarities, or the sum of these 8 similarities, is used as the third target similarity corresponding to the first preset question.
[0154] The above embodiment is merely an example of how to determine the third target similarity corresponding to the first preset question sample. In this embodiment, any method of determining the third target similarity corresponding to the preset question sample by comprehensively considering the similarities between multiple target center subvectors and the sub-semantic vectors at the corresponding segmented positions of the preset question sample is feasible. This embodiment does not further elaborate on the implementation method of determining the third target similarity of other preset question samples.
[0155] Step 505: Based on the determined third target similarity of the preset question samples, select a first type of question from a plurality of preset question samples.
[0156] As mentioned above, the third target similarity can also more accurately reflect the semantic proximity between the target question and the preset question sample.
[0157] above Figures 4A-4Das well as Figure 6 These are all exemplary descriptions, and parameters such as the segmentation method and the number of preset question samples can be set according to the actual application scenario.
[0158] Step 506: Based on the multiple first-category questions, determine the question with the highest matching degree with the target question, and determine the answer corresponding to the question with the highest matching degree as the reply to the target question.
[0159] The implementation of step 506 is the same as that of step 104, and will not be repeated here.
[0160] In this embodiment, the semantic vector representing the complete semantics of the target question is segmented to obtain the target sub-semantic vector, and the semantic vector representing the complete semantics of the preset question sample is segmented to obtain the corresponding sub-semantic vector. If the preset question sample and the preset segmentation method remain unchanged, the sub-semantic vector and the center vector of the preset question sample are fixed. In this way, the similarity between the sub-semantic vectors of these preset question samples and the center vectors corresponding to the same segment position can be determined in advance; by selecting the target center sub-vector with the greatest similarity to the target sub-semantic vector from the center sub-vectors at the same segment position, according to the above-mentioned predetermined similarity, the similarity between the above-mentioned target center sub-vector and the sub-semantic vector at the corresponding segment position of the preset question sample can be quickly and efficiently obtained, and then the third target similarity that can accurately reflect the degree of semantic proximity between the target question and the preset question sample is determined. Based on this, the preset question sample that is semantically close to the target question is selected, which can reduce the amount of calculation, improve the timeliness of feedback, and is more suitable for scenarios such as intelligent question answering.
[0161] The above embodiment can determine the preset question samples that are semantically close to the target question by comparing semantic vectors. The semantic vector similarity of two texts only reflects the semantic closeness of the two texts. Whether the two texts match needs to be judged based on multiple aspects. In order to more comprehensively determine the preset question samples related to the target question, this embodiment provides a fifth data processing method, which can be implemented on the basis of any of the above embodiments. Figure 7 Example Figure 1 The method is described based on the embodiment, which includes:
[0162] Step 701: Obtain a target question and determine a target semantic vector for the target question.
[0163] Step 702: Determine the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library.
[0164] Step 703: Select multiple first-category questions with semantics close to the target question from multiple preset question samples based on the similarities corresponding to each semantic vector set.
[0165] The implementation of steps 701-703 is the same as that of the above steps 101-103, and will not be repeated here.
[0166] Step 704: Perform character retrieval processing based on the target question and multiple preset question samples to determine a second type of question related to the target question.
[0167] In this embodiment, in addition to the semantic library for semantic search, there is also a character library for character search. The semantic library and the character library contain the same preset question samples. The character library also contains character indexes for searching the preset question samples. This embodiment does not limit the specific implementation method for determining the second type of question. For example:
[0168] Pre-process the characters contained in the preset question samples to remove characters without actual meaning, obtain the preset characters (with actual meaning) contained in each preset question sample, and determine the preset question samples corresponding to each preset character (the preset character appears in these preset question samples);
[0169] Preprocess the target question to obtain target characters contained in the target question, find the preset characters that are identical to each target character, and then determine the preset question samples corresponding to these identical preset characters;
[0170] The character relevance between the target question and these predetermined preset question samples can be calculated using methods such as the BM25 algorithm. Taking the BM25 algorithm as an example, the character relevance A1 between the preset question sample 1 and the target question is calculated using the following formula:
[0171] The above qi is the i-th word in the preset question sample 1; f(qi,D) is the number of times qi appears in all the above preset question samples; fieldLen is the length of the preset question sample 1 (the preset number of characters contained in the preset question 1); avgFieldLen is the average length of all preset question samples (the average preset number of characters of all preset question samples); k1 and b are both preset parameters. In some embodiments, k1 is 1.2 and b is 0.75; IDF(qi) is the inverse document frequency of qi in all preset question samples. If qi appears in most preset question samples, the IDF value is low, otherwise it is high. docCount is the total number of preset question samples; f(qi): is the number of preset question samples in which qi appears;
[0172] The preset question samples that are highly relevant to the target question characters are taken as the second type of questions mentioned above.
[0173] There is no necessary logical order between the above steps 701-703 and step 704. Steps 701-703 can be performed first, or step 704 can be performed first, or steps 701-703 and step 704 can be performed simultaneously.
[0174] Step 705: De-duplicate the first and second category questions to obtain candidate questions; determine the question with the highest degree of match to the target question from the candidate questions; and determine the answer corresponding to the question with the highest degree of match as the reply to the target question.
[0175] In this embodiment, among the above-mentioned multiple preset question samples, there may be preset question samples that are semantically close to the target question and have similar characters. These preset question samples may be selected as both the first category and the second category. Therefore, it is necessary to deduplicate the first category and the second category questions to avoid subsequent repeated calculation of the matching degree between the same preset question sample and the target question, which increases the amount of calculation.
[0176] In this embodiment, after deduplication of the first and second category questions, some candidate questions are obtained. It is also necessary to select a question with the highest matching degree with the target question from these candidate questions. By comparing the matching degree of the target question and each candidate question one by one, the higher the matching degree, the closer the candidate question is to the target question. The answer corresponding to the candidate question with the highest matching degree is determined as the reply to the target question.
[0177] This embodiment determines the first type of questions that are semantically close to the target question by means of semantic vectors, and determines the second type of questions that are character-close to the target question by means of characters, thereby obtaining a more comprehensive sample of preset questions related to the target question; by deduplicating these two types of questions, repeated calculation of the matching degree between the preset question samples and the target question is avoided when these two types of questions contain the same preset question samples, thereby reducing the amount of calculation.
[0178] like Figure 8 As shown, based on the same inventive concept, an embodiment of the present disclosure provides a data processing device 800, including:
[0179] A vector determination module 801 is used to obtain a target question and determine a target semantic vector of the target question;
[0180] A similarity determination module 802 is configured to determine the similarity between a target semantic vector and a central vector of each semantic vector set in a semantic library, wherein the semantic vector set is obtained by clustering the semantic vectors of a plurality of preset question samples, and the central vector of the semantic vector set represents each semantic vector in the semantic vector set;
[0181] A question selection module 803 is configured to select a plurality of first-category questions that are semantically close to the target question from a plurality of preset question samples based on the similarities corresponding to the semantic vector sets;
[0182] The reply determination module 804 is configured to determine the question with the highest matching degree with the target question based on the multiple first-category questions, and determine the reply corresponding to the question with the highest matching degree as the reply to the target question.
[0183] In some embodiments, the question selection module 803 selects a plurality of first-category questions that are semantically close to the target question from a plurality of preset question samples based on the similarities corresponding to the semantic vector sets, including:
[0184] According to the similarities corresponding to each semantic vector set, select the N semantic vector sets with the largest similarity, where N is an integer greater than or equal to 1; determine the first target similarity between the target semantic vector and each semantic vector in the selected semantic vector set respectively; based on the first target similarity, select the first type of question from the preset question samples corresponding to the semantic vectors in the selected semantic vector set; wherein the semantic vector of the preset question sample is a semantic vector that represents the complete semantics of the preset question sample, and the target semantic vector is a semantic vector that represents the complete semantics of the target question.
[0185] In some implementations, the vector determination module 801 determines the target semantic vector of the target question including:
[0186] Determine a semantic vector representing the complete semantics of the target question; and segment the semantic vector representing the complete semantics of the target question according to a preset segmentation method to obtain multiple target sub-semantic vectors of the target question;
[0187] The similarity determination module 802 determines the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library, including:
[0188] The similarity between the target sub-semantic vector and the central sub-vector of each sub-semantic vector set corresponding to the same segmentation position in the semantic library is determined respectively; wherein, the sub-semantic vector set is obtained by clustering the sub-semantic vectors of the same segmentation position of multiple preset question samples, the central sub-vector of the sub-semantic vector set represents each sub-semantic vector in the sub-semantic vector set, and the sub-semantic vector of the preset question sample is obtained by segmenting the semantic vector representing the complete semantics of the preset question sample based on a preset segmentation method.
[0189] In some embodiments, the question selection module 803 selects a plurality of first-category questions that are semantically close to the target question from a plurality of preset question samples based on the similarities corresponding to the semantic vector sets, including:
[0190] For any preset question sample, the second target similarity corresponding to the preset question sample is determined based on the similarities corresponding to each sub-semantic vector of the preset question sample, wherein the similarity corresponding to each sub-semantic vector of the preset question sample is the similarity corresponding to the sub-semantic vector set in which the sub-semantic vector is located; based on the determined second target similarity of the preset question sample, the first type of question is selected from multiple preset question samples.
[0191] In some embodiments, the question selection module 803 determines a second target similarity corresponding to the preset question sample based on the similarities corresponding to the sub-semantic vectors of the preset question sample, including:
[0192] Determine the sum of the similarities corresponding to the sub-semantic vectors of the preset question sample as the second target similarity corresponding to the preset question sample; or
[0193] An average value of the similarities corresponding to the sub-semantic vectors of the preset question sample is determined as the second target similarity corresponding to the preset question sample.
[0194] In some embodiments, the question selection module 803 selects a plurality of first-category questions that are semantically close to the target question from a plurality of preset question samples based on the similarities corresponding to the semantic vector sets, including:
[0195] For any target sub-semantic vector, select the target central sub-vector with the highest similarity to the target sub-semantic vector from the central sub-vectors corresponding to the same segment position;
[0196] For any preset question sample, determine the third target similarity corresponding to the preset question sample according to the similarity between the target center subvector and the sub-semantic vector of the corresponding segment position of the preset question sample;
[0197] Based on the determined third target similarity of the preset question samples, a first type of question is selected from a plurality of preset question samples.
[0198] In some embodiments, the question selection module 803 determines a third target similarity corresponding to the preset question sample based on the similarity between the target center subvector and the sub-semantic vectors at the corresponding segmented positions of the preset question sample, including:
[0199] Determine the sum of the similarities between the target center sub-vector and the sub-semantic vectors at the corresponding segmented positions of the preset question sample as the third target similarity corresponding to the preset question sample; or
[0200] The average value of the similarities between the target center sub-vector and the sub-semantic vectors at the corresponding segmented positions of the preset question sample is determined as the third target similarity corresponding to the preset question sample.
[0201] In some implementations, the target question includes: a received user question, and an expanded question obtained by performing synonym replacement on at least one content word included in the user question.
[0202] In some embodiments, the vector determination module 801 determines the target semantic vector of the target question, including: inputting the target question into a target vector model to obtain a target semantic vector; wherein the target vector model is trained based on the semantic vectors of two selected sample texts, the difference between the semantic vectors of the two sample texts, and the product of the semantic vectors of the two sample texts.
[0203] In some embodiments, the target vector model is obtained by:
[0204] Two sample texts and their annotation information are used as inputs to a first model, where the annotation information represents the degree of similarity between the two sample texts. The first model includes a vector sub-model and a linear layer, where the inputs to the linear layer include the semantic vectors of the two sample texts output by the vector model, the difference between the semantic vectors of the two sample texts, and the product of the semantic vectors of the two sample texts. Based on the output results and the annotation information of the first model, the first model is iteratively trained to obtain a second model. The vector sub-model in the second model is determined as the target vector model.
[0205] In some embodiments, the question selection module is further configured to: perform character retrieval processing based on the target question and a plurality of preset question samples to determine a second type of question related to the target question;
[0206] The reply determination module is specifically used to: remove duplicates from the first and second category questions to obtain candidate questions; and determine the question with the highest matching degree with the target question from the candidate questions.
[0207] Since the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0208] like Figure 9 As shown, based on the same inventive concept, an embodiment of the present disclosure provides an electronic device 900. In some specific embodiments, the electronic device 900 includes: a processor 901 and a memory 902;
[0209] Memory 902 is used to store computer programs executed by processor 901. Memory 902 can be a volatile memory (volatile memory), such as random-access memory (RAM); memory 902 can also be a non-volatile memory (non-volatile memory), such as read-only memory, flash memory (flash memory), hard disk drive (HDD) or solid-state drive (SSD), or memory 902 is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. Memory 902 can be a combination of the above memories. Processor 901 may include one or more central processing units (CPU), graphics processing units (GPU) or digital processing units, etc.
[0210] The specific connection medium between the memory 902 and the processor 901 is not limited in the embodiment of the present disclosure. Figure 9 The memory 902 and the processor 901 are connected via a bus 903. Figure 9 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus 903 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 9 The bus is represented by only one thick line, but it does not mean that there is only one bus or one type of bus. The memory stores program code. When the program code is executed by the processor, the processor 901 performs the following process:
[0211] Obtain a target question and determine a target semantic vector for the target question; determine the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library, respectively, where the semantic vector set is obtained by clustering the semantic vectors of multiple preset question samples, and the central vector of the semantic vector set represents each semantic vector in the semantic vector set; select multiple first-category questions with semantics close to the target question from multiple preset question samples based on the similarity corresponding to each semantic vector set; based on the multiple first-category questions, determine the question with the highest match to the target question, and determine the answer corresponding to the question with the highest match as the reply to the target question.
[0212] The above-mentioned electronic device is a device with certain computing capabilities, and can be one or more groups of servers, or a smart device that directly interacts with a user, etc. This embodiment does not make specific limitations on this.
[0213] Since the electronic device is the electronic device that executes the method in the embodiment of the present disclosure, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0214] The present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned data processing method.
[0215] The present disclosure is described above with reference to block diagrams and / or flow charts illustrating methods, apparatus (systems) and / or computer program products according to embodiments of the present disclosure. It should be understood that a block of a block diagram and / or flow chart, as well as a combination of blocks of a block diagram and / or flow chart, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer and / or other programmable data processing device to produce a machine such that the instructions executed by the computer processor and / or other programmable data processing device create a method for implementing the functions / actions specified in the block diagram and / or flow chart block.
[0216] Accordingly, the present disclosure may also be implemented in hardware and / or software (including firmware, resident software, microcode, etc.). Furthermore, the present disclosure may take the form of a computer program product on a computer-usable or computer-readable storage medium having a computer-usable or computer-readable program code implemented in the medium for use by or in conjunction with an instruction execution system. In the context of the present disclosure, a computer-usable or computer-readable medium may be any medium that can contain, store, communicate, transmit, or convey a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0217] Although the preferred embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present disclosure.
[0218] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A data processing method, characterized in that: The method comprises: Obtaining a target question and determining a target semantic vector for the target question; Determining the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library, wherein the semantic vector set is obtained by clustering the semantic vectors of a plurality of preset question samples, and the central vector of the semantic vector set represents each semantic vector in the semantic vector set; Selecting, from the plurality of preset question samples, a plurality of first-category questions that are semantically close to the target question according to the similarities corresponding to the respective semantic vector sets; Based on the multiple first-category questions, determining the question with the highest matching degree with the target question, and determining the answer corresponding to the question with the highest matching degree as the reply to the target question; Determining a target semantic vector for the target question includes: Determine a semantic vector representing the complete semantics of the target question; and segment the semantic vector representing the complete semantics of the target question according to a preset segmentation method to obtain multiple target sub-semantic vectors of the target question; Determining the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library respectively includes: Determine respectively the similarity between the target sub-semantic vector and the central sub-vector of each sub-semantic vector set corresponding to the same segmentation position in the semantic library; wherein, the sub-semantic vector set is obtained by clustering the sub-semantic vectors at the same segmentation position of the multiple preset question samples, the central sub-vector of the sub-semantic vector set represents each sub-semantic vector in the sub-semantic vector set, and the sub-semantic vector of the preset question sample is obtained by segmenting the semantic vector representing the complete semantics of the preset question sample based on the preset segmentation method; Selecting a plurality of first-category questions that are semantically close to the target question from the plurality of preset question samples according to the similarities corresponding to the semantic vector sets, including: For any preset question sample, determine the second target similarity corresponding to the preset question sample based on the similarity corresponding to each sub-semantic vector of the preset question sample, wherein the similarity corresponding to each sub-semantic vector of the preset question sample is the similarity corresponding to the sub-semantic vector set to which the sub-semantic vector belongs; based on the determined second target similarity of the preset question sample, select the first type of question from the multiple preset question samples; or, For any target sub-semantic vector, select the target central sub-vector with the highest similarity to the target sub-semantic vector from the central sub-vectors corresponding to the same segmented position; for any preset question sample, determine the third target similarity corresponding to the preset question sample based on the similarity between the target central sub-vector and the sub-semantic vectors at the corresponding segmented position of the preset question sample; based on the determined third target similarity of the preset question sample, select the first type of question from the multiple preset question samples.
2. The method according to claim 1, characterized in that Selecting a plurality of first-category questions that are semantically close to the target question from the plurality of preset question samples according to the similarities corresponding to the semantic vector sets, including: According to the similarities corresponding to the semantic vector sets, select N semantic vector sets with the greatest similarity, where N is an integer greater than or equal to 1; respectively determining a first target similarity between the target semantic vector and each semantic vector in the selected semantic vector set; Based on the first target similarity, the first type of question is selected from the preset question samples corresponding to the semantic vectors in the selected semantic vector set; wherein the semantic vector of the preset question sample is a semantic vector that represents the complete semantics of the preset question sample, and the target semantic vector is a semantic vector that represents the complete semantics of the target question.
3. The method according to claim 1, characterized in that Determining a second target similarity corresponding to the preset question sample according to the similarities corresponding to the sub-semantic vectors of the preset question sample includes: Determining the sum of the similarities corresponding to the sub-semantic vectors of the preset question sample as the second target similarity corresponding to the preset question sample; or An average value of the similarities corresponding to the sub-semantic vectors of the preset question sample is determined as the second target similarity corresponding to the preset question sample.
4. The method according to claim 1, wherein Determining a third target similarity corresponding to the preset question sample according to similarities between the target center subvector and the sub-semantic vectors at corresponding segmented positions of the preset question sample, including: Determine the sum of the similarities between the target center sub-vector and the sub-semantic vectors at the corresponding segmented positions of the preset question sample as the third target similarity corresponding to the preset question sample; or An average value of the similarities between the target center sub-vector and the sub-semantic vectors at the corresponding segmented positions of the preset question sample is determined as a third target similarity corresponding to the preset question sample.
5. The method according to claim 1, wherein The target question includes: a received user question, and an expanded question obtained by performing synonymous replacement on at least one content word included in the user question.
6. The method according to claim 1, characterized in that Determining a target semantic vector for the target question includes: Inputting the target question into a target vector model to obtain the target semantic vector; The target vector model is obtained by training based on the semantic vectors of two selected sample texts, the difference between the semantic vectors of the two sample texts, and the product of the semantic vectors of the two sample texts.
7. The method according to claim 6, characterized in that The target vector model is obtained by: The two sample texts and their annotation information are used as inputs of a first model, wherein the annotation information represents the degree of similarity between the two sample texts. The first model includes a vector model and a linear layer, wherein the inputs of the linear layer include the semantic vectors of the two sample texts output by the vector model, the difference between the semantic vectors of the two sample texts, and the product of the semantic vectors of the two sample texts; Iteratively training the first model according to the output result of the first model and the labeled information to obtain a second model; The vector sub-model in the second model is determined as the target vector model.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Performing character retrieval processing on the target question and the plurality of preset question samples to determine a second type of question related to the target question; Determining, based on the multiple first-category questions, a question that has the highest degree of matching with the target question, including: The first and second types of questions are deduplicated to obtain candidate questions; and the question with the highest matching degree with the target question is determined from the candidate questions.
9. A data processing device, characterized in that: The device includes: A vector determination module, configured to obtain a target question and determine a target semantic vector for the target question; a similarity determination module, configured to respectively determine the similarity between the target semantic vector and a central vector of each semantic vector set in a semantic library, wherein the semantic vector set is obtained by clustering the semantic vectors of a plurality of preset question samples, and the central vector of the semantic vector set represents each semantic vector in the semantic vector set; a question selection module, configured to select, from the plurality of preset question samples, a plurality of first-category questions that are semantically close to the target question based on the similarities corresponding to the respective semantic vector sets; a reply determination module, configured to determine, based on the plurality of first-category questions, a question that has the highest degree of matching with the target question, and determine an answer corresponding to the question with the highest degree of matching as a reply to the target question; The vector determination module determines the target semantic vector of the target problem including: Determine a semantic vector representing the complete semantics of the target question; and segment the semantic vector representing the complete semantics of the target question according to a preset segmentation method to obtain multiple target sub-semantic vectors of the target question; The similarity determination module determines the similarity between the target semantic vector and the central vector of each semantic vector set in the semantic library, including: Determine the similarity between the target sub-semantic vector and the central sub-vector of each sub-semantic vector set corresponding to the same segment position in the semantic library; wherein the sub-semantic vector set is obtained by clustering the sub-semantic vectors at the same segment position of multiple preset question samples, the central sub-vector of the sub-semantic vector set represents each sub-semantic vector in the sub-semantic vector set, and the sub-semantic vector of the preset question sample is obtained by segmenting the semantic vector representing the complete semantics of the preset question sample based on a preset segmentation method; The question selection module selects multiple first-category questions with semantic similarity to the target question from multiple preset question samples based on the similarity corresponding to each semantic vector set, including: For any preset question sample, determine a second target similarity corresponding to the preset question sample based on the similarity corresponding to each sub-semantic vector of the preset question sample, wherein the similarity corresponding to each sub-semantic vector of the preset question sample is the similarity corresponding to the sub-semantic vector set to which the sub-semantic vector belongs; based on the determined second target similarity of the preset question sample, select a first type of question from multiple preset question samples; or, For any target sub-semantic vector, select the target central sub-vector with the highest similarity to the target sub-semantic vector from the central sub-vectors corresponding to the same segmented position; for any preset question sample, determine the third target similarity corresponding to the preset question sample based on the similarity between the target central sub-vector and the sub-semantic vectors at the corresponding segmented positions of the preset question sample; based on the determined third target similarity of the preset question sample, select the first type of question from multiple preset question samples.
10. An electronic device, characterized in that: comprising one or more processors, and a memory for storing instructions executable by said processors; The processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 8.
11. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the data processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Semantic retrieval method, device and equipment and storage medium
CN111753069A