Semantic recognition method based on professional vocabularies

Through the semantic recognition method based on professional vocabulary and combined with the comprehensive calculation of semantic distance and professional distance, the problem of difficulty in matching user input problems in professional fields is solved, and more accurate matching effect of professional fields is achieved, and the calculation complexity is reduced.

CN120068878AActive Publication Date: 2025-05-30HUBEI TAIYUE SATELLITE TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510137876.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

In the professional field, due to the existence of a large number of professional vocabulary, the prior art is difficult to accurately identify and match user input, which makes it difficult to get the best answers in the professional field.

Method used

A semantic recognition method based on professional vocabulary is provided. By obtaining the collection of current problems and target professional vocabulary, the data set corresponding to the current problems, including the overall semantic word vector and professional vocabulary. Then, the semantic distance and professional distance from the candidate criteria question are comprehensively calculated, and the candidate criteria question with the smallest comprehensive distance is selected as the answer.

Benefits of technology

Through the comprehensive calculation of semantic distance and professional distance, the semantic differences of professional vocabulary in specific fields can be more accurately reflected, the matching effect of professional field Q&A, and significantly reduce the complexity and resource consumption of real-time calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068878A_ABST
    Figure CN120068878A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic recognition method based on professional vocabularies, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining a current question and a target professional vocabulary set, and determining a first data set corresponding to the current question according to the target professional vocabulary set; obtaining a second data set corresponding to each candidate standard question; for each candidate standard question, determining a semantic distance corresponding to the candidate standard question according to the second overall semantic word vector and the first overall semantic word vector, and determining a professional distance corresponding to the candidate standard question according to the first professional vocabulary set and the second professional vocabulary set; according to the semantic distance and the professional distance corresponding to each candidate standard question, the comprehensive distance between each candidate standard question and the current question is determined, and the standard answer corresponding to the candidate standard question with the minimum comprehensive distance serves as the answer of the current question. The semantic difference of the professional vocabularies in the specific field can be reflected more accurately, and the matching effect of questions and answers in the professional field is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and particularly to a semantic recognition method based on professional vocabulary. Background Art

[0002] In a digital human system, a series of standard questions and their corresponding answers are usually built-in. When a user inputs a question, the BERT model is usually used to convert the user's question into word vectors, and then compare them with the word vectors of the built-in standard questions, and select the answer of the standard question closest to the current question for reply.

[0003] By using the above method, the best answer can be obtained in the general field. However, in the professional field, due to the existence of a large number of professional vocabulary, when the question contains these professional vocabulary, the word vectors of these vocabulary in the general corpus may be similar, but their meanings in the professional field are quite different. Taking the agricultural field as an example, "crop rotation" and "intercropping", "irrigation" and "drainage" are different concepts. Among them, "irrigation" refers to supplying water to farmland to meet the growth needs of crops, while "drainage" is to drain excessive water in the farmland to prevent crops from being waterlogged. These two words represent opposite operations in agriculture, but in the general corpus, because they are both related to "water", the model may think that their word vectors are similar, resulting in it being difficult to obtain the best answer in the professional field. Therefore, there is an urgent need for a semantic recognition method for professional vocabulary at present. Summary of the Invention

[0004] The purpose of the present application aims to solve at least one of the above technical defects.

[0005] On the one hand, the embodiments of the present application provide a semantic recognition method based on professional vocabulary, and the method includes:

[0006] Obtain the current question and the target professional vocabulary set, and determine the first data set corresponding to the current question according to the target professional vocabulary set. The first data set includes the first overall semantic word vector corresponding to the current question and the first professional vocabulary set. The first professional vocabulary set includes the first professional vocabulary included in the current question and the first word vector corresponding to the first professional vocabulary;

[0007] Obtain the second data set corresponding to each candidate standard question. Each second data set includes the second overall semantic word vector corresponding to the corresponding candidate standard question and the second professional vocabulary set. The second professional vocabulary set includes the second professional vocabulary included in the corresponding candidate standard question and the second word vector corresponding to the second professional vocabulary;

[0008] For each candidate standard question, determine the semantic distance corresponding to the candidate standard question according to the second overall semantic word vector and the first overall semantic word vector corresponding to the candidate standard question, and determine the professional distance corresponding to the candidate standard question according to the first set of professional vocabulary corresponding to the current question and the second set of professional vocabulary corresponding to the candidate standard question;

[0009] According to the semantic distance and professional distance corresponding to each candidate standard question, determine the comprehensive distance between each candidate standard question and the current question, and use the standard answer corresponding to the candidate standard question with the smallest comprehensive distance as the answer to the current question.

[0010] Optionally, the target set of professional vocabulary is the set of professional vocabulary corresponding to the field to which the current question belongs. Determine the first data set corresponding to the current question according to the target set of professional vocabulary, including:

[0011] Input the current question into the BERT model to obtain the first overall semantic word vector corresponding to the current question;

[0012] Extract professional vocabulary from the current question based on the target set of professional vocabulary to obtain the first professional vocabulary included in the current question;

[0013] Input the first professional vocabulary into the BERT model to obtain the first word vector corresponding to the first professional vocabulary.

[0014] Optionally, according to the semantic distance and professional distance corresponding to each candidate standard question, determine the comprehensive distance between each candidate standard question and the current question, including:

[0015] Determine the number of words in the first professional vocabulary included in the current question, and determine the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words;

[0016] For each candidate standard question, determine the comprehensive distance between the candidate standard question and the current question according to the first weight coefficient, the second weight coefficient, and the semantic distance and professional distance corresponding to the candidate standard question.

[0017] Optionally, determine the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words, including:

[0018] Determine whether the number of words meets the preset conditions;

[0019] If the number of words does not meet the preset conditions, set the first weight coefficient corresponding to the semantic distance to 1 and set the second weight coefficient corresponding to the professional distance to 0;

[0020] If the number of words meets the preset condition, then according to the number of words, determine the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance.

[0021] Optionally, the sum of the first weight and the second weight is 1. Determining the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words includes:

[0022] Determine the second weight coefficient corresponding to the professional distance according to the number of words;

[0023] Determine the first weight coefficient corresponding to the semantic distance according to the second weight coefficient;

[0024] Among them, the second weight coefficient is determined by the following method:

[0025]

[0026] Among them, β is the second weight coefficient, is the number of words, a is an adjustment parameter, and k is a non-linear adjustment parameter.

[0027] Optionally, for each candidate standard question, determine the professional distance corresponding to the candidate standard question according to the first professional vocabulary set corresponding to the current question and the second professional vocabulary set corresponding to the candidate standard question, including:

[0028] Determine the vocabulary intersection and vocabulary union according to the first professional vocabulary included in the current question and the second professional vocabulary included in the candidate standard question;

[0029] Determine the first quantity of professional vocabulary in the vocabulary intersection, the second quantity of professional vocabulary in the vocabulary union, and the ratio of the first quantity to the second quantity;

[0030] Determine the professional distance corresponding to the candidate standard question according to the ratio, the second quantity, the first professional vocabulary set, and the second professional vocabulary set.

[0031] Optionally, determine the professional distance corresponding to the candidate standard question according to the ratio, the first professional vocabulary set, and the second professional vocabulary set, including:

[0032] Determine the semantic distance corresponding to the vocabulary intersection and the semantic distance corresponding to the vocabulary union respectively according to the first professional vocabulary set and the second professional vocabulary set;

[0033] Determine the professional distance corresponding to the candidate standard question according to the ratio and the preset threshold relationship, the second quantity, the semantic distance corresponding to the vocabulary intersection, and the semantic distance corresponding to the vocabulary union.

[0034] Optionally, determine the professional distance corresponding to the candidate standard question according to the ratio and the preset threshold relationship, the second quantity, the semantic distance corresponding to the lexical intersection, and the semantic distance corresponding to the lexical union, including:

[0035] If the ratio meets the first threshold, obtain the first amplification factor, and determine the professional distance corresponding to the candidate standard question according to the first amplification factor, the second quantity, and the semantic distance corresponding to the lexical union. The first threshold indicates that the lexical intersection is the same as the lexical union;

[0036] If the ratio meets the second threshold, obtain the second amplification factor, and determine the professional distance corresponding to the candidate standard question according to the second amplification factor, the second quantity, and the semantic distance corresponding to the lexical intersection. The second threshold indicates that the lexical intersection is an empty set, the first threshold is less than the second threshold, and the first amplification factor is greater than the second amplification factor;

[0037] If the ratio does not meet the first threshold and the second threshold, then determine the third amplification factor according to the ratio, and determine the professional distance corresponding to the candidate standard question according to the third amplification factor, the second quantity, the semantic distance corresponding to the lexical union, and the semantic distance corresponding to the lexical intersection.

[0038] Optionally, determine the semantic distance corresponding to the lexical intersection and the semantic distance corresponding to the lexical union according to the first professional vocabulary set and the second professional vocabulary set, respectively, including:

[0039] Determine the word vectors corresponding to each professional vocabulary in the lexical intersection. According to the word vectors corresponding to each professional vocabulary in the lexical intersection, determine the semantic distance corresponding to the lexical intersection. The word vectors corresponding to each professional vocabulary in the lexical intersection include at least one of the first word vector corresponding to the professional vocabulary in the first professional vocabulary set and the second word vector corresponding to the professional vocabulary in the second professional vocabulary set;

[0040] Determine the target professional vocabulary in the lexical union that does not belong to the lexical intersection. According to the word vector corresponding to the target professional vocabulary, determine the semantic distance corresponding to the lexical union. The word vector corresponding to the target professional vocabulary is the first word vector corresponding to the target professional vocabulary in the first professional vocabulary set or the second word vector corresponding to the target professional vocabulary in the second professional vocabulary set.

[0041] Optionally, the comprehensive distance between each candidate standard question and the current question is determined by the following formula:

[0042]

[0043] where D(Q i ) represents the comprehensive distance, α is the first weight coefficient, β is the second weight coefficient, represents the second overall semantic word vector Q of the candidate standard question i of the component in the j-th dimension, Represents the second overall semantic word vector Q of the current problem c The component in the j-th dimension, U represents the dimension of the vector, p is the norm coefficient, A is the vocabulary intersection, B is the vocabulary union, γ is the amplification coefficient, and N B Is the second quantity, w j Is the first word vector or the corresponding second word vector corresponding to the professional vocabulary w And Are respectively the j-th components of the professional vocabulary w in the corresponding first word vector and second word vector

[0044] On the other hand, the embodiment of the present application provides a semantic recognition device based on professional vocabulary, and the device includes:

[0045] A current problem acquisition module, configured to acquire the current problem and the target professional vocabulary set, and determine the first data set corresponding to the current problem according to the target professional vocabulary set. The first data set includes the first overall semantic word vector corresponding to the current problem and the first professional vocabulary set. The first professional vocabulary set includes the first professional vocabulary included in the current problem and the first word vector corresponding to the first professional vocabulary

[0046] A candidate problem acquisition module, configured to acquire the second data set corresponding to each candidate standard problem. Each second data set includes the second overall semantic word vector corresponding to the corresponding candidate standard problem and the second professional vocabulary set. The second professional vocabulary set includes the second professional vocabulary included in the corresponding candidate standard problem and the second word vector corresponding to the second professional vocabulary

[0047] A distance determination module, configured to, for each candidate standard problem, determine the semantic distance corresponding to the candidate standard problem according to the second overall semantic word vector corresponding to the candidate standard problem and the first overall semantic word vector, and determine the professional distance corresponding to the candidate standard problem according to the first professional vocabulary set corresponding to the current problem and the second professional vocabulary set corresponding to the candidate standard problem

[0048] A result determination module, configured to determine the comprehensive distance between each candidate standard problem and the current problem according to the semantic distance and the professional distance corresponding to each candidate standard problem, and use the standard answer corresponding to the candidate standard problem with the smallest comprehensive distance as the answer to the current problem

[0049] On yet another aspect, the embodiment of the present application provides an electronic device, including a processor and a memory:

[0050] The memory is configured to store machine-readable instructions. When the instructions are executed by the processor, the processor executes any one of the methods in the semantic recognition based on professional vocabulary

[0051] The beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0052] In this application, for the current problem input by the user, first determine the professional vocabulary included in the current problem and the corresponding word vectors, then calculate the comprehensive semantic distance based on the professional vocabulary and the corresponding word vectors included in the current problem and each candidate standard problem. Finally, select the problem with the smallest comprehensive semantic distance from all candidate standard problems, and output the corresponding answer as the final result. At this time, the obtained comprehensive semantic distance includes the overall semantic distance and the professional vocabulary distance between the current problem and the candidate standard problem, and the professional vocabulary distance can fully reflect the word vector characteristics of the professional vocabulary, representing the degree of fit between the current problem and the candidate standard problem in the professional field. Therefore, the final result determined based on the comprehensive semantic distance can more accurately reflect the semantic differences of professional vocabulary in a specific field and improve the matching effect of question and answer in the professional field.

[0053] In addition, in this application, by constructing a professional vocabulary list and using a pre-trained language model, hierarchical semantic vector extraction and storage are performed on the candidate standard problems and their professional vocabulary, significantly reducing the complexity of real-time calculation and resource consumption, and laying a foundation for subsequent efficient and accurate semantic matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0055] Figure 1 It is a flowchart of the semantic recognition method based on professional vocabulary provided by the embodiment of this application;

[0056] Figure 2 It is the overall flowchart of the semantic recognition method based on professional vocabulary provided by the embodiment of this application;

[0057] Figure 3 It is a flowchart of the preprocessing steps for candidate standard problems provided by the embodiment of this application;

[0058] Figure 4 It is a flowchart of the processing steps for the current problem provided by the embodiment of this application;

[0059] Figure 5 It is a flowchart of the step of outputting the best answer provided by the embodiment of this application;

[0060] Figure 6 It is a schematic structural diagram of the semantic recognition device based on professional vocabulary provided by the embodiment of this application;

[0061] Figure 7 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0062] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present invention.

[0063] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means that there are the described features, integers, steps, operations, elements and / or components, but does not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0064] To make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0065] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0066] Specifically, as Figure 1 shown, the method may include:

[0067] Step S101, obtaining the current problem and the target professional vocabulary set, and determining the first data set corresponding to the current problem according to the target professional vocabulary set. The first data set includes the first overall semantic word vector corresponding to the current problem and the first professional vocabulary set. The first professional vocabulary set includes the first professional vocabulary included in the current problem and the first word vector corresponding to the first professional vocabulary.

[0068] The current question refers to the question that needs to be matched. The user can input the current question through the digital human. The digital human can input the current question in voice or text. If it is voice input, the voice content needs to be converted into text with the help of a voice-to-text tool to obtain the current question.

[0069] A professional vocabulary set refers to a set of professional vocabulary in a certain field, and a target professional vocabulary set is a professional vocabulary set corresponding to the field to which the current question belongs. For example, if the current question is "What should be paid attention to when irrigating rice?", and the current question belongs to the agricultural field, then the corresponding target professional vocabulary set is the professional vocabulary set corresponding to the agricultural field.

[0070] In an optional embodiment of the present application, the professional vocabulary set corresponding to each field is determined in the following manner:

[0071] A document collection corresponding to the field is obtained, and professional vocabulary is extracted from the document collection to form a professional vocabulary collection corresponding to the field.

[0072] Optionally, when determining the professional vocabulary set corresponding to a certain field, a manual method can be used to analyze the relevant literature, data and standard issues in the field, and then extract common professional terms to build a professional vocabulary table (i.e. professional vocabulary set) covering the field. This vocabulary table will serve as the basis for identifying professional vocabulary in subsequent steps to ensure accurate extraction and processing of key terms in candidate standard issues.

[0073] Furthermore, after obtaining the current question and the target professional vocabulary set, the first overall semantic word vector corresponding to the current question, the first professional vocabulary included in the current question and the first word vector corresponding to the first professional vocabulary can be determined according to the target professional vocabulary set.

[0074] In an optional embodiment of the present application, the target professional vocabulary set is a professional vocabulary set corresponding to the field to which the current question belongs, and determining the first data set corresponding to the current question according to the target professional vocabulary set includes:

[0075] Input the current question into the BERT model to obtain the first overall semantic word vector corresponding to the current question;

[0076] Extract professional vocabulary from the current question based on the target professional vocabulary set to obtain the first professional vocabulary included in the current question;

[0077] The first professional vocabulary is input into the BERT model to obtain the first word vector corresponding to the first professional vocabulary.

[0078] Optionally, the received current question can be input into a pre-trained BERT model to generate an overall semantic word vector of the current question input by the user (i.e., the first overall semantic word vector). Then, the current question is matched through a target professional vocabulary set to extract the professional vocabulary therein, obtaining the professional vocabulary included in the current question (i.e., the first professional vocabulary). For the professional vocabulary included in the current question, the BERT model is used to generate semantic word vectors of these first professional vocabulary (i.e., the first word vectors). Further, the first overall semantic word vector of the current question, the included professional vocabulary, and their corresponding first word vectors are integrated to generate a first data set corresponding to the current question (i.e., the first data set), providing support for subsequent calculations and processing. That is to say, in this application, through the processing and semantic analysis of the user input question, the transformation from the question text to the structured data set is gradually completed.

[0079] Step S102: Obtain second data sets corresponding to each candidate standard question. Each second data set includes a second overall semantic word vector corresponding to the corresponding candidate standard question and a second professional vocabulary set. The second professional vocabulary set includes second professional vocabulary included in the corresponding candidate standard question and second word vectors corresponding to the second professional vocabulary.

[0080] Optionally, second data sets corresponding to each candidate standard question stored in advance can be obtained. The second data set includes the second overall semantic word vector of the candidate standard question, the second professional vocabulary included in the corresponding candidate standard question, and the second word vector corresponding to the second professional vocabulary.

[0081] In an optional embodiment of the present application, the second data set corresponding to each candidate standard question is determined by the following method:

[0082] Determine the second overall semantic word vector corresponding to the candidate standard question through the BERT model;

[0083] According to the professional vocabulary set corresponding to the candidate standard question, determine the second professional vocabulary included in the candidate standard question, and based on the BERT model, determine the second word vector corresponding to each second professional vocabulary;

[0084] Structurally manage and store the second data set corresponding to the candidate standard question and the candidate standard question.

[0085] Optionally, in practical applications, each candidate standard question will have a corresponding answer. Therefore, unique encoding processing will be performed on the candidate standard question and its corresponding answer, so that the candidate standard question and the answer can be associated by the encoding during subsequent semantic matching.

[0086] Among them, for each candidate standard question, the BERT model can be used to determine the overall semantic word vector of each candidate standard question (i.e., the second overall semantic word vector). At this time, the global semantic information of the question can be captured, which can provide support for subsequent matching and analysis. Further, for each candidate standard question, according to the set of professional vocabulary to which the candidate standard question belongs, the candidate standard question can be matched with professional vocabulary, the professional terms therein can be extracted, and the professional vocabulary included in the candidate standard question (i.e., the second professional vocabulary) can be determined. For each professional vocabulary included in the candidate standard question, the BERT model is used to obtain the word vector of each professional vocabulary included in the candidate standard question (i.e., the second word vector).

[0087] Correspondingly, after determining that each candidate standard question has been processed, the overall word vectors of all candidate standard questions and the word vectors of their corresponding professional vocabulary are stored in a structured manner, providing support for subsequent calculation and comparison with the current question input by the user. The stored data includes the number of each candidate standard question, its overall word vector, the professional vocabulary it includes, and the word vectors of the professional vocabulary it includes.

[0088] In this application, by constructing a professional vocabulary table and using a pre-trained language model (such as BERT), hierarchical semantic vector extraction and storage are performed on candidate standard questions and their professional vocabulary, significantly reducing the complexity of real-time calculation and resource consumption, and laying a foundation for subsequent efficient and accurate semantic matching.

[0089] Step S103, for each candidate standard question, according to the second overall semantic word vector corresponding to the candidate standard question and the first overall semantic word vector, determine the semantic distance corresponding to the candidate standard question, and according to the first set of professional vocabulary corresponding to the current question and the second set of professional vocabulary corresponding to the candidate standard question, determine the professional distance corresponding to the candidate standard question.

[0090] Optionally, the semantic distance between the current question and each candidate standard question can be determined according to the overall semantic word vector corresponding to the current question and the overall semantic word vector corresponding to each candidate standard question. Among them, SemanticDist(Q i ,Q c ) is used to calculate the overall semantic vector Q of the current question c and the overall semantic vector Q of the candidate standard question i The distance between them is calculated as follows:

[0091]

[0092] Among them, represents the overall semantic vector Q of the candidate standard question iThe component in the j-th dimension represents the overall semantic vector Q of the current problem c The component in the j-th dimension, U represents the dimension of the vector, and p is the norm coefficient. In this application, when the dimension U = 768 and p = 2, it is the Euclidean distance

[0093] Furthermore, since the BERT model is used to extract word vectors, and the BERT model can encode professional vocabulary by combining context information, even the same professional vocabulary may have differences in the word vectors generated in different problems. To measure this difference, the norm distance formula DomainDist(Q i , Q c ) is used to calculate the vector distance (i.e., professional distance) of professional vocabulary in the current problem and each candidate standard problem. This professional distance is used to measure the current problem Q c and the candidate standard problem Q i in the degree of difference in the professional field

[0094] In an alternative embodiment of the present application, for each candidate standard problem, according to the first set of professional vocabulary corresponding to the current problem and the second set of professional vocabulary corresponding to the candidate standard problem, determining the professional distance corresponding to the candidate standard problem includes:

[0095] Determine the vocabulary intersection and vocabulary union according to the first professional vocabulary included in the current problem and the second professional vocabulary included in the candidate standard problem

[0096] Determine the first quantity of professional vocabulary in the vocabulary intersection, the second quantity of professional vocabulary in the vocabulary union, and the ratio of the first quantity and the second quantity

[0097] Determine the professional distance corresponding to the candidate standard problem according to the ratio, the second quantity, the first set of professional vocabulary, and the second set of professional vocabulary

[0098] Optionally, when determining the professional distance between the current problem and each candidate standard problem, the vocabulary intersection and vocabulary union between the two can be determined according to the first professional vocabulary included in the current problem and the second professional vocabulary included in the candidate standard problem. Then, according to the number of professional vocabulary in the vocabulary intersection (i.e., the first quantity) and the number of professional vocabulary in the vocabulary union (i.e., the second quantity), the ratio of the number of vocabulary between the two is obtained. Further, according to the determined ratio, the second quantity, the first set of professional vocabulary, and the second set of professional vocabulary, the professional distance between the current problem and the candidate standard problem is determined

[0099] For example, let the set of professional vocabulary of the current problem Q c be V c , and the candidate standard problem Q iThe set of professional terms is V i, At this time, the vocabulary intersection is A = V c ∩V i and the vocabulary union is B = V c ∪V i It should be noted that if the current problem contains professional terms while the candidate standard problem does not contain any professional terms, the set A may be an empty set

[0100] In an alternative embodiment of the present application, according to the ratio, the first set of professional terms, and the second set of professional terms, determining the professional distance corresponding to the candidate standard problem includes:

[0101] According to the first set of professional terms and the second set of professional terms, respectively determining the semantic distance corresponding to the vocabulary intersection and the semantic distance corresponding to the vocabulary union;

[0102] According to the ratio and the preset threshold relationship, the second quantity, the semantic distance corresponding to the vocabulary intersection, and the semantic distance corresponding to the vocabulary union, determining the professional distance corresponding to the candidate standard problem.

[0103] Optionally, after determining the vocabulary intersection and the vocabulary union, the semantic distance corresponding to the vocabulary intersection and the semantic distance corresponding to the vocabulary union can be determined according to the first set of professional terms and the second set of professional terms, respectively.

[0104] In an alternative embodiment of the present application, the determining the semantic distance corresponding to the vocabulary intersection and the semantic distance corresponding to the vocabulary union according to the first set of professional terms and the second set of professional terms:

[0105] Determining the word vector corresponding to each professional term in the vocabulary intersection, and according to the word vector corresponding to each professional term in the vocabulary intersection, determining the semantic distance corresponding to the vocabulary intersection, where the word vector corresponding to each professional term in the vocabulary intersection includes at least one of the first word vector corresponding to the professional term in the first set of professional terms and the second word vector corresponding to the professional term in the second set of professional terms;

[0106] Determining the target professional terms in the vocabulary union that do not belong to the vocabulary intersection, and according to the word vector corresponding to the target professional terms, determining the semantic distance corresponding to the vocabulary union, where the word vector corresponding to the target professional terms is the first word vector corresponding to the target professional terms in the first set of professional terms or the second word vector corresponding to the target professional terms in the second set of professional terms.

[0107] Optionally, for the lexical intersection, each professional term in the lexical intersection exists in both the current question and the candidate standard question. At this time, there will be two word vectors corresponding to this professional term, namely the first word vector corresponding to it in the current question and the second word vector corresponding to it in the candidate standard question. Correspondingly, the semantic distance of each professional term in the lexical intersection in the current question and the candidate standard question can be determined, and then according to the semantic distance of each professional term in the lexical intersection in the current question and the candidate standard question, the semantic distance corresponding to the lexical intersection is determined. Optionally, if a professional term appears multiple times, the weighted average of its multiple semantic distances is used as the semantic distance representing this professional term in the current question and the candidate standard question. Among them, for a certain professional term w (w ∈ A) in the lexical intersection A, this professional term w exists in Q c and Q i both. At this time, the semantic distance of this professional term in the current question and the candidate standard question is:

[0108]

[0109] Among them, and are respectively the j-th component in the word vectors representing the professional term w in Q i and Q c . U represents the dimension of the vector, p is the norm coefficient, and A represents the lexical intersection. It can be understood that if A is an empty set, then DA(w) = 0.

[0110] Optionally, for the lexical union, a certain professional term in the lexical union may only exist in the current question or the candidate standard question. At this time, the professional terms in the lexical union that do not belong to the lexical intersection (i.e., the target professional terms) can be determined. For each target professional term, there will be one word vector corresponding to this target professional term, namely the first word vector corresponding to it in the current question or the second word vector corresponding to it in the candidate standard question. Correspondingly, the semantic distance of each target professional term in the current question and the candidate standard question can be determined, and then according to the semantic distance of each target professional term in the current question and the candidate standard question, the semantic distance corresponding to the lexical intersection is determined. Among them, assuming a target professional term w, this target professional term w only exists in Q i or Q c , that is but w ∈ B. At this time, the semantic distance of this target professional term w in the current question and the candidate standard question is:

[0111]

[0112] Among them, w j is the target professional term w in Q i or Qc The j-th component of the word vector in it, U represents the dimension of the vector, p is the norm coefficient, A represents the lexical intersection, B represents the lexical intersection, and Z is a predefined empty vector, which is set to an all-zero vector in this application. It can be understood that if A = B, then DB(w) = 0.

[0113] Furthermore, a preset threshold relationship can be obtained, and then according to the ratio of the number of words in the lexical intersection and the lexical union, the preset threshold relationship, the second quantity, and the semantic distances corresponding to the determined lexical intersection and the lexical union, the professional distance corresponding to the candidate standard question can be determined.

[0114] In an alternative embodiment of the present application, determining the professional distance corresponding to the candidate standard question according to the ratio, the preset threshold relationship, the second quantity, the semantic distance corresponding to the lexical intersection, and the semantic distance corresponding to the lexical union includes:

[0115] If the ratio meets the first threshold, obtain a first amplification coefficient, and determine the professional distance corresponding to the candidate standard question according to the first amplification coefficient, the second quantity, and the semantic distance corresponding to the lexical union. The first threshold indicates that the lexical intersection is the same as the lexical union;

[0116] If the ratio meets the second threshold, obtain a second amplification coefficient, and determine the professional distance corresponding to the candidate standard question according to the second amplification coefficient, the second quantity, and the semantic distance corresponding to the lexical intersection. The second threshold indicates that the lexical intersection is an empty set. The first threshold is less than the second threshold, and the first amplification coefficient is greater than the second amplification coefficient;

[0117] If the ratio does not meet the first threshold and the second threshold, then determine a third amplification coefficient according to the ratio, and determine the professional distance corresponding to the candidate standard question according to the third amplification coefficient, the second quantity, the semantic distance corresponding to the lexical union, and the semantic distance corresponding to the lexical intersection.

[0118] Optionally, the ratio of the number of words in the lexical intersection and the lexical union reflects the situation of the amplification coefficient. This amplification coefficient is to more reasonably measure the influence of the professional word matching degree on the semantic distance, adjust the relationship between the lexical intersection and the lexical union non-linearly, and can dynamically adjust the weight of the semantic distance in scenarios of complete match, partial match, and complete non-match.

[0119] Among them, if there is a large overlap in the professional vocabulary included in the current problem and the candidate standard problem (that is, when the sizes of set A and set B are close), it indicates a high semantic similarity between the two in the professional field. At this time, the amplification factor can be set to 1, making the calculation result of the semantic distance smaller, so as to reflect the high matching degree between the two. When the overlap of the professional vocabulary included in the candidate standard problem and the current problem is small (for example, set A is empty or much smaller than set B), it indicates a large semantic difference between the two in the professional field. In this case, the amplification factor should be increased to amplify the result of the semantic distance to reflect the low matching degree between the two. Optionally, the calculation formula of the amplification factor is:

[0120]

[0121] Among them, N A is the size of the number of professional vocabulary in set A, N B is the size of the number of professional vocabulary in set B, a is a constant term, k is a parameter that controls the attenuation relationship between the vocabulary intersection and the vocabulary union, and when a = 2 and k = 1.45, it can better satisfy the properties of γ.

[0122] Among them, the properties of the amplification factor γ are:

[0123] 1. When N A = N B (that is, the vocabulary intersection is equal to the vocabulary union), γ = 1, indicating that the vocabulary intersection coincides with the vocabulary union, and no amplification is required at this time;

[0124] 2. When N A = 0 (that is, the vocabulary intersection A is an empty set), regardless of the size of N B , the amplification factor γ = 2;

[0125] 3. When N A = 0.5N B (that is, the vocabulary intersection is half of the vocabulary union), the amplification factor γ = 1.2.

[0126] In this application, in order to more accurately measure the differences in the professional field, when determining the professional distance, for the same professional vocabulary, it is required that its distance should be as small as possible, while for different professional vocabulary or missing professional vocabulary, it is required that its distance should be as large as possible. Based on this, the overall semantic vector Q c of the current problem and the overall semantic vector Q i of the candidate standard problem, the professional distance between them is defined as the average distance after the sum of the overall semantic distances DA(w) and DB(w):

[0127]

[0128] Among them, N Bis the quantity size of the professional vocabulary in the vocabulary union B (i.e., the second quantity), and D(w) represents the semantic distance between a certain professional vocabulary in the current question and the candidate standard question, specifically including the semantic distance DA(w) corresponding to the vocabulary intersection and the semantic distance DA(w) corresponding to the vocabulary union. The amplification factor γ is used to adjust the weights of the vocabulary intersection and the vocabulary union.

[0129] Optionally, in this application, different amplification factors are set according to the comparison between the ratio of the number of vocabulary in the vocabulary intersection and the vocabulary union and a preset threshold. Specifically, a first threshold is set to represent the situation where the vocabulary intersection and the vocabulary union are the same. If the ratio meets the first threshold, it means the vocabulary intersection and the vocabulary union are the same, and the amplification factor is the first amplification factor. At this time, according to the first amplification factor, the second quantity, and the semantic distance corresponding to the vocabulary union, the professional distance corresponding to the candidate standard question is determined; a second threshold is set to represent the situation where the vocabulary intersection is an empty set (the first threshold is less than the second threshold). When the ratio meets the second threshold, it means the vocabulary intersection is an empty set, and the amplification factor is the second amplification factor (the first amplification factor is greater than the second amplification factor). At this time, according to the second amplification factor, the second quantity, and the semantic distance corresponding to the vocabulary union, the professional distance corresponding to the candidate standard question is determined; conversely, if the ratio does not meet the first threshold and the second threshold, the ratio is substituted into the formula of the amplification factor to determine the amplification factor (i.e., the third amplification factor), and then according to the third amplification factor, the second quantity, and the semantic distance corresponding to the vocabulary union, the professional distance corresponding to the candidate standard question is determined.

[0130] Among them, the ratio of the number of vocabulary in the vocabulary intersection and the vocabulary union is N A / N B , the first threshold is 0, and the second threshold is 1. If N A / N B = 0, the amplification factor is the first amplification factor, and the specific value is γ = 2; if N A / N B = 1, the amplification factor is the second amplification factor, and the specific value is γ = 1; if N A / N B is not 0 and not 1 either, the amplification factor is the third amplification factor, and the third amplification factor needs to be specifically calculated according to the calculation formula of the amplification factor.

[0131] Step S104, according to the semantic distance and professional distance corresponding to each candidate standard question, determine the comprehensive distance between each candidate standard question and the current question, and use the standard answer corresponding to the candidate standard question with the smallest comprehensive distance as the answer to the current question.

[0132] Optionally, after obtaining the semantic distance and professional distance corresponding to the current question and each candidate standard question, for each candidate question, the comprehensive distance between the current question and the candidate standard question can be determined based on the semantic distance and professional distance between the current question and the candidate standard question, and then the comprehensive distance between the current question and the candidate standard question is compared, and the standard answer corresponding to the candidate standard question with the smallest comprehensive distance is taken as the answer to the current question.

[0133] In an optional embodiment of the present application, determining the comprehensive distance between each candidate standard question and the current question according to the semantic distance and professional distance corresponding to each candidate standard question includes:

[0134] Determine the number of first professional vocabulary included in the current question, and determine a first weight coefficient corresponding to the semantic distance and a second weight coefficient corresponding to the professional distance according to the number of vocabulary;

[0135] For each candidate standard question, the comprehensive distance between the candidate standard question and the current question is determined according to the first weight coefficient and the second weight coefficient, and the semantic distance and the professional distance corresponding to the candidate standard question.

[0136] In order to comprehensively measure the matching degree between the current question and the candidate standard question in terms of semantics and professional fields, a weighted formula is designed to combine the semantic distance and professional field distance.

[0137] In the prior art, when selecting the appropriate standard question answer for the user input question, the shortest distance between the current question and the standard question was originally calculated to select the answer to the standard question with the closest distance. However, in the professional field, since the BERT model fails to fully reflect the word vector characteristics of professional vocabulary, and the professional vocabulary is constantly updated, it is not appropriate to retrain the BERT model. Therefore, on the basis of calculating the overall distance between the current question and the standard question, the present application adds the calculation of the word vector distance between the professional vocabulary in the current question and the vocabulary in the standard candidate question. The original shortest distance calculation measures the closeness of the standard semantics, and the distance between the newly added professional vocabulary represents the degree of fit between the current question and the candidate standard question in the professional field. Based on this, by weighted summing these two distances according to the weights, the candidate standard question with the smallest comprehensive distance is found, and its corresponding answer is the answer required for the current question.

[0138] Optionally, in the present application, the weight coefficient corresponding to the semantic distance is the first weight coefficient, and the weight coefficient corresponding to the professional distance is the second weight coefficient. The first weight coefficient and the second weight coefficient are used to adjust the influence of the two on the comprehensive distance. The first weight coefficient and the second weight coefficient can be determined according to the number of professional words included in the previous question. At this time, the current question Q c Question Q with candidate criteriai The comprehensive distance can be expressed as:

[0139] D(Q i ) = α·SemanticDist(Q i , Q c ) + β·DomainDist(Q i , Q c )

[0140] Among them, Q i is the word vector of the candidate standard question, Q c is the word vector of the current question, D(Q i ) represents the comprehensive distance between the current question Q c and the candidate standard question Q i , α is the first weight coefficient (i.e., the weight coefficient of the semantic distance), β is the second weight coefficient (i.e., the weight coefficient of the professional distance), SemanticDist(Q i , Q c ) represents the semantic distance between the current question and the candidate standard question, DomainDist(Q i , Q c ) represents the professional domain distance between the current question and the candidate standard question.

[0141] Furthermore, combining the specific formula of the semantic distance and the formula of the professional distance between the current question and the candidate standard question above, the comprehensive distance between each candidate standard question and the current question can be determined by the following formula:

[0142]

[0143] Among them, D(Q y ) represents the comprehensive distance, α is the first weight coefficient, β is the second weight coefficient, represents the j-th component of the second overall semantic word vector Q i (i.e., the second overall semantic word vector) of the candidate standard question, represents the word vector Q c (i.e., the second overall semantic word vector) of the current question, U represents the dimension of the vector, p is the norm coefficient, A is the vocabulary intersection, B is the vocabulary union, γ is the amplification coefficient, N B is the second quantity, w j is the first word vector or the corresponding second word vector corresponding to the professional vocabulary w, and are respectively the j-th components of the professional vocabulary w in the corresponding first word vector and second word vector.

[0144] In an alternative embodiment of the present application, the sum of the first weight and the second weight is 1. Determining the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words includes:

[0145] Determining the second weight coefficient corresponding to the professional distance according to the number of words;

[0146] Determining the first weight coefficient corresponding to the semantic distance according to the second weight coefficient.

[0147] Optionally, in order to flexibly balance the importance of semantic matching and professional field matching, parameters α and β need to be dynamically adjusted. The values of these two parameters are not fixed, but should vary dynamically according to the number of professional words involved in the current problem. The basic requirements for dynamic adjustment include:

[0148] (1) Both α and β are decimals greater than zero, and the maximum value is 1;

[0149] (2) If the number of professional words involved in the current problem is 0, it means that it depends entirely on semantic matching. At this time, α = 1 and β = 0;

[0150] (3) α and β are complementary, and satisfy the relationship of α P +β P = 1, where p is a positive integer. In the present application, p = 1 is set.

[0151] Correspondingly, when performing dynamic adjustment, if the number of professional words involved in the current problem is small, it means that the professional semantic features of the current problem are weak, and the professional words will be covered by the semantics. At this time, more reliance on professional field matching is required. Therefore, the weight coefficient β of the professional distance should be larger, and the weight coefficient α of the semantic distance should be smaller, but not less than 0.5 at the lowest. On the contrary, if the number of professional words involved in the current problem is large, it means that the professional features of the current problem are stronger. At this time, various professional characteristics can be included in the semantics, and it is not necessary to rely solely on professional field matching. Therefore, the weight coefficient α of the semantic distance should be increased, and the weight coefficient β of the professional distance should be decreased. It can be seen that the weight coefficient β of the professional distance in the present application satisfies non-linearity. To implement the above dynamic adjustment logic, in the present application, β (i.e., the second weight coefficient) is set as a non-linear function of the number of professional words in the following form:

[0152]

[0153] where β is the second weight coefficient, is the number of words, a is an adjustment parameter, k is a non-linear adjustment parameter. In the present application, a > 0 is used to control the growth rate, and k > 1 is used to adjust the decreasing trend of β.

[0154] It can be seen that the parameters of this non-linear function have some constraints, and the parameters a and k can be determined according to the following conditions specifically:

[0155] 1) When β = 0, indicating complete dependence on semantic matching, and at this time β = 1.

[0156] 2) When β = 0.5, indicating that the weight coefficients of semantic matching and professional field matching are equal.

[0157] 3) When β is close to 0.1, indicating that the weight of the professional field is small and semantic matching dominates.

[0158] According to these conditions, it is calculated that a = 1 and k = 2. Of course, in actual applications, the values of a and k can also be recalculated according to other conditions.

[0159] In actual applications, when the number of professional terms involved in the current problem is greater than 10 (i.e., ), it can be considered that β is close to 0. At this time, it mainly depends on semantic matching. When is a positive integer and takes different values, the values of α and β are shown in the following table:

[0160]

[0161] In this application, the dynamically adjusted values of α and β can better adapt to the changes in the number of professional terms in different problems, so as to achieve the balance between semantic matching and professional field matching.

[0162] In the embodiments of this application, determining the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words includes:

[0163] Determine whether the number of words meets the preset conditions;

[0164] If the number of words does not meet the preset conditions, set the first weight coefficient corresponding to the semantic distance to 1 and set the second weight coefficient corresponding to the professional distance to 0;

[0165] If the number of words meets the preset conditions, determine the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words.

[0166] Optionally, in practical applications, when there are too many or no professional terms involved in the current problem, the determined comprehensive distance mainly depends on the semantic distance. At this time, the first weight coefficient corresponding to the semantic distance can be set to 1, and the second weight coefficient corresponding to the professional distance can be set to 0. For example, when the number of professional terms in the current problem is 0, it indicates that the problem is not a professional problem. If the number of professional terms exceeds 10, although the problem belongs to a professional problem, its overall semantics already fully contains the characteristics of the professional problem. In both of these cases, the similarity of the problem can be directly characterized by the overall semantic distance without the need to additionally introduce the distance in the professional field, that is, α = 1 and β = 0.

[0167] In this application, for the current problem input by the user, first, the professional terms included in the current problem and their corresponding word vectors are determined. Then, the comprehensive semantic distance is calculated based on the professional terms and their corresponding word vectors included in the current problem and each candidate standard problem. Finally, the problem with the smallest comprehensive semantic distance is selected from all candidate standard problems, and its corresponding answer is output as the final result. The comprehensive semantic distance obtained at this time includes the overall semantic distance and the professional term distance between the current problem and the candidate standard problem, and the professional term distance can fully reflect the word vector characteristics of the professional terms and represent the degree of fit between the current problem and the candidate standard problem in the professional field. Therefore, the final result determined based on the comprehensive semantic distance can more accurately reflect the semantic differences of professional terms in a specific field and improve the matching effect of question and answer in the professional field.

[0168] To better understand the method provided in the embodiments of this application, the execution process of this method will be described in detail below. As Figure 2 shown, the method mainly includes three steps: preprocessing the candidate standard problems, processing the current problem, and outputting the best answer. Among them,

[0169] I. Preprocessing of candidate standard problems: By constructing a professional term table and combining syntactic analysis, the professional terms included in the candidate standard problems are identified to form a corresponding set of professional terms. Subsequently, the BERT model is used to generate the overall word vector of the candidate standard problem and the word vectors of each professional term therein, and this vector information is stored for subsequent use. By preprocessing the candidate standard problems, the resource consumption and time overhead during real-time calculation can be effectively reduced, and the real-time response efficiency can be improved.

[0170] II. Processing of the current problem: When the user inputs the current problem, the system first uses the BERT model to generate the overall semantic word vector of this problem. Then, the system matches the current problem through the pre-constructed professional term table and extracts the set of professional terms therein. If the set of professional terms is not empty, the word vectors of each professional term in the current problem are further generated to provide support for subsequent processing.

[0171] III. Output the best answer: According to the current question input by the user, calculate the comprehensive distance with the candidate standard questions in turn, and finally select the answer corresponding to the candidate standard question with the smallest comprehensive distance for output.

[0172] The following is a detailed description of these three steps.

[0173] As Figure 3 shown, the preprocessing of the candidate standard questions includes the following steps:

[0174] Step 1: Construct a professional vocabulary: Analyze relevant literature, data, and standard questions in each field manually, extract common professional terms, and construct a professional vocabulary for each field. This vocabulary will serve as the basis for identifying professional terms in subsequent steps to ensure accurate extraction and processing of key terms in the standard questions.

[0175] Step 2: Encode the standard questions and answers: Perform unique encoding on the candidate standard questions and their corresponding answers so that the candidate standard questions and answers can be associated with the encoding.

[0176] Step 3: Select a standard question: Select a question from the candidate standard question set in turn. If it is the first selection, select the first candidate standard question; if not, select the next standard question of the current candidate standard question in turn.

[0177] Step 4: Obtain word vectors: Use the BERT model to generate an overall semantic word vector for the currently selected candidate standard question, capture the global semantic information of the question, and provide support for subsequent matching and analysis.

[0178] Step 5: Obtain the set of professional vocabulary: Combine the constructed professional vocabulary to perform professional vocabulary matching on the current candidate standard question, extract the professional terms therein, and form the set of professional vocabulary for this candidate standard question.

[0179] Step 6: Whether the set is empty: Judge whether the set of professional vocabulary of the current candidate standard question is empty. If the set is empty, it means that the question does not contain professional terms, and go to Step 8; otherwise, go to Step 7.

[0180] Step 7: Obtain the word vectors of professional vocabulary: For each word in the set of professional vocabulary, use the BERT model to obtain the word vectors of all professional vocabulary in the set in the current candidate standard question.

[0181] Step 8: Whether all standard questions have been traversed: Check whether all candidate standard questions have been traversed. If all questions have been processed, go to Step 8; otherwise, it means that the remaining candidate standard questions need to be processed continuously, and go to Step 3.

[0182] Step 9: Store the word vectors of standard questions and professional vocabulary: Store the overall word vectors of all candidate standard questions and the word vectors of their corresponding professional vocabulary in a structured manner to support subsequent calculations and comparisons with the current question input by the user. The stored data includes the number of each candidate standard question, its overall word vector, the corresponding set of professional vocabulary, and the word vectors of these professional vocabulary.

[0183] It can be seen that the preprocessing process of candidate standard questions, by constructing a professional vocabulary table and using a pre-trained language model (such as BERT), as well as performing hierarchical semantic vector extraction and storage on candidate standard questions and their professional vocabulary, significantly reduces the complexity and resource consumption of real-time calculations by traversing and storing the word vectors of all candidate standard questions and their professional vocabulary in advance, laying a foundation for subsequent efficient and accurate semantic matching.

[0184] As Figure 4 shown, the processing of the current question includes the following steps:

[0185] Step 1: User inputs the current question: The user inputs the current question through the digital human, with input methods of voice or text. After the input is completed, proceed to Step 2.

[0186] Step 2: Obtain the current word vector: The system uses the BERT model to generate the overall semantic word vector of the user's current question. After the processing is completed, proceed to Step 3.

[0187] Step 3: Obtain the collection of professional vocabulary: Match the current question with the professional vocabulary table corresponding to the field of the user's current question, extract the professional terms therein, and form the set of professional vocabulary of the current question. After completion, proceed to Step 4.

[0188] Step 4: Whether the set is empty: Check whether the set of professional vocabulary of the current question is empty. If the set is empty, it means that the question input by the user does not contain professional terms, and directly proceed to Step 6. If the set is non-empty, then proceed to Step 5.

[0189] Step 5: Obtain the word vectors of professional vocabulary: For the set of professional vocabulary of the current question, use the BERT model to generate the word vectors of these professional vocabulary. After the processing is completed, proceed to Step 6.

[0190] Step 6: Form the current question dataset: Integrate the overall word vector of the current question, the set of professional vocabulary, and the word vectors corresponding to the included professional vocabulary to generate a complete dataset of the current question to support subsequent calculations and processing.

[0191] It can be seen that the processing of the current problem has gradually completed the transformation from the problem text to the structured data set through the processing and semantic analysis of the user's input problem, laying a foundation for subsequent calculations and processing.

[0192] As Figure 5 shown, the steps for outputting the best answer are as follows:

[0193] Step 1: Data set of the current problem: Continuing from the second step, generate the corresponding data set for the current problem in real-time processing of the user's input, including the overall word vector of the current problem and the word vectors of the professional vocabulary included in the current problem. After completion, proceed to Step 2.

[0194] Step 2: Read the candidate standard problem set: Read the processing results of all stored candidate standard problems, including the unique encoding of each candidate standard problem, its overall word vector, and the professional vocabulary included in the candidate standard problem and their corresponding word vectors. After completion, proceed to Step 3.

[0195] Step 3: Set Mindis = 9999, MinNo = -1: Set the variable Mindis = 9999 as the initial value of the distance between the current problem and the candidate standard problem. Here, a relatively large number is taken. At the same time, set MinNo to -1 to store the standard problem number corresponding to the minimum distance, with the initial value set to an impossible number. After completion, proceed to Step 4.

[0196] Step 4: Select a standard problem: Sequentially select a candidate standard problem from the stored candidate standard problem set. If it is the first selection, select the first record in the set; if it is not the first selection, select the next record of the current record. After completion, proceed to Step 5.

[0197] Step 5: Calculate SemanticDis: Calculate the semantic distance SemanticDis between the overall semantic word vector of the current problem and the overall semantic word vector of the candidate standard problem using the formula. After completion, proceed to Step 6.

[0198] Step 6: The number of professional vocabulary in the current problem = 0 or > 10: Check the number of professional vocabulary in the current problem and determine whether it is 0 or exceeds 10. If the judgment result is "yes", proceed to Step 20; otherwise, proceed to Step 7.

[0199] Step 7: Determine α and β: Determine the corresponding α and β values according to the number of professional vocabulary included in the current problem. After completion, proceed to Step 8.

[0200] Step 8: DA(w) = 0, DB(w) = 0: Set the initial values of both DA(w) and DB(w) to 0. After completion, proceed to Step 9.

[0201] Step 9: Determine sets A and B: Calculate the intersection and union of the professional vocabulary included in the current problem and the professional vocabulary included in the candidate standard problems respectively. Define the intersection as set A and the union as set B. After completion, proceed to Step 10.

[0202] Step 10: Determine N A and N B : N A represents the number of elements in set A. If A is an empty set, then N A = 0, and N B represents the number of elements in set B. After completion, proceed to Step 11.

[0203] Step 11: N A / N B = 0: Calculate the value of N A / N B and determine whether it is equal to 0. If it is equal to 0, it means set A is empty, and proceed to Step 12; otherwise, proceed to Step 14.

[0204] Step 12: γ = 2: Set the value of γ to 2. After completion, proceed to Step 13.

[0205] Step 13: Calculate DB(w): According to the calculation formula of DB(w), use set B to find the professional vocabulary vectors corresponding to its elements. After completion, proceed to Step 19.

[0206] Step 14: N A / N B = 1: Calculate the value of N A / N B and determine whether it is equal to 1. If it is equal to 1, proceed to Step 15; otherwise, proceed to Step 17.

[0207] Step 15: γ = 1: Set the value of γ to 1. After completion, proceed to Step 16.

[0208] Step 16: Calculate DA(w): According to the calculation formula of DA(w), use set A to find the professional vocabulary vectors corresponding to its elements. After completion, proceed to Step 19.

[0209] Step 17: Calculate γ: Calculate the value of γ using the formula. After completion, proceed to Step 18.

[0210] Step 18: Calculate DA(w) and DB(w): Calculate DA(w) and DB(w) respectively according to set A and set B. After completion, proceed to Step 19.

[0211] Step 19: Calculate DomainDis: Use N B, the values of γ, DA(w), and DB(w), calculate the professional distance DomainDis, and after completion, proceed to step 21.

[0212] Step 20: D(Q i ) = SemanticDis: Set the comprehensive distance D(Q i ) to the overall semantic distance SemanticDis between the current problem and the candidate standard problem. After completion, proceed to step 22.

[0213] Step 21: Calculate D(Q i ): According to the formula, calculate the comprehensive distance D(Q i ). After completion, proceed to step 22.

[0214] Step 22: D(Q i ) < Mindis: Determine whether D(Q i ) is less than Mindis. If D(Q i ) is smaller, proceed to step 23; otherwise, proceed to step 24.

[0215] Step 23: Mindis = D(Q i ), MinNo = the current standard record number: Assign the value of D(Q i ) to Mindis, and assign the record number of the current candidate standard problem to MinNo. At this time, Mindis stores the current minimum comprehensive distance value, and MinNo corresponds to the standard problem number with the minimum distance. After completion, proceed to step 24.

[0216] Step 24: Have all standard problem sets been traversed: Check whether all candidate standard problems have been traversed. If all have been traversed, it means the best - matching candidate standard problem number has been found, proceed to step 25; otherwise, return to step 4 and continue searching.

[0217] Step 25: Select and output the answer corresponding to MinNo: According to the candidate standard problem number stored in MinNo, find its corresponding answer and output the answer for the digital human to use to answer the current question.

[0218] It can be seen that the best answer is output by calculating the semantic distance between the current problem and the candidate standard problem set, gradually traversing all candidate standard problems, finding the matching item with the minimum comprehensive distance, and outputting the corresponding answer. The entire process focuses on set - relationship judgment, distance calculation, and optimal - matching selection, ensuring the accuracy and efficiency of the final result.

[0219] The embodiment of the present application provides a semantic recognition device based on professional vocabulary, such as Figure 6As shown, the device may include: a current problem acquisition module 601, a candidate problem acquisition module 602, a distance determination module 603, and a result determination module 604, where

[0220] The current problem acquisition module is configured to acquire a current problem and a target professional vocabulary set, and determine a first data set corresponding to the current problem according to the target professional vocabulary set. The first data set includes a first overall semantic word vector corresponding to the current problem and a first professional vocabulary set. The first professional vocabulary set includes a first professional vocabulary included in the current problem and a first word vector corresponding to the first professional vocabulary.

[0221] The candidate problem acquisition module is configured to acquire a second data set corresponding to each candidate standard problem. Each second data set includes a second overall semantic word vector corresponding to the corresponding candidate standard problem and a second professional vocabulary set. The second professional vocabulary set includes a second professional vocabulary included in the corresponding candidate standard problem and a second word vector corresponding to the second professional vocabulary.

[0222] The distance determination module is configured to, for each candidate standard problem, determine a semantic distance corresponding to the candidate standard problem according to the second overall semantic word vector corresponding to the candidate standard problem and the first overall semantic word vector, and determine a professional distance corresponding to the candidate standard problem according to the first professional vocabulary set corresponding to the current problem and the second professional vocabulary set corresponding to the candidate standard problem.

[0223] The result determination module is configured to determine a comprehensive distance between each candidate standard problem and the current problem according to the semantic distance and the professional distance corresponding to each candidate standard problem, and use the standard answer corresponding to the candidate standard problem with the smallest comprehensive distance as the answer to the current problem.

[0224] Optionally, the target professional vocabulary set is a professional vocabulary set corresponding to the field to which the current problem belongs. When the current problem acquisition module determines the first data set corresponding to the current problem according to the target professional vocabulary set, it is specifically configured to:

[0225] Input the current problem into the BERT model to obtain a first overall semantic word vector corresponding to the current problem;

[0226] Extract professional vocabulary from the current problem based on the target professional vocabulary set to obtain a first professional vocabulary included in the current problem;

[0227] Input the first professional vocabulary into the BERT model to obtain a first word vector corresponding to the first professional vocabulary.

[0228] Optionally, when the result determination module determines the comprehensive distance between each candidate standard problem and the current problem according to the semantic distance and the professional distance corresponding to each candidate standard problem, it is specifically configured to:

[0229] Determine the number of words in the first professional vocabulary included in the current problem, and based on the number of words, determine the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance;

[0230] For each candidate standard problem, determine the comprehensive distance between the candidate standard problem and the current problem according to the first weight coefficient and the second weight coefficient, as well as the semantic distance and the professional distance corresponding to the candidate standard problem.

[0231] Optionally, when the result determination module determines the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words, it is specifically used for:

[0232] Determine whether the number of words meets the preset conditions;

[0233] If the number of words does not meet the preset conditions, set the first weight coefficient corresponding to the semantic distance to 1 and the second weight coefficient corresponding to the professional distance to 0;

[0234] If the number of words meets the preset conditions, determine the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words.

[0235] Optionally, the sum of the first weight and the second weight is 1. When the result determination module determines the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words, it is specifically used for:

[0236] Determine the second weight coefficient corresponding to the professional distance according to the number of words;

[0237] Determine the first weight coefficient corresponding to the semantic distance according to the second weight coefficient;

[0238] Among them, the second weight coefficient is determined by the following method:

[0239]

[0240] Among them, β is the second weight coefficient, is the number of words, a is the adjustment parameter, and k is the non-linear adjustment parameter.

[0241] Optionally, when the distance determination module determines the professional distance corresponding to each candidate standard problem according to the first professional vocabulary set corresponding to the current problem and the second professional vocabulary set corresponding to the candidate standard problem, it is specifically used for:

[0242] Determine the vocabulary intersection and vocabulary union according to the first professional vocabulary included in the current problem and the second professional vocabulary included in the candidate standard problem;

[0243] Determine the first quantity of professional terms in the lexical intersection and the second quantity of professional terms in the lexical union, as well as the ratio of the first quantity to the second quantity;

[0244] Determine the professional distance corresponding to the candidate standard question according to the ratio, the second quantity, the first set of professional terms, and the second set of professional terms.

[0245] Optionally, when determining the professional distance corresponding to the candidate standard question according to the ratio, the first set of professional terms, and the second set of professional terms, the distance determination module is specifically used for:

[0246] Determine the semantic distance corresponding to the lexical intersection and the semantic distance corresponding to the lexical union according to the first set of professional terms and the second set of professional terms respectively;

[0247] Determine the professional distance corresponding to the candidate standard question according to the ratio and the preset threshold relationship, the second quantity, the semantic distance corresponding to the lexical intersection, and the semantic distance corresponding to the lexical union.

[0248] Optionally, when determining the professional distance corresponding to the candidate standard question according to the ratio and the preset threshold relationship, the second quantity, the semantic distance corresponding to the lexical intersection, and the semantic distance corresponding to the lexical union, the distance determination module is specifically used for:

[0249] If the ratio satisfies the first threshold, obtain the first amplification factor, and determine the professional distance corresponding to the candidate standard question according to the first amplification factor, the second quantity, and the semantic distance corresponding to the lexical union. The first threshold indicates that the lexical intersection is the same as the lexical union;

[0250] If the ratio satisfies the second threshold, obtain the second amplification factor, and determine the professional distance corresponding to the candidate standard question according to the second amplification factor, the second quantity, and the semantic distance corresponding to the lexical intersection. The second threshold indicates that the lexical intersection is an empty set. The first threshold is less than the second threshold, and the first amplification factor is greater than the second amplification factor;

[0251] If the ratio does not satisfy the first threshold and the second threshold, then determine the third amplification factor according to the ratio, and determine the professional distance corresponding to the candidate standard question according to the third amplification factor, the second quantity, the semantic distance corresponding to the lexical union, and the semantic distance corresponding to the lexical intersection.

[0252] Optionally, when determining the semantic distance corresponding to the lexical intersection and the semantic distance corresponding to the lexical union according to the first set of professional terms and the second set of professional terms respectively, the distance determination module is specifically used for:

[0253] Determine the word vectors corresponding to each professional term in the lexical intersection. According to the word vectors corresponding to each professional term in the lexical intersection, determine the semantic distance corresponding to the lexical intersection. The word vectors corresponding to each professional term in the lexical intersection include at least one of the first word vector corresponding to the professional term in the first set of professional terms and the second word vector corresponding to the professional term in the second set of professional terms;

[0254] Determine the target professional terms in the lexical union that do not belong to the lexical intersection. According to the word vectors corresponding to the target professional terms, determine the semantic distance corresponding to the lexical union. The word vector corresponding to the target professional term is the first word vector corresponding to the target professional term in the first set of professional terms or the second word vector corresponding to the target professional term in the second set of professional terms.

[0255] Optionally, the comprehensive distance between each candidate standard question and the current question is determined by the following formula:

[0256]

[0257] where D(Q i ) represents the comprehensive distance, α is the first weight coefficient, β is the second weight coefficient, represents the second overall semantic word vector Q of the candidate standard question i at the j-th dimension component, represents the second overall semantic word vector Q of the current question c at the j-th dimension component, U represents the dimension of the vector, p is the norm coefficient, A is the lexical intersection, B is the lexical union, γ is the amplification coefficient, N B is the second quantity, w j is the first word vector corresponding to the professional term w or the corresponding second word vector, and are respectively the j-th components representing the professional term w in the corresponding first word vector and the second word vector.

[0258] A semantic recognition device based on professional terms in this embodiment can execute a semantic recognition method based on professional terms shown in the embodiments of the present application. The implementation principle is similar and will not be elaborated here.

[0259] The embodiments of the present application provide an electronic device. The electronic device in the embodiments of the present application includes: a processor; and a memory configured to store machine-readable instructions, which when executed by the processor, cause the processor to execute a semantic recognition method based on professional terms.

[0260] The embodiments of the present application provide an electronic device, as Figure 7 shown, Figure 7The electronic device shown includes: a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are connected, such as connected through a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in practical applications, the transceiver 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation on the embodiments of the present application.

[0261] The processor 2001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0262] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI bus or an EISA bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0263] The memory 2003 may be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM, a CD-ROM or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0264] The memory 2003 is used to store the application program code for executing the solution of the present application, and is controlled by the processor 2001 to execute. The processor 2001 is used to execute the application program code stored in the memory 2003 to implement Figure 6 the actions of a semantic recognition device based on professional vocabulary provided by the illustrated embodiment.

[0265] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0266] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A semantic recognition method based on professional vocabulary, characterized in that: include: Obtaining a current question and a target professional vocabulary set, and determining a first data set corresponding to the current question according to the target professional vocabulary set, wherein the first data set includes a first overall semantic word vector and a first professional vocabulary set corresponding to the current question, and the first professional vocabulary set includes a first professional vocabulary included in the current question and a first word vector corresponding to the first professional vocabulary; Obtain a second data set corresponding to each candidate standard question, each of the second data sets including a second overall semantic word vector corresponding to the corresponding candidate standard question and a second professional vocabulary set, the second professional vocabulary set including a second professional vocabulary included in the corresponding candidate standard question and a second word vector corresponding to the second professional vocabulary; For each candidate standard question, determine the semantic distance corresponding to the candidate standard question according to the second overall semantic word vector corresponding to the candidate standard question and the first overall semantic word vector, and determine the professional distance corresponding to the candidate standard question according to the first professional vocabulary set corresponding to the current question and the second professional vocabulary set corresponding to the candidate standard question; According to the semantic distance and professional distance corresponding to each of the candidate standard questions, the comprehensive distance between each of the candidate standard questions and the current question is determined, and the standard answer corresponding to the candidate standard question with the smallest comprehensive distance is used as the answer to the current question.

2. The method according to claim 1, characterized in that: The target professional vocabulary set is a professional vocabulary set corresponding to the field to which the current question belongs, and determining the first data set corresponding to the current question according to the target professional vocabulary set includes: Input the current question into the BERT model to obtain a first overall semantic word vector corresponding to the current question; Extracting professional vocabulary from the current question based on the target professional vocabulary set to obtain a first professional vocabulary included in the current question; The first professional vocabulary is input into the BERT model to obtain a first word vector corresponding to the first professional vocabulary.

3. The method according to claim 1, characterized in that Determining the comprehensive distance between each of the candidate standard questions and the current question according to the semantic distance and the professional distance corresponding to each of the candidate standard questions includes: Determine the number of first professional vocabulary included in the current question, and determine a first weight coefficient corresponding to the semantic distance and a second weight coefficient corresponding to the professional distance according to the number of vocabulary; For each of the candidate standard questions, the comprehensive distance between the candidate standard question and the current question is determined according to the first weight coefficient and the second weight coefficient, and the semantic distance and the professional distance corresponding to the candidate standard question.

4. The method according to claim 3, characterized in that The determining, according to the number of words, a first weight coefficient corresponding to the semantic distance and a second weight coefficient corresponding to the professional distance includes: Determining whether the number of words meets a preset condition; If the number of words does not meet the preset condition, the first weight coefficient corresponding to the semantic distance is set to 1, and the second weight coefficient corresponding to the professional distance is set to 0; If the number of words meets a preset condition, a first weight coefficient corresponding to the semantic distance and a second weight coefficient corresponding to the professional distance are determined according to the number of words.

5. The method according to claim 3, characterized in that: The sum of the first weight and the second weight is 1, and determining the first weight coefficient corresponding to the semantic distance and the second weight coefficient corresponding to the professional distance according to the number of words includes: Determining a second weight coefficient corresponding to the professional distance according to the number of words; Determine the first weight coefficient corresponding to the semantic distance according to the second weight coefficient; The second weight coefficient is determined by the following method: Wherein, β is the second weight coefficient, is the number of words, a is an adjustment parameter, and k is a nonlinear adjustment parameter.

6. The method according to claim 1, characterized in that For each candidate standard question, determining the professional distance corresponding to the candidate standard question according to the first professional vocabulary set corresponding to the current question and the second professional vocabulary set corresponding to the candidate standard question includes: Determine a vocabulary intersection and a vocabulary union according to the first professional vocabulary included in the current question and the second professional vocabulary included in the candidate standard question; Determine a first number of professional words in the vocabulary intersection and a second number of professional words in the vocabulary union, and a ratio of the first number to the second number; The professional distance corresponding to the candidate standard question is determined according to the ratio, the second number, the first professional vocabulary set, and the second professional vocabulary set.

7. The method according to claim 6, characterized in that Determining the professional distance corresponding to the candidate standard question according to the ratio, the first professional vocabulary set, and the second professional vocabulary set includes: Determining, according to the first professional vocabulary set and the second professional vocabulary set, respectively, a semantic distance corresponding to the vocabulary intersection and a semantic distance corresponding to the vocabulary union; The professional distance corresponding to the candidate standard question is determined based on the relationship between the ratio and a preset threshold, the second number, the semantic distance corresponding to the vocabulary intersection, and the semantic distance corresponding to the vocabulary union.

8. The method according to claim 7, characterized in that The determining, according to the relationship between the ratio and a preset threshold value, the second number, the semantic distance corresponding to the vocabulary intersection, and the semantic distance corresponding to the vocabulary union, of the professional distance corresponding to the candidate standard question includes: If the ratio satisfies a first threshold, a first magnification coefficient is obtained, and the professional distance corresponding to the candidate standard question is determined according to the first magnification coefficient, the second number, and the semantic distance corresponding to the vocabulary union, wherein the first threshold indicates that the vocabulary intersection is the same as the vocabulary union; If the ratio satisfies a second threshold, a second magnification coefficient is obtained, and the professional distance corresponding to the candidate standard question is determined according to the second magnification coefficient, the second number, and the semantic distance corresponding to the vocabulary intersection, the second threshold indicates that the vocabulary intersection is an empty set, the first threshold is less than the second threshold, and the first magnification coefficient is greater than the second magnification coefficient; If the ratio does not meet the first threshold and the second threshold, a third amplification factor is determined based on the ratio, and the professional distance corresponding to the candidate standard question is determined based on the third amplification factor, the second number, the semantic distance corresponding to the vocabulary union, and the semantic distance corresponding to the vocabulary intersection.

9. The method according to claim 7, characterized in that: The determining, based on the first professional vocabulary set and the second professional vocabulary set, respectively, the semantic distance corresponding to the vocabulary intersection and the semantic distance corresponding to the vocabulary union comprises: Determine a word vector corresponding to each professional vocabulary in the vocabulary intersection, and determine a semantic distance corresponding to the vocabulary intersection according to the word vector corresponding to each professional vocabulary in the vocabulary intersection, wherein the word vector corresponding to each professional vocabulary in the vocabulary intersection includes at least one of a first word vector corresponding to the professional vocabulary in the first professional vocabulary set and a second word vector corresponding to the professional vocabulary in the second professional vocabulary set; Determine the target professional vocabulary in the vocabulary union that does not belong to the vocabulary intersection, and determine the semantic distance corresponding to the vocabulary union based on the word vector corresponding to the target professional vocabulary, the word vector corresponding to the target professional vocabulary is the first word vector corresponding to the target professional vocabulary in the first professional vocabulary set or the second word vector corresponding to the target professional vocabulary in the second professional vocabulary set.

10. The method according to claim 8, characterized in that The comprehensive distance between each candidate standard question and the current question is determined by the following formula: Among them, D(Q i ) represents the comprehensive distance, α is the first weight coefficient, β is the second weight coefficient, The second overall semantic word vector Q representing the candidate standard question i The component in the jth dimension, The second overall semantic word vector Q representing the current question c The component in the jth dimension, U represents the dimension of the vector, p is the norm coefficient, A is the vocabulary intersection, B is the vocabulary union, γ is the magnification coefficient, N B is the second quantity, w j is the first word vector or the second word vector corresponding to the professional vocabulary w, and They represent the j-th component of the professional vocabulary w in the corresponding first word vector and second word vector respectively.

Citation Information

Patent Citations

  • RPA and AI combined dialogue question answering method and device, equipment and storage medium

    CN111966808A

  • Chinese text matching method and device, electronic equipment and readable storage medium

    CN118503723A

  • Method and apparatus for constructing object relationship network, and electronic device

    US20230004715A1