Natural language processing method and device, equipment and storage medium
By combining knowledge base retrieval and weighted fusion probability distribution of prior knowledge in a large language model, the problem of inaccurate output in unfamiliar knowledge domains is solved, and the reliability of the output is improved.
Patent Information
- Application Number
- CN202410585882.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-11
- Publication Date
- 2025-11-18
AI Technical Summary
When faced with problems in unfamiliar knowledge domains, large language models tend to output information that is inconsistent with the facts or unverified, lacking effective supervision information.
By retrieving information related to the input statement from the existing knowledge base and combining it with the prior knowledge of the natural language processing model, the probability distribution of words at the output position is determined. Multiple probability distributions are then weighted and fused using weight parameters to determine the final output words.
It increases the accuracy of output statements, reduces the risk of output being inconsistent with facts or containing unverified information, and improves the credibility of generated answers.
Smart Images

Figure CN120973886A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular relates to a natural language processing method and device, equipment and a storage medium. BACKGROUND
[0002] With the development of large language models (LLM), the demand for implementing natural language question answering through large language models is increasing.
[0003] In related technologies, a large language model can understand the semantics of natural language, input a user input question sentence into the large language model, and predict coherent and logical output text, which is an answer sentence to the question sentence.
[0004] However, due to the training set in the training process of the large language model, the prior knowledge obtained by the large language model in the training process is limited. When the large language model faces a problem in an unfamiliar knowledge field, it will answer the problem according to the prior knowledge, and there is a high probability that it will output unverified or doubtful information inconsistent with the facts in the output sentence. SUMMARY
[0005] The present application provides a natural language processing method, device, equipment and storage medium, and the technical solution is as follows:
[0006] According to an aspect of the present application, a natural language processing method is provided, which is executed by a computer device, and the method comprises:
[0007] An input sentence is obtained, which is a character sequence with natural language semantics;
[0008] In an existing knowledge base, retrieval information is retrieved based on the input sentence, which is information associated with the semantics of at least one word in the input sentence in the existing knowledge base;
[0009] Natural language processing is performed on the input sentence to determine a first probability distribution, which is a probability distribution of a first group of candidate words at a current output position of an output sentence;
[0010] The natural language processing is performed on the input sentence and the retrieval information to determine a second probability distribution, which is a probability distribution of a second group of candidate words at the current output position;
[0011] According to the first probability distribution and the second probability distribution, an output word at the current output position is determined.
[0012] According to another aspect of the present application, there is provided a natural language processing apparatus, the apparatus comprising:
[0013] an obtaining module configured to obtain an input sentence, the input sentence being a sequence of characters having a natural language semantic;
[0014] a retrieving module configured to retrieve, based on the input sentence, retrieval information from an existing knowledge base, the retrieval information being information in the existing knowledge base that is semantically associated with at least one word in the input sentence;
[0015] a processing module configured to perform a natural language processing on the input sentence to determine a first probability distribution, the first probability distribution being a probability distribution of a first set of candidate words at a current output position of an output sentence;
[0016] the processing module is further configured to perform the natural language processing on the input sentence and the retrieval information to determine a second probability distribution, the second probability distribution being a probability distribution of a second set of candidate words at the current output position;
[0017] a determining module configured to determine an output word at the current output position according to the first probability distribution and the second probability distribution.
[0018] In an optional design of the present application, the determining module is further configured to:
[0019] obtain a first weight parameter and a second weight parameter; perform weighting on the first probability distribution based on the first weight parameter, and perform weighting on the second probability distribution based on the second weight parameter;
[0020] determine, according to the weighted first probability distribution and the weighted second probability distribution, a fusion probability value corresponding to each candidate word in the first set of candidate words and the second set of candidate words;
[0021] determine the output word at the current output position from the first set of candidate words and the second set of candidate words based on the fusion probability value.
[0022] In an optional design of the present application, a union of the first set of candidate words and the second set of candidate words comprises j candidate words, any two of the j candidate words being different from each other;
[0023] the determining module is further configured to:
[0024] subtract a first probability value in the weighted first probability distribution from a second probability value of an i-th candidate word in the weighted second probability distribution to obtain an i-th fusion probability value corresponding to the i-th candidate word.
[0025] updating i to i+1, and starting to execute the step of subtracting the first probability value in the weighted first probability distribution from the second probability value of the ith candidate word in the weighted second probability distribution to obtain an ith fusion probability value corresponding to the ith candidate word, until j fusion probability values are calculated;
[0026] wherein i is a positive integer, and j is a positive integer greater than or equal to i.
[0027] In an optional design of the present application, the determining module is further configured to:
[0028] obtain a priori knowledge corpus, the a priori knowledge corpus comprising at least one phrase in a general knowledge field;
[0029] in a case where a repetition number of the phrase in the a priori knowledge corpus and the input word in the input sentence does not exceed a quantity threshold, determining the first weight parameter as a first numerical value and the second weight parameter as a second numerical value, the first numerical value and the second numerical value being positive numbers, and the first numerical value being less than the second numerical value;
[0030] in a case where the repetition number of the phrase in the a priori knowledge corpus and the input word exceeds the quantity threshold, determining the first weight parameter as a third numerical value and the second weight parameter as a fourth numerical value, the third numerical value being a negative number and the fourth numerical value being a positive number.
[0031] In an optional design of the present application, in the case where the repetition number of the phrase in the a priori knowledge corpus and the input word does not exceed the quantity threshold, a difference between the second numerical value and the first numerical value is positively correlated with the repetition number;
[0032] in the case where the repetition number of the phrase in the a priori knowledge corpus and the input word exceeds the quantity threshold, a difference between the fourth numerical value and the third numerical value is negatively correlated with the repetition number.
[0033] In an optional design of the present application, the processing module is further configured to:
[0034] invoke a natural language model to predict the first probability distribution according to the input sentence and the existing word;
[0035] invoke the natural language model to predict the second probability distribution according to the input sentence, the retrieval information, and the existing word;
[0036] wherein the existing word is a word in the output sentence located before the current output position.
[0037] In an optional design of the present application, the determining module is further configured to:
[0038] obtain a training data set of the natural language model;
[0039] construct the obtained prior knowledge base according to feature word groups in the training data set;
[0040] wherein a document frequency of the feature word groups in the training data set is greater than a first frequency threshold, and an inverse document frequency is less than a second frequency threshold.
[0041] In an optional design of the present application, the determining module is further configured to:
[0042] perform normalization on the j fusion probability values corresponding to the j candidate words respectively, to obtain j normalized probability values, and the output word is determined based on the j normalized probability values.
[0043] In an optional design of the present application, the retrieving module is further configured to:
[0044] determine relevance information between the candidate information and the input sentence according to text content of the candidate information in the existing knowledge base and the input sentence;
[0045] determine the candidate information with relevance information exceeding a first relevance threshold as the retrieval information.
[0046] In an optional design of the present application, the retrieving module is further configured to:
[0047] encode the first hidden layer representation of the input sentence, and encode the second hidden layer representation of the text content of the candidate information in the existing knowledge base;
[0048] in a case where a similarity between the first hidden layer representation and the second hidden layer representation exceeds a first similarity threshold, determine the candidate information as the retrieval information.
[0049] In an optional design of the present application, the retrieval information includes at least two subparts.
[0050] The determining module is further configured to:
[0051] determine at least one associated subpart from the at least two subparts according to a degree of association between the at least two subparts and the input sentence respectively.
[0052] wherein the second probability distribution is determined by performing the natural language processing on the input sentence and the at least one associated subpart.
[0053] In an optional design of the present application, the determining module is further configured to perform at least one of the following:
[0054] split the search information into the at least two subparts according to a delimiter in the search information;
[0055] split the search information into the at least two subparts according to a preset string length, wherein at least one overlapping character is included between two adjacent subparts;
[0056] split the search information into the at least two subparts based on a maximum input length of the natural language model.
[0057] In an optional design of the present application, the determining module is further configured to:
[0058] calculate at least two sub-relevance information corresponding to the at least two subparts respectively according to the input sentence, wherein the sub-relevance information is used to indicate the relevance between one subpart in the search information and the input sentence;
[0059] determine the first subpart corresponding to the first sub-relevance information as the relevant subpart when the first sub-relevance information in the at least two sub-relevance information exceeds a second relevance threshold.
[0060] In an optional design of the present application, the determining module is further configured to:
[0061] obtain a first hidden layer representation of the input sentence by encoding, and obtain at least two sub-hidden layer representations corresponding to the at least two subparts by encoding;
[0062] determine the first subpart corresponding to the first sub-hidden layer representation as the relevant subpart when the similarity between the first hidden layer representation and the first sub-hidden layer representation in the at least two sub-hidden layer representations exceeds a second similarity threshold.
[0063] According to another aspect of the present application, a computer device is provided, which comprises a processor and a memory, and the memory stores at least one instruction, at least one program, a code set or an instruction set, which are loaded and executed by the processor to implement the natural language processing method according to the above aspect.
[0064] According to another aspect of the present application, a computer readable storage medium is provided, the readable storage medium storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by a processor to implement the natural language processing method according to the above aspect.
[0065] According to another aspect of the present application, a computer program product is provided, the computer program product comprising computer instructions stored in a computer readable storage medium, the computer instructions being read and executed by a processor from the computer readable storage medium to implement the natural language processing method according to the above aspect.
[0066] The technical scheme provided by the present application has at least the following beneficial effects:
[0067] By using the first probability distribution and the second probability distribution, the current output word in the output sentence is determined, and in addition to the prior knowledge of natural language processing, the retrieval information is added as the basis for determining the current output word in the output sentence. When generating the output sentence, the prior knowledge of natural language processing and the retrieval information obtained based on the input sentence are taken into account. Compared with the related art, by using the expanded retrieval information as additional supervision information of the current output word, the current output word can be determined based on more effective information, and the risk of including inconsistent or unverified suspicious information in the output sentence is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0068] In order to more clearly illustrate the technical schemes in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0069] Figure 1 is a schematic diagram of a computer system provided by an exemplary embodiment of the present application;
[0070] Figure 2 is a schematic diagram of a natural language processing method provided by an embodiment of the present application;
[0071] Figure 3 is a flowchart of a natural language processing method provided by an embodiment of the present application;
[0072] Figure 4 is a flowchart of a natural language processing method provided by an embodiment of the present application;
[0073] Figure 5is a flowchart of a natural language processing method provided by an embodiment of the present application;
[0074] Figure 6 is a flowchart of a natural language processing method provided by an embodiment of the present application;
[0075] Figure 7 is a flowchart of a natural language processing method provided by an embodiment of the present application;
[0076] Figure 8 is a flowchart of a natural language processing method provided by an embodiment of the present application;
[0077] Figure 9 is a structural block diagram of a natural language processing apparatus provided by an embodiment of the present application;
[0078] Figure 10 is a structural block diagram of a server provided by an example embodiment of the present application.
[0079] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the present application. DETAILED DESCRIPTION
[0080] For the purpose of clarity, technical solutions and advantages of the present application will be further described in detail below with reference to the accompanying drawings.
[0081] The example embodiments will be described in detail herein with reference to the accompanying drawings. The following description is presented for purposes of illustration and description, and is not intended to limit the application, as understood by persons of ordinary skill in the art. The implementation described in the following example embodiments is merely representative of the devices and methods consistent with aspects of the present application, as detailed in the appended claims.
[0082] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0083] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions. For example, the input sentence, existing knowledge base and other information involved in the present application are obtained under sufficient authorization.
[0084] It should be understood that although the terms first, second, etc. can be employed in this disclosure to describe various information, the information should not be limited to these terms. These terms are only used to distinguish one type of information from another type of information. For example, a first parameter can also be referred to as a second parameter, and similarly, a second parameter can also be referred to as a first parameter, without departing from the scope of the present disclosure. Depending on the context, the word "if' as used herein can be interpreted as "when" or "upon" or "in response to determining".
[0085] Figure 1 A schematic diagram of a computer system provided by an embodiment of the present application is shown. The computer system can be implemented as a system architecture of a natural language processing method. The computer system can include a terminal 100 and a server 200.
[0086] The terminal 100 can be an electronic device such as a mobile phone, a tablet computer, a vehicle-mounted terminal (car machine), a wearable device, a PC (Personal Computer), an access control device, a self-service terminal, etc. A client running a target application program can be installed in the terminal 100, which can be a game application program or other application program providing a natural language processing function, which is not limited by the present application. In addition, the present application does not limit the form of the target application program, including but not limited to App (Application) installed in the terminal 100, mini program, etc., which can also be in the form of a web page.
[0087] The server 200 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The server 200 can be a background server of the above-mentioned target application program, used to provide background services for the client of the target application program.
[0088] The natural language processing method provided by the embodiment of the present application, the execution subject of each step can be a computer device, which refers to an electronic device with data calculation, processing and storage capabilities. For example, the terminal 100 and the server 200 can be computer devices. Figure 1Taking the computer system shown as an example, the natural language processing method can be executed by the terminal 100 (such as by the client of the target application installed and running in the terminal 100), or by the server 200, or by the interaction and cooperation between the terminal 100 and the server 200. This application does not limit this.
[0089] Furthermore, the technical solution of this application can be combined with blockchain technology. For example, some data involved in the natural language processing method disclosed in this application (such as 3D medical images, predicted instance segmentation results, etc.) can be stored on the blockchain. Terminal 100 and server 200 can communicate via a network, such as a wired or wireless network.
[0090] Figure 2 A schematic diagram of a natural language processing method provided in an exemplary embodiment of this application is shown.
[0091] The input statement 302 is obtained. The input statement 302 is a character sequence with natural language semantics. In this embodiment, the input statement 302 is a medical question as an example. The input statement 302 has the semantics of raising a medical question.
[0092] The encoder 352 is invoked to encode the first hidden layer representation 303 of the input statement 302, and the second hidden layer representation 313 of the text content of the candidate information 312 in the existing knowledge base 310 is also encoded. For example, the existing knowledge base 310 is a knowledge base of medical knowledge, such as at least one of medical books, medical papers, Internet medical websites, medical institution business databases, etc.
[0093] If the similarity 315 between the first hidden layer representation 303 and the second hidden layer representation 313 exceeds a first similarity threshold, the candidate information 312 is determined as the retrieval information 318. For example, the similarity 315 between the first hidden layer representation 303 and the second hidden layer representation 313 is the cosine similarity between the two hidden layer representations.
[0094] The statements in the search information are further filtered; for example, the search information 318 includes a first sub-part 318a to an Nth sub-part 318b, for a total of N sub-parts, where N is a positive integer greater than 1.
[0095] It should be noted that, considering the maximum input length of the natural language model 354, the retrieved information 318 needs further filtering. Similar to the encoding of candidate information above, the encoding yields N hidden layer representations corresponding to the aforementioned N sub-parts. Among the N sub-parts, the associated sub-part 319 is obtained through filtering. The similarity between the hidden layer representation corresponding to the associated sub-part 319 and the first hidden layer representation 303 exceeds the second similarity threshold.
[0096] The input sentence 302 is input into the natural language model 354 to predict a first probability distribution 332 of a first set of candidate words; the input sentence 302 and the associated subpart 319 are input into the natural language model 354 to predict a second probability distribution 334 of a second set of candidate words. The first set of candidate words includes at least one candidate word, and the second set of candidate words includes at least one candidate word; the number of candidate words in the above two sets of candidate words is usually independent of each other, and usually has no association.
[0097] For a candidate word, a fusion probability value 340 corresponding to the candidate word is constructed, wherein the fusion probability value is a difference value of a first value minus a second value, wherein the first value is a product of the second weight parameter 335 and a probability value corresponding to the candidate word in the second probability distribution 334; and the second value is a product of the first weight parameter 333 and a probability value corresponding to the candidate word in the first probability distribution 332.
[0098] For example, the first set of candidate words and the second set of candidate words include j candidate words, according to the above introduction of the fusion probability value 340, similarly, j fusion probability values corresponding to the j candidate words are constructed. The j fusion probability values corresponding to the j candidate words are normalized to obtain j normalized probability values 342.
[0099] The current output word 345 is determined according to the j normalized probability values 342 among the first set of candidate words and the second set of candidate words; the output sentence is constructed based on the current output word 345, and the output sentence is the answer sentence of the input sentence 302.
[0100] It can be seen that the current output word 345 is determined by the first probability distribution 332 and the second probability distribution 334, in addition to the prior knowledge possessed by the natural language model 354, the associated subpart 319 is added as a basis for determining the current output word in the output sentence; the prior knowledge possessed by the natural language model 354 and the retrieval information 318 based on the input sentence are considered when generating the output sentence, and the associated subpart 319 is further screened out from the retrieval information 318, so as to avoid the adverse effects of the part of characters in the retrieval information 318 irrelevant to the input sentence 302 in the process of predicting the second probability distribution 334 based on the retrieval information 318 and the input sentence 302 by the natural language model 354. By setting the first weight parameter 333 and the second weight parameter 335, the fusion probability value 340 is determined by subtracting the weighted first probability value from the weighted second probability value, which realizes flexible adjustment of the dependence degree of the current output word 345 on the prior knowledge in the natural language model 354 or the retrieval information 318.
[0101] Next, the natural language processing method will be introduced through the following examples.
[0102] Figure 3 A flowchart of the natural language processing method provided by an example embodiment of the present application is shown. The method can be executed by a computer device. The method comprises:
[0103] Step 510: obtaining an input sentence;
[0104] The input sentence is a character sequence with natural language semantics. The input sentence is a sentence input by a human-computer interaction operation; for example, the sentence type of the input sentence can be a question or a statement. In the case of a question, the input sentence indicates a semantic of asking at least one question, and the embodiment is used to obtain an answer to the at least one question. In the case of a statement, the embodiment is used to obtain an explanatory answer to the meaning of the words in the input sentence.
[0105] For example, the input sentence is a sentence input by a human-computer interaction operation.
[0106] Step 520: retrieving information based on the input sentence in an existing knowledge base;
[0107] For example, the existing knowledge base usually includes at least one of the following, but is not limited to: at least one of a thesis, a book, and a website;
[0108] For example, the existing knowledge base is a publicly available knowledge base, and the retrieval of the existing knowledge base is subject to the consent of the authorized party and complies with relevant laws, regulations and standards in the relevant region. The existing knowledge base can be a knowledge base in a knowledge field, such as a medical knowledge base or a legal knowledge base; or a general knowledge base including extensive basic knowledge and basic skills. In an example, the existing knowledge base in the medical knowledge field is a medical book, a medical thesis, a term of a medical knowledge field name in a search engine, a data set published by a medical open source platform (such as a biomedical language understanding evaluation data set (Biomedical Language Understanding Evaluation)), etc.
[0109] For example, the retrieval information is obtained by retrieving the input sentence in the existing knowledge base; the retrieval information is information in the existing knowledge base that is semantically associated with at least one word in the input sentence; for example, the retrieval information includes at least one of the following: information including at least one word in the input sentence, information including a translation of at least one word in the input sentence, and information having the same / similar semantics as at least one word in the input sentence.
[0110] As introduced above, the input sentence can be a question sentence, and the input sentence includes an interrogative word and an interrogated word. For example, the input question sentence is "What is the survival rate of pancreatic cancer?". The interrogative word is "What is the survival rate of pancreatic cancer?", which has the semantic meaning of asking. The interrogated word is "pancreatic cancer", and the question sentence asks about the survival rate of pancreatic cancer. In the case of the input sentence being a question sentence, the retrieved information is information in the existing knowledge base that is semantically associated with the interrogated word, such as at least one of information including the interrogated word, information including a translation of the interrogated word, and information having the same / similar semantic meaning as the interrogated word. As introduced above, the input sentence can be a statement sentence, and the retrieved information is information in the existing knowledge base that is semantically associated with at least one word in the statement sentence.
[0111] For example, the retrieval of the input sentence in the existing knowledge base can be direct retrieval of a word in the input sentence, or can be encoding of the input sentence to obtain encoded information including semantic information of the input sentence, and then retrieving the encoded information. The retrieval method in the existing knowledge base is not limited in the present application.
[0112] Step 530: performing natural language processing on the input sentence to determine a first probability distribution;
[0113] The first probability distribution is the processing result of performing natural language processing on the input sentence. For example, the natural language processing is used to obtain a corresponding answer according to the input sentence, and the natural language processing is usually implemented based on an artificial neural network model with a question and answer function, also known as a natural language model. The natural language model has the ability to output a natural language sentence. For example, the first probability distribution is predicted during the process of performing natural language processing on the input sentence.
[0114] For example, the types of natural language models include at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM), and a generative pre-trained transformer (GPT). The natural language processing can also be implemented as a mapping relationship between the input sentence and the output sentence, and the first probability distribution is obtained based on the preset mapping relationship. The manner of obtaining the probability distribution by natural language processing is not limited in the present application.
[0115] The first probability distribution is, for example, a probability distribution of the first set of candidate words at the current output position of the output sentence. The current output position is usually a position (also referred to as a slot) corresponding to a word in a natural language sentence. The natural language sentence is the answer to the input sentence, i.e., the output sentence.
[0116] The first set of candidate words is a set of candidate words predicted by performing natural language processing on the input sentence and filled in at the current output position. The first set of candidate words includes, for example, at least one candidate word, and the first probability distribution is a set of at least one probability value. In the first probability distribution, each of the at least one probability value corresponds to one of the at least one candidate word in the first set of candidate words.
[0117] The first probability value in the at least one probability value is, for example, used to indicate a probability that the output sentence obtained by performing natural language processing on the input sentence includes the first candidate word at the current output position (also referred to as a confidence level of filling the candidate word in the output sentence). The first candidate word is any one of the candidate words in the first set of candidate words, and the first probability value is the probability value corresponding to the first candidate word in the at least one probability value.
[0118] In one example, the input sentence is “What is the survival rate of pancreatic cancer?”. One of the probability values included in the first probability distribution is 0.8, and the corresponding candidate word is “9 percent”, indicating that the probability that the output sentence includes the word “9 percent” is 0.8. Another of the probability values included in the first probability distribution is 0.2, and the corresponding candidate word is “2 percent”, indicating that the probability that the output sentence includes the word “2 percent” is 0.2.
[0119] In some examples, the first probability distribution is normalized, and the sum of the probability values corresponding to each of the candidate words in the first set of candidate words is 1.
[0120] Step 540: performing natural language processing on the input sentence and the retrieved information to determine a second probability distribution;
[0121] The second probability distribution is, for example, a result of performing natural language processing on the input sentence and the retrieved information. The natural language processing is, for example, used to obtain a corresponding answer based on the input sentence and the retrieved information, and the specific manner of the natural language processing is as described above. The second probability distribution is, for example, predicted during the process of performing natural language processing on the input sentence and the retrieved information.
[0122] The second probability distribution is, for example, a probability distribution of the second set of candidate words at the current output position. The second set of candidate words is a set of candidate words predicted by performing natural language processing on the input sentence and the retrieved information and filled in at the current output position.
[0123] Exemplarily, the second set of candidate words includes at least one candidate word, and the second probability distribution is a set of at least one probability value; in the second probability distribution, the at least one probability value and the at least one candidate word in the second set of candidate words are in one-to-one correspondence. The intersection of the first set of candidate words and the second set of candidate words can be empty or not empty, and no limitation is made in this regard.
[0124] Exemplarily, a second probability value in the at least one probability value is used to indicate a probability that, in the case of performing natural language processing on the input sentence, the output sentence obtained includes the second candidate word at the current output position; the second candidate word is any one candidate word in the second set of candidate words, and the second probability value is a probability value corresponding to the first candidate word in the at least one probability value.
[0125] It should be noted that steps 530 and 540 can be executed simultaneously or sequentially, and the application does not limit the execution timing of the two steps.
[0126] Step 550: determining an output word at the current output position according to the first probability distribution and the second probability distribution.
[0127] The output word belongs to at least one of the first set of candidate words and the second set of candidate words.
[0128] Exemplarily, the output word is sampled from the first set of candidate words and the second set of candidate words; the probability of sampling the output word is determined according to the first probability distribution and the second probability distribution. Exemplarily, the first probability distribution and the second probability distribution can be subjected to four arithmetic operations, calculation of extreme values, etc., to obtain a probability value of including a candidate word in the output sentence, and then the output word is sampled from the first set of candidate words and the second set of candidate words according to the probability value.
[0129] To sum up, the method provided by the embodiment determines the current output word in the output sentence through the first probability distribution and the second probability distribution, and increases the retrieval information as the basis for determining the current output word in the output sentence in addition to the prior knowledge of natural language processing; both the prior knowledge of natural language processing and the retrieval information obtained based on the input sentence are taken into account when generating the output sentence; compared with the related art, the expanded retrieval information is used as additional supervision information of the current output word, which can determine the current output word based on more effective information and reduce the risk of including information inconsistent with facts or unverified and doubtful information in the output sentence.
[0130] In some optional examples, the first probability distribution and the second probability distribution need to be weighted, and the output word is determined according to the weighted first probability distribution and the weighted second probability distribution; next, the weighting of the first probability distribution and the second probability distribution is introduced:
[0131] Figure 4 A flowchart of a natural language processing method provided by an example embodiment of the present application is shown. The method can be executed by a computer device. That is, in Figure 3 In the shown embodiment, step 550 can be implemented as step 552, step 554, and step 556:
[0132] Step 552: obtaining a first weight parameter and a second weight parameter; performing weighting on the first probability distribution based on the first weight parameter, and performing weighting on the second probability distribution based on the second weight parameter;
[0133] For example, the first weight parameter and the second weight parameter can be a pre-set fixed value, can be set based on human-computer interaction operation, or can be a weight parameter determined according to an input sentence or other information, which will be described separately in the following embodiments, and the present application does not limit the setting mode of the weight parameter.
[0134] The values of the first weight parameter and the second weight parameter are usually different, so as to set different weights for the first probability distribution and the second probability distribution. For example, the first weight parameter is used to perform weighting on the first probability distribution; the second weight parameter is used to perform weighting on the second probability distribution. In an example, the first weight parameter is used to perform weighting on the first probability distribution in a multiplication manner; similarly, the second weight parameter is used to perform weighting on the second probability distribution in a multiplication manner. Taking the weighting of the first probability distribution as an example, the product of the first weight parameter and each probability value in the first probability distribution is calculated to obtain the weighted first probability distribution.
[0135] In an example, the value of the first weight parameter is 1, and the second weight parameter is different from the first weight parameter; that is, only the second weight parameter is used to perform weighting on the second probability distribution, and the first probability distribution remains unchanged. Similarly, the value of the second weight parameter is 1, and the first weight parameter is different from the second weight parameter; that is, only the first weight parameter is used to perform weighting on the first probability distribution, and the second probability distribution remains unchanged.
[0136] In some examples, the values of the first weight parameter and the second weight parameter are positive numbers.
[0137] Exemplarily, the first weight parameter and the second weight parameter can be independent of each other, for example, the first weight parameter and the second weight parameter are respectively set through human-computer interaction operation. The first weight parameter and the second weight parameter can also be associated with each other, and according to any one of the obtained first weight parameter and the second weight parameter, the other weight parameter can be determined.
[0138] Step 554: determining a fusion probability value corresponding to each candidate word in the first group of candidate words and the second group of candidate words according to the weighted first probability distribution and the weighted second probability distribution;
[0139] Taking the i th candidate word in the first group of candidate words and the second group of candidate words as an example, the i th candidate word is any one of the candidate words in the first group of candidate words and the second group of candidate words. The fusion probability value corresponding to the i th candidate word is determined according to the probability value of the i th candidate word in the first probability distribution and the probability value of the i th candidate word in the second probability distribution.
[0140] By setting different weights for the first probability distribution and the second probability distribution in step 552, it is realized that the first probability distribution and the second probability distribution have different importance when determining the fusion probability value. By adjusting the first weight parameter and the second weight parameter, it is possible to realize personalized adjustment of the information source of the word dependence in the output sentence of the answer input sentence, for example, the output sentence is more dependent on the retrieved information obtained through retrieval or more dependent on the prior knowledge obtained by the natural language model in the training stage.
[0141] Step 556: determining an output word at a current output position in the first group of candidate words and the second group of candidate words based on the fusion probability value;
[0142] Exemplarily, the output word is sampled from the first group of candidate words and the second group of candidate words. The fusion probability value is used to indicate the probability of determining the candidate word corresponding to the fusion probability value as the output word in the first group of candidate words and the second group of candidate words.
[0143] Exemplarily, the output word is located at the current output position of the output sentence, and the output sentence is an answer sentence to the input sentence. In one example, the output word is filled into the current output position in the output sentence to construct the output sentence.
[0144] To sum up, the method provided in the embodiment considers both the prior knowledge of natural language processing and the retrieval information retrieved based on the input sentence when generating the output sentence; by setting the weights of the first probability distribution and the second probability distribution, the dependence of the output sentence on the prior knowledge or the retrieval information can be flexibly controlled; the first weight parameter indicates the dependence of the current output word on the prior knowledge, and the second weight parameter indicates the dependence of the current output word on the retrieval information; compared with the related art, by expanding the retrieval information as additional supervision information of the current output word and setting the second weight parameter to indicate the dependence of the current output word on the retrieval information, the current output word can be determined based on more effective information, and the risk of including suspicious information that does not conform to the facts or has not been verified in the output sentence is reduced.
[0145] Next, the weighting process of the probability distribution is introduced.
[0146] In an optional implementation manner, the step 554 in the above can be implemented as the following sub-steps.
[0147] · the second probability value of the i th candidate word in the second probability distribution after weighting is subtracted from the first probability value in the first probability distribution after weighting to obtain an i th fusion probability value corresponding to the i th candidate word;
[0148] In the embodiment, the union set of the first group of candidate words and the second group of candidate words includes j candidate words, and any two candidate words in the j candidate words are not repeated.
[0149] For example, if the i th candidate word is not included in the first group of candidate words, the first probability value of the i th candidate word in the first probability distribution is 0; if the i th candidate word is not included in the second group of candidate words, the second probability value of the i th candidate word in the second probability distribution is 0. i is a positive integer, and j is a positive integer greater than or equal to i.
[0150] For example, since the second probability value of the second probability distribution is obtained by inputting the input sentence and the retrieval information into the natural language model, the natural language model utilizes both the retrieval information retrieved and the prior knowledge obtained in the training phase of the natural language model when generating the second probability distribution.
[0151] By subtracting the second probability value from the first probability value, the influence of the prior knowledge on the output sentence can be eliminated, and the degree of eliminating the influence of the prior knowledge in the second probability distribution can be controlled by adjusting the first weight parameter and the second weight parameter.
[0152] In one example, the input sentence is: "What is the survival rate of pancreatic cancer?". The first probability distribution is predicted by the natural language model based on the input sentence, and the first probability distribution indicates that the probability of the output sentence including the word "9 percent" is 0.8, and the probability of the output sentence including the word "2 percent" is 0.2. The second probability distribution is predicted by the natural language model based on the input sentence and the search information, and the second probability distribution indicates that the probability of the output sentence including the candidate word "9 percent" is 0.6, and the probability of the output sentence including the candidate word "2 percent" is 0.4.
[0153] wherein the first weighting parameter is 2, the second weighting parameter is 3, the first probability distribution and the second probability distribution are weighted respectively; the second probability value of the candidate word in the weighted second probability distribution is subtracted by the first probability value in the weighted first probability distribution to obtain a fusion probability value corresponding to the candidate word; that is, the fusion probability value of the candidate word "9 percent" is 3*0.6-2*0.8=0.2, and the fusion probability value of the candidate word "2 percent" is 3*0.4-2*0.2=0.8.
[0154] By comparing the first probability distribution and the second probability distribution, it can be seen that in the presence of search information, the probability of the output sentence including the candidate word "9 percent" decreases; by weighting the probability distribution and subtracting the second probability value from the first probability value, the influence of prior knowledge on the prediction of the natural language model is further eliminated, so that the output sentence is more dependent on the search information, and the probability of the output sentence including the candidate word "9 percent" is further reduced.
[0155] • updating i to i+1, and starting to execute the above step of subtracting the first probability value in the weighted first probability distribution from the second probability value of the i th candidate word in the weighted second probability distribution to obtain the i th fusion probability value corresponding to the i th candidate word, until j fusion probability values are calculated;
[0156] By re-entering the above calculation step of the fusion probability value, j fusion probability values are obtained, and the j fusion probability values and the j candidate words correspond one by one. For example, the j fusion probability values are used to indicate the probability that the j corresponding candidate words are determined as output words.
[0157] In a further optional implementation manner, the natural language processing method provided by the present application further includes the following steps based on the above sub-steps:
[0158] • performing normalization on the j fusion probability values corresponding to the j candidate words to obtain j normalized probability values;
[0159] Exemplarily, the manner of performing normalization on the j fusion probability values includes, but is not limited to, at least one of minimum-maximum normalization, Z-Score standardization, logarithmic normalization, and exponential normalization. Exemplarily, after the normalization, the j normalized probability values all belong to a first value range, for example, the first value range is less than 1 and greater than 0.
[0160] In this embodiment, the output word is determined based on the j normalized probability values. The normalized probability value is used to indicate the probability of the corresponding candidate word being determined as the output word. It should be noted that the present application does not require the sum of the j normalized probability values, or the sum of the j fusion probability values, to be a fixed value (such as 1).
[0161] In summary, the method provided in this embodiment takes into account both the prior knowledge of natural language processing and the retrieval information retrieved based on the input sentence when generating the output sentence; the second probability value weighted by the second probability value is predicted based on the prior knowledge and the retrieval information; the first probability value weighted by the first probability value is predicted based only on the prior knowledge;
[0162] Based on the weighted second probability value, by setting the first weight parameter and the second weight parameter, in the case that the second weight coefficient is a positive number, the weighted second probability value is subtracted from the weighted first probability value. Since the first probability value is predicted based only on the prior knowledge, it is realized that on the basis of the second probability value predicted by simultaneously relying on the prior knowledge and the retrieval information, the influence caused by the prediction result relying only on the prior knowledge is eliminated, and the fusion probability value can be determined in a manner of more relying on the retrieval information and less relying on the prior knowledge.
[0163] In the case that the second weight coefficient is a negative number, the weighted second probability value is subtracted from the weighted first probability value, which realizes that on the basis of the second probability value predicted by simultaneously relying on the prior knowledge and the retrieval information, the degree of dependence on the prior knowledge is increased, and the fusion probability value can be determined in a manner of more relying on the prior knowledge and less relying on the retrieval information.
[0164] In the above, it is realized that the degree of dependence of the current output word on the prior knowledge or the retrieval information is flexibly adjusted.
[0165] Figure 5 A schematic diagram of a natural language processing method provided by an example embodiment of the present application is shown.
[0166] In an example of the present application, the input sentence 302 is taken as an example: "What is the survival rate of pancreatic cancer?" Since the input sentence 302 involves the medical knowledge field; the corresponding existing knowledge base is a knowledge base of medical knowledge, including at least one of medical books, medical papers, Internet medical websites, medical institution business databases, etc.
[0167] In the existing knowledge base, the retrieval information 318 obtained based on the input sentence 302 is: "According to the statistics of the World Health Organization, about 56,000 people die of pancreatic cancer every year, and the five-year survival rate of pancreatic cancer is about 2%." It should be noted that the specific manner of retrieving the retrieval information 318 in the existing knowledge base will be shown separately in the embodiments below.
[0168] The embodiment takes a natural language model 354 specifically implemented as a large language model (LLM) as an example for illustration. The natural language model 354 is called to predict the first probability distribution 332 based on the input sentence 302. The natural language model 354 is called to predict the second probability distribution 334 based on the input sentence 302 and the retrieval information 318.
[0169] As introduced above, the first probability distribution 332 indicates that the probability of including the word "9%" in the output sentence is 0.8, and the probability of including the word "2%" in the output sentence is 0.2. The second probability distribution is predicted by the natural language model based on the input sentence and the retrieval information. The second probability distribution 334 indicates that the probability of including the candidate word "9%" in the output sentence is 0.6, and the probability of including the candidate word "2%" in the output sentence is 0.4.
[0170] The first weight parameter and the second weight parameter are obtained, and the first probability distribution and the second probability distribution are respectively weighted based on the first weight parameter and the second weight parameter. In the embodiment, the first weight parameter and the second weight parameter can also be associated with each other; that is, the second weight parameter is determined according to the first weight parameter. For example, the difference between the first weight parameter and the second weight parameter is a first value, and the first value is a positive integer. In one example, the first value is 1.
[0171] In the embodiment, the fusion probability value 340 is determined according to the weighted first probability distribution and the weighted second probability distribution. For example, the ith fusion probability value corresponding to the ith candidate word is obtained by subtracting the first probability value in the weighted first probability distribution from the second probability value of the ith candidate word in the weighted second probability distribution. In the case where the first value is 1, the implementation manner of retaining only the first distribution probability or the second distribution probability to determine the output word can be retained by setting the weight parameter.
[0172] In an optional implementation manner, the natural language processing in various embodiments of the present application is implemented by calling a natural language model.
[0173] The first probability distribution is predicted by calling the natural language model according to the input sentence and the existing words;
[0174] The second probability distribution is predicted by calling the natural language model according to the input sentence, the search information and the existing words;
[0175] The existing words are the words in the output sentence before the current output position. As introduced above, the current output position is the position of any one of the first group of candidate words and the second group of candidate words in the output sentence.
[0176] As introduced above, the types of the natural language model include, but are not limited to, at least one of a convolutional neural network, a recurrent neural network, a long short-term memory network and a generative pre-training conversion network. It should be noted that the natural language model called to predict the first probability distribution and the second probability distribution is the same natural language model, or can be implemented as two natural language models with the same parameters and the same structure to predict the first probability distribution and the second probability distribution in parallel.
[0177] In one example, the current output position is the t-th position in the output sentence (t is a positive integer greater than 1), that is, the output word is the word located at the t-th position in the output sentence.
[0178] The first probability distribution and the second probability distribution are predicted by calling the natural language model, as introduced above.
[0179] The first probability distribution is denoted as p θ (Y t |Q,Y <t ), which is output information predicted by calling the natural language model according to the input sentence (denoted as Q) and the words before the t-th position in the output sentence (denoted as Y <t ). θ (Y t |Q,Y <t ) is a probability distribution predicted by the natural language model in the presence of Q and Y <t .
[0180] The second probability distribution is denoted as p θ (Y t |C,Q,Y <t ), which is output information predicted by calling the natural language model according to the search information (denoted as C), the input sentence (denoted as Q) and the words before the t-th position in the output sentence (denoted as Y <t ). θ (Y t |C,Q,Y <t ) is a probability distribution predicted by the natural language model in the presence of C, Q and Y<t the probability distribution predicted by the natural language model.
[0181] It should be noted that only the determination process of the word at the tth position after the first position in the output sentence by the natural language model is introduced above. The determination of the word at the first position is introduced as follows:
[0182] For the word at the first position in the output sentence, since there is no existing word when predicting the word at the first position, the following is performed: the natural language model is called to predict a first probability distribution at the first position according to the input sentence; and the natural language model is called to predict a second probability distribution at the first position according to the input sentence and the search information.
[0183] And for the word at the t+u position in the output sentence, the natural language model is called to predict a first probability distribution at the t+u position according to the input sentence and the existing words before the t+u position; the natural language model is called to predict a second probability distribution at the t+u position according to the input sentence, the search information and the existing words before the t+u position;
[0184] When the output word at the t+u position determined according to the first probability distribution at the t+u position and the second probability distribution at the t+u position is an end symbol, the natural language model is no longer called, that is, the prediction of the output sentence is completed, and the output sentence is constructed according to the output words at a total of t+u positions predicted by the natural language model.
[0185] Further, for the output word at the tth position in the output sentence, the probability distribution is:
[0186] Y t ~ softmax [(1 + a) * p θ (Y t | C, Q, Y <t ) - a * p θ (Y t | Q, Y <t ] ;
[0187] Y t is the probability distribution of the output word at the tth position, and softmax is a normalization function. The first weight parameter is denoted as a, and the second weight parameter is denoted as 1+a; the difference between the second weight parameter and the first weight parameter is 1.
[0188] In one example, the value of a in the first weight parameter is 0, and the corresponding second weight parameter is 1+a = 1; the fusion probability value is determined according to the second probability distribution weighted only according to the second weight parameter. In another example, the value of a in the first weight parameter is -1, and the corresponding second weight parameter is 1+a = 0; the fusion probability value is determined according to the first probability distribution weighted only according to the first weight parameter.
[0189] The manner of performing natural language processing by calling a natural language model introduced above does not need to perform supplementary training on the natural language model, nor does it need to fine-tune the network parameters of the natural language model; that is, it can take into account both the prior knowledge possessed by the natural language model and the retrieval information in the medical knowledge field retrieved based on the input sentence when generating the output sentence; wherein the first probability distribution is predicted only relying on prior knowledge, and the second probability distribution is predicted relying on prior knowledge and retrieved retrieval information. The weight coefficient can be used to individually adjust the degree of dependence on the above two ways.
[0190] Next, the acquisition process of the first weight parameter and the second weight parameter is introduced:
[0191] As introduced above, the second probability value of the ith candidate word in the weighted second probability distribution is subtracted from the first probability value in the weighted first probability distribution to obtain the ith fusion probability value corresponding to the ith candidate word; on this basis, Figure 4 The step 552 in the embodiment shown can be implemented as the following three sub-steps:
[0192] · Acquire the prior knowledge vocabulary;
[0193] The prior knowledge vocabulary includes at least one word group in the general knowledge field; for example, general knowledge includes basic knowledge in multiple knowledge fields, such as but not limited to basic knowledge in the fields of mathematics, basic science, literature, and art.
[0194] Next, the acquisition of the prior knowledge library is introduced.
[0195] As introduced above, the natural language processing is introduced by calling a natural language model. In the case of predicting the first probability distribution and the second probability distribution by calling the natural language model, the above natural language model is an artificial neural network model obtained by training.
[0196] For example, the artificial neural network model is trained based on the training data set, the semantics of the natural language sentence in the training data set is converted into the model parameters of the artificial neural network model, and the natural language model is obtained. That is, the prior knowledge possessed by the natural language model is obtained from the training data set.
[0197] Correspondingly, the acquisition of the prior knowledge library in the above can be implemented as:
[0198] Acquire a training data set of a natural language model;
[0199] According to the feature word group in the training data set, the obtained prior knowledge library is constructed;
[0200] Illustratively, the training data set is as described above, which is a data set used for training the natural language model. The feature word group is a word group in the training data set, which is screened in the training data set. Further, the feature word group belongs to the first data in the training data set, and the first data is the data input to the natural language model.
[0201] Illustratively, the feature word group has a document frequency greater than a first frequency threshold and an inverse document frequency less than a second frequency threshold in the training data set. Illustratively, the feature word group has a document frequency greater than a first frequency threshold in the training data set, which ensures that the semantic information of the feature word group is converted into prior knowledge through the training of the natural language model, that is, the natural language model has prior knowledge covering the feature word group and can understand the semantics of the feature word group. The inverse document frequency of the feature word group in the training data set is less than the second frequency threshold, which avoids determining the conjunction words or mood words such as "huh", "why", and "of" in the training data as feature word groups.
[0202] · In the case where the number of repetitions of the word group in the prior knowledge library and the input word in the input sentence does not exceed the number threshold, the first weight parameter is determined as the first value, and the second weight parameter is determined as the second value;
[0203] In the case where the number of repetitions of the word group in the prior knowledge library and the input word in the input sentence does not exceed the number threshold, the prior knowledge possessed in the natural language processing and the knowledge field to which the input sentence belongs are quite different.
[0204] The first weight parameter is used to weight the first probability distribution, and the second weight parameter is used to weight the second probability distribution. According to the above description, it can be seen that the process of obtaining the second probability distribution through natural language processing simultaneously utilizes the prior knowledge possessed by natural language processing and the semantic knowledge in the retrieval information, which is retrieved according to the input sentence and belongs to the same knowledge field as the input sentence. While the process of obtaining the first probability distribution through natural language processing only utilizes the prior knowledge possessed by natural language processing.
[0205] In the embodiment, the first value and the second value are both positive numbers, and the first value is less than the second value; the influence of the prior knowledge is eliminated in the process of obtaining the fusion probability value by subtracting the first probability distribution weighted by the first value from the second probability distribution weighted by the second value on the basis of the second probability distribution weighted by the second value; and the output word is more dependent on the information provided by the search information and belonging to the same knowledge field as the input sentence.
[0206] In one example, the input words in the input sentence include words in the medical knowledge field such as "pancreatic cancer" and "survival rate". The word groups in the general knowledge field in the prior knowledge base only involve the basic knowledge of each knowledge field and do not involve the knowledge of specific diseases in the medical knowledge field.
[0207] Correspondingly, the first value and the second value are both positive numbers, and the first value is less than the second value; in combination with the introduction in the above, the ith fusion probability value is obtained by subtracting the first probability value in the first probability distribution weighted by the first value from the second probability value in the second probability distribution weighted by the second value.
[0208] The dependence degree of the ith fusion probability value on the search information obtained by searching is the second value, and the dependence degree on the prior knowledge is the second value minus the first value. The first value and the second value are both positive numbers, and the first value is less than the second value, which can reduce the dependence degree on the prior knowledge, eliminate the influence of the prior knowledge in the process of obtaining the fusion probability value, and make the output word more dependent on the information provided by the search information and belonging to the same knowledge field as the input sentence, thereby avoiding the deviation of the output sentence caused by the lack of prior knowledge.
[0209] Further, in the case where the number of repetitions of the word groups in the prior knowledge base and the input words in the input sentence does not exceed the number threshold, as the number of repetitions increases, there are more repetitions of the words in the prior knowledge base and the input words, and the prior knowledge has more reference value.
[0210] Correspondingly, the difference between the second value and the first value is positively correlated with the number of repetitions. That is, as the number of repetitions increases, the difference between the second value and the first value increases. The dependence degree of the fusion probability value on the information provided by the search information and belonging to the same knowledge field as the input sentence is reduced; the influence of the prior knowledge can be eliminated to a smaller extent in the process of obtaining the fusion probability value, and more probability distributions based on the prior knowledge are retained.
[0211] In the case where the number of repetitions of the word groups in the prior knowledge base and the input words exceeds the number threshold, the first weight parameter is determined as a third value, and the second weight parameter is determined as a fourth value;
[0212] In the embodiment, the third value is negative, and the fourth value is positive. On the basis of the second probability distribution weighted by the fourth value, the first probability distribution weighted by the third value is subtracted, and the dependence on the prior knowledge is increased by setting the third value as negative in the process of obtaining the fusion probability value; when the output word is obtained, the information in the general knowledge field provided by the prior knowledge is more relied on.
[0213] In one example, the input words in the input sentence include "cosine", "Pythagorean theorem", and other words in the mathematical knowledge field. This is the basic knowledge in the mathematical knowledge field (such as knowledge not involving higher education stage or knowledge frequently searched in a public search engine), and the word group in the general knowledge field in the prior knowledge base only involves the basic knowledge in each knowledge field.
[0214] Correspondingly, the third value is negative, and the fourth value is positive. In combination with the introduction in the above, the ith fusion probability value is the second probability value in the second probability distribution weighted by the fourth value, minus the first probability value in the first probability distribution weighted by the third value.
[0215] The dependence of the ith fusion probability value on the retrieved information is the fourth value, and the dependence on the prior knowledge is the fourth value minus the third value. The third value is negative, and the fourth value is positive, which ensures that the dependence on the retrieved information is less than the dependence on the prior knowledge; the influence of the retrieved information is reduced in the process of obtaining the fusion probability value; when the output word is obtained, the prior knowledge of the natural language model is more relied on. The limited number of the retrieved information retrieved causes the deviation of the output sentence to be avoided.
[0216] Further, in the case where the repetition number of the word group in the prior knowledge base and the input word exceeds the number threshold, as the repetition number increases, there are more repeated words in the prior knowledge base and the input word, and the prior knowledge has more reference value.
[0217] Correspondingly, the first ratio is negatively correlated with the repetition number, and the first ratio is the ratio between the fourth value and the first difference value, and the first difference value is the difference between the fourth value and the third value. That is, as the repetition number increases, the dependence of the fusion probability value on the prior knowledge is increased. Only the prior knowledge is used in the first probability distribution, and the prior knowledge and the retrieved information are used in the second probability distribution. The weight (fourth value) of the second probability distribution needs to be reduced to increase the dependence of the fusion probability value on the prior knowledge; correspondingly, the difference between the fourth value and the third value is also reduced, and the difference between the fourth value and the third value is negatively correlated with the repetition number, and more probability distributions based on the prior knowledge are retained.
[0218] In summary, the method provided in the embodiment takes into account both the prior knowledge of natural language processing and the search information retrieved based on the input sentence when generating the output sentence; determines the degree of dependence on the prior knowledge and the search information according to the degree of association between the prior knowledge and the input sentence; for example, in the case where the knowledge field of the input sentence is similar to the prior knowledge possessed in natural language processing, the degree of dependence of the current output word on the prior knowledge is increased by setting the first value and the second value; and in the case where the knowledge field of the input sentence is significantly different from the prior knowledge possessed in natural language processing, the degree of dependence of the current output word on the search information is increased by setting the third value and the fourth value; and the risk of including unverified or doubtful information in the output sentence is reduced.
[0219] Next, the search information is further introduced:
[0220] Figure 6 A flowchart of a natural language processing method provided in an example embodiment of the present application is shown. The method can be executed by a computer device. That is, in the example embodiment shown in FIG. 5, the method can be executed by the computer device 100 shown in FIG. 1. Figure 3 In the example embodiment shown, the step 520 can be implemented as a step 522 and a step 524.
[0221] The step 522: determining the relevance information between the candidate information and the input sentence according to the text content of the candidate information in the existing knowledge base and the input sentence.
[0222] For example, the number of candidate information included in the existing knowledge base can be one or more, and the relevance information between the candidate and the input sentence includes at least one of the document frequency (DF) and the inverse document frequency (IDF) of the input sentence and the text content of the candidate information.
[0223] In an example, the relevance information between the candidate information and the input sentence is:
[0224]
[0225] where D is the text content of the candidate information, Q is the input sentence, q i is the i-th word in the input sentence, and there are n words in the input sentence. IDF(q i ) is the inverse document probability between the i-th word in the input sentence and the text content of the candidate information. f(q iD) is the frequency of the ith word in the input sentence in the text content of the candidate information. |D| is the character length of the text content of the candidate information, avgdl is the average length of all candidate information in the existing knowledge base. k and b are adjustable parameters, usually k is between 1.2 and 2, and b is usually set to 0.75.
[0226] In one example, the inverse document probability is implemented as:
[0227]
[0228] Wherein, N is the total number of candidate information in the existing knowledge base, and n(qi) is the number of candidate information containing the word qi.
[0229] Step 524: determining the candidate information with the relevance information exceeding the first relevance threshold as the retrieval information;
[0230] Exemplarily, the first relevance threshold can be a preset value, or a value set according to human-computer interaction operation.
[0231] In one example, the first relevance threshold is related to the mth relevance information in the order queue. For example, the step 522 is performed on each candidate information in the existing knowledge base respectively, and N relevance information corresponding to N candidate information in the existing knowledge base is calculated. The above N relevance information is sorted from large to small to obtain the order queue. The first relevance threshold is the mth relevance information in the order queue, and m is less than N. N and m are integers.
[0232] Exemplarily, the way of retrieving the retrieval information from the existing knowledge base through the above relevance information is also called sparse retrieval.
[0233] In summary, the method provided by the embodiment determines the relevance information between the candidate information and the input sentence according to the frequency of the word in the input sentence in the candidate information through the sparse retrieval, and further determines the retrieval information in the existing knowledge base; the relevance information includes at least one of the document frequency and the inverse document frequency, which considers the importance of the word in the input sentence in the document; the prior knowledge of natural language processing and the retrieval information retrieved based on the input sentence are considered when generating the output sentence; compared with related technologies, by expanding the retrieval information as the basis for the current output word, the current output word can be determined based on more information sources, and the risk of including unverified or doubtful information in the output sentence is reduced.
[0234] Figure 7 A flowchart of a natural language processing method provided by an example embodiment of the present application is shown. The method can be executed by a computer device. That is, in the embodiment, the computer device can execute the method to generate the output sentence based on the input sentence. Figure 3In the illustrated embodiment, step 520 can be implemented as steps 526 and 528:
[0235] Step 526: Encode the first hidden layer representation of the input statement and the second hidden layer representation of the text content of the candidate information in the existing knowledge base;
[0236] For example, the first hidden layer feature encodes the natural language in the input statement into at least one of feature values, feature vectors, and feature matrices to describe the feature information of the natural language in the input statement in the hidden layer space. Similarly, the second hidden layer feature is used to describe the feature information of the text content of the candidate information in the hidden layer space. For example, the first and second hidden layer representations are obtained based on encoding by an encoder (also called an encoding network or encoding layer); further, the same encoder is used to encode the input statement to obtain the first hidden layer representation and to encode the text content of the candidate information to obtain the second hidden layer representation.
[0237] For example, the encoder can be implemented as a Bidirectional Encoder Representations from Transformers (BERT); taking candidate information as an example, the candidate information is denoted as D, and the sequence is constructed as: [CLS,d1,d2,…,d n [,SEP]; where CLS and SEP are identifiers located at the beginning and end of the natural language in the BERT model. d1 to dn are n words in the candidate information. Inputting the above sequence into the BERT model yields the second hidden layer representation:
[0238] h D =BERT([CLS,d1,d2,…,d n ,SEP]);
[0239] Among them, h D It is the second hidden layer representation, h D It is to set the sequence [CLS,d1,d2,…,d] n [SEP] is obtained from the input BERT model. Similarly, the input statement is denoted as Q, and the corresponding first hidden layer is represented as h. Q .
[0240] Step 528: If the similarity between the first hidden layer representation and the second hidden layer representation exceeds the first similarity threshold, the candidate information is determined as the retrieval information;
[0241] For example, the similarity between the first hidden layer representation and the second hidden layer representation is at least one of cosine similarity, Euclidean distance similarity, and Pearson correlation coefficient.
[0242] Exemplarily, the first similarity threshold can be a preset value or a value set according to the human-computer interaction operation.
[0243] In an example, the first similarity threshold is related to the mth similarity information in the order queue. For example, the step 526 is performed on each candidate information in the existing knowledge base to obtain N pieces of similarity information corresponding to N candidate information in the subsequent knowledge base. The N pieces of related information are sorted from large to small to obtain the order queue. The first similarity threshold is the mth similarity information in the order queue, and m is less than N. N and m are integers.
[0244] Exemplarily, the way of retrieving the retrieval information from the existing knowledge base through the similarity information is also called dense retrieval.
[0245] In summary, the method provided in the embodiment determines the similarity between the first hidden layer representation and the second hidden layer representation obtained through the encoding of the input sentence and the candidate information respectively, and further determines the retrieval information in the existing knowledge base through the dense retrieval manner. The hidden layer representation obtained through the encoding can indicate the feature information of the semantic of the natural language in the hidden layer space. The prior knowledge of the natural language processing and the retrieval information retrieved based on the input sentence are considered when generating the output sentence. Compared with the related technology, the current output word can be determined based on more information sources by expanding the retrieval information as the basis of the current output word, thereby reducing the risk of including the suspicious information that does not conform to the facts or has not been verified in the output sentence.
[0246] Figure 8 A flowchart of a natural language processing method provided by an example embodiment of the present application is shown. The method can be executed by a computer device. That is, in the Figure 3 Based on the embodiment shown, steps 521a to 521d are further included:
[0247] Step 521a: splitting the retrieval information into at least two subparts according to the separator in the retrieval information;
[0248] The retrieval information in the embodiment includes at least two subparts.
[0249] The separator includes, but is not limited to, punctuation marks such as commas and periods in the retrieval information, format identifiers such as line breaks and page breaks. The retrieval information can be split into at least two subparts, which can be one or more sentences, one or more paragraphs, one or more chapters, etc. according to the difference of the separator.
[0250] Step 521b: splitting the retrieval information into at least two subparts according to a preset string length;
[0251] According to the preset string length, the search information is split into at least two subparts, which are also referred to as at least two blocks, and the size of each subpart is the same. According to the character order in the search information, at least one overlapping character is included between two adjacent subparts obtained by the splitting; that is, some overlapping area is included between two adjacent subparts to ensure the continuity of the context in the subparts obtained by splitting the search information.
[0252] Step 521c: splitting the search information into at least two subparts based on the maximum input length of the natural language model;
[0253] In some examples, the natural language model has a maximum input length due to the design of the natural language model, and the natural language model can only process characters within the maximum input length.
[0254] For example, the obtained subpart is smaller than the maximum input length. In some examples, the obtained subpart is smaller than one half of the maximum input length, so that the spliced at least one subpart and input sentence is smaller than the maximum input length.
[0255] It should be noted that steps 521a to 521c in the embodiment can be executed alternatively or at least one of them can be executed.
[0256] Step 521d: determining at least one associated subpart from the at least two subparts according to the degree of association between the at least two subparts and the input sentence;
[0257] The at least one associated subpart is part of the content in the search information. The degree of association between the subpart and the input sentence can be implemented as at least one of the relevance information or the similarity, which will be introduced separately through examples.
[0258] For example, the at least one associated subpart is a part of the subpart with a stronger degree of association with the input sentence. For example, the associated subpart can be determined in the subpart by a preset threshold.
[0259] In the embodiment, the second probability distribution is determined by performing natural language processing on the input sentence and the at least one associated subpart. For example, the second probability distribution is output information of the natural language model based on the input sentence and the at least one associated subpart.
[0260] For example, the second set of candidate words includes at least one candidate word, and the first probability distribution is a set of at least one probability value, and the at least one probability value and the at least one candidate word in the second set of candidate words are one-to-one corresponding.
[0261] Exemplarily, step 521d in this embodiment can be implemented alone in combination with step 510, step 520, step 530 and step 550 in the foregoing to form a new embodiment, and the present application does not limit this.
[0262] To sum up, the method provided in this embodiment includes at least two subparts in the search information, at least one associated subpart is determined, the character range that needs to be input into the natural language model is further filtered and reduced on the basis of the search information, and the associated subpart that is more concise relative to the search information is provided to the natural language model, so that the irrelevant character part in the search information does not have an adverse effect on the process in which the natural language model predicts the second probability distribution based on the search information and the input sentence.
[0263] In an optional implementation manner, step 521d in the foregoing can be implemented as the following two substeps:
[0264] calculating at least two sub-relevance information corresponding to the at least two subparts respectively according to the at least two subparts and the input sentence;
[0265] Exemplarily, the sub-relevance information is used to indicate the relevance of one subpart in the search information to the input sentence, and the sub-relevance information corresponding to one subpart includes at least one of a document frequency and an inverse document frequency. For details, refer to step 522 in the foregoing.
[0266] determining the first subpart corresponding to the first sub-relevance information as the associated subpart in a case where the first sub-relevance information in the at least two sub-relevance information exceeds a second relevance threshold;
[0267] Similar to the first relevance threshold, the second relevance threshold can be a preset value or a value set according to the human-computer interaction operation. Similarly, the second relevance threshold can also be determined from the order queue, and the order queue is obtained from the sub-relevance information corresponding to each subpart in the search information from large to small.
[0268] In another optional implementation manner, step 521d in the foregoing can be implemented as the following two substeps:
[0269] obtaining a first hidden layer representation of the input sentence by encoding, and obtaining at least two sub-hidden layer representations corresponding to the at least two subparts by encoding;
[0270] The first hidden layer feature is used to describe the feature information of the natural language in the input sentence in the hidden layer space. The sub-hidden layer representation is used to describe the feature information of the natural language in one subpart in the search information in the hidden layer space. For the introduction of the at least two sub-hidden layer representations obtained by encoding, refer to step 526 in the foregoing.
[0271] • determining the first sub-part corresponding to the first sub-hidden layer representation as the relevant sub-part in a case where the similarity between the first hidden layer representation and the first sub-hidden layer representation exceeds a second similarity threshold;
[0272] Similar to the first similarity threshold, the second similarity threshold can be a preset value or a value set according to the human-computer interaction operation. Similarly, the second similarity threshold can also be determined from the order queue, which is obtained from the similarity of each sub-part in the search information from large to small.
[0273] In summary, the method provided in the embodiment determines the similarity between the first hidden layer representation and the sub-hidden layer representation according to the frequency of the words in the input sentence appearing in the sub-part in the candidate information, or respectively encodes the input sentence and the sub-part in the candidate information, and filters out the relevant sub-part in the search information, thereby avoiding the adverse effects of the irrelevant parts of the search information on the input sentence in the process of predicting the second probability distribution based on the search information and the input sentence by the natural language model.
[0274] In an application scenario, the input sentence involves the medical knowledge field, and the output sentence obtained by the natural language processing provided in the application is an answer involving the medical knowledge field. For example, answering questions in the medical knowledge field can be used for popularizing medical knowledge, assisting in symptom diagnosis, etc., and high accuracy is required for the answers, and the risk of including unverified or doubtful information in the answer sentence needs to be reduced.
[0275] • obtaining a medical sentence, the medical sentence being a character sequence with natural language semantics;
[0276] The medical sentence is a natural language character sequence related to the medical knowledge field. The input sentence is a sentence input based on a human-computer interaction operation, such as a sentence input through a human-computer interaction operation interface of an application program having at least one of a medical auxiliary diagnosis function and a medical knowledge popularization function.
[0277] • retrieving relevant medical information based on the medical sentence in a medical knowledge base;
[0278] The medical knowledge base includes at least one of a medical book, a medical paper, an Internet medical website, and a medical institution business database; the relevant medical information is a retrieval result obtained in the medical knowledge base based on the medical sentence. With the development of the medical field, new medical books and medical papers are published, and new medical knowledge can be retrieved by updating the medical knowledge base, and then used to generate an output sentence based on the new medical knowledge in the prediction process of the natural language model. Without updating the network parameters of the natural language model, the new medical knowledge can be used in the prediction process.
[0279] performing natural language processing on the medical sentence to determine a first probability distribution of a first set of candidate words;
[0280] For example, the natural language processing is used to obtain a corresponding answer according to the input sentence, and the natural language processing is usually implemented based on an artificial neural network model with a question and answer function, such as a large language model (LLM). The first set of candidate words includes at least one candidate word.
[0281] performing natural language processing on the medical sentence and the related medical information to determine a second probability distribution of a second set of candidate words;
[0282] For example, the second set of candidate words includes at least one candidate word. In one example, the candidate words include words in the field of medical knowledge.
[0283] determining an output word at a current output position according to the first probability distribution and the second probability distribution;
[0284] The current output word belongs to at least one of the first set of candidate words and the second set of candidate words, and the output sentence is an answer sentence to the medical sentence.
[0285] In an optional implementation, the determination of the current output word is introduced as follows:
[0286] obtaining a first weight parameter and a second weight parameter, and performing weighting on the first probability distribution and the second probability distribution based on the first weight parameter and the second weight parameter;
[0287] determining the current output word from the first set of candidate words and the second set of candidate words based on the fusion probability value;
[0288] For example, a product calculation is performed on the first weight parameter and the first probability distribution to obtain the weighted first probability distribution; and a product calculation is performed on the second weight parameter and the second probability distribution to obtain the weighted second probability distribution.
[0289] determining a fusion probability value corresponding to each candidate word in the first set of candidate words and the second set of candidate words according to the weighted first probability distribution and the weighted second probability distribution;
[0290] For the fusion probability value of the i th candidate word, a second probability value of the i th candidate word in the weighted second probability distribution is subtracted from a first probability value in the weighted first probability distribution to obtain an i th fusion probability value corresponding to the i th candidate word.
[0291] Further, the first weight parameter is determined as a first numerical value, and the second weight parameter is determined as a second numerical value; the first numerical value and the second numerical value are both positive numbers, and the first numerical value is less than the second numerical value.
[0292] By subtracting the second probability value from the first probability value, the influence of prior knowledge on the output sentence can be eliminated, and by adjusting the first weight parameter and the second weight parameter, the degree of eliminating the influence of prior knowledge in the second probability distribution can be controlled. In the case where the input sentence involves the field of medical knowledge, the field of medical knowledge is significantly different from general knowledge, which can achieve the control that the current output word of the output sentence is more dependent on the relevant medical information retrieved, is conducive to ensuring that the generated current output word has an information source, and reduces the risk of appearing in the output sentence with information that is inconsistent with the facts or unverified suspicious information.
[0293] In another application scenario, the input sentence involves the field of legal knowledge, and the output sentence obtained by the natural language processing provided by the present application is an answer involving the field of legal knowledge. For example, answering questions in the field of legal knowledge can be used for popularizing legal knowledge, assisting in providing legal advice, etc., which requires high accuracy of answers and reduces the risk of appearing in the answer sentence with information that is inconsistent with the facts or unverified suspicious information.
[0294] · obtaining a legal sentence, the legal sentence being a character sequence with natural language semantics;
[0295] The legal sentence is a natural language character sequence involving the field of legal knowledge. The input sentence is a sentence input based on human-computer interaction operation, such as a sentence input by a human-computer interaction operation entrance of an application program with at least one of the functions of providing legal advice and popularizing legal knowledge.
[0296] · retrieving relevant legal information based on the legal sentence in a legal knowledge base;
[0297] The legal knowledge base includes at least one of legal books, legal papers, relevant regional laws and regulations, industry specifications, and business databases of legal service agencies; the relevant legal information is a retrieval result obtained based on the legal sentence in the legal knowledge base. With the update of legal knowledge, new legal books, legal papers are published, new laws and regulations, judicial interpretations are solicited for opinions / published / implemented, etc., by updating the legal knowledge base, new legal knowledge can be retrieved, and then the output sentence can be generated based on the new legal knowledge in the prediction process of the natural language model. Without updating the network parameters of the natural language model, the new legal knowledge can be used in the prediction process.
[0298] · performing natural language processing on the legal sentence to determine a first probability distribution of a first set of candidate words;
[0299] Exemplarily, the natural language processing is used to obtain a corresponding answer according to the input sentence, and the natural language processing is usually implemented based on an artificial neural network model with a question and answer function, such as a large language model (LLM). The first set of candidate words includes at least one candidate word.
[0300] performing natural language processing on the legal sentence and related legal information to determine a second probability distribution of a second set of candidate words;
[0301] Exemplarily, the second set of candidate words includes at least one candidate word. In one example, the candidate words include words in the field of legal knowledge.
[0302] determining an output word at the current output position according to the first probability distribution and the second probability distribution;
[0303] The current output word belongs to at least one of the first set of candidate words and the second set of candidate words, and the output sentence is an answer sentence to the legal sentence.
[0304] In an optional implementation, the determination of the current output word is introduced as follows:
[0305] obtaining a first weight parameter and a second weight parameter, and performing weighting on the first probability distribution and the second probability distribution based on the first weight parameter and the second weight parameter, respectively;
[0306] determining the current output word from the first set of candidate words and the second set of candidate words based on the fusion probability value;
[0307] determining a fusion probability value corresponding to each candidate word in the first set of candidate words and the second set of candidate words according to the weighted first probability distribution and the weighted second probability distribution;
[0308] wherein, for the fusion probability value of the i th candidate word, the second probability value of the i th candidate word in the weighted second probability distribution is subtracted from the first probability value in the weighted first probability distribution to obtain the i th fusion probability value corresponding to the i th candidate word.
[0309] Further, the first weight parameter is determined as a first numerical value, and the second weight parameter is determined as a second numerical value; the first numerical value and the second numerical value are both positive numbers, and the first numerical value is smaller than the second numerical value.
[0310] The influence of prior knowledge on the output sentence can be eliminated by subtracting the second probability value from the first probability value, and the degree of eliminating the influence of prior knowledge in the second probability distribution can be controlled by adjusting the first weight parameter and the second weight parameter. In the case where the input sentence involves the field of legal knowledge, the field of legal knowledge is significantly different from general knowledge, which can achieve the control that the current output word of the output sentence is more dependent on the relevant legal information retrieved, is conducive to ensuring that the generated current output word has an information source, and reduces the risk of appearing in the output sentence with inconsistent or unverified suspicious information.
[0311] In yet another application scenario, the input sentence involves the field of programming knowledge, and the output sentence obtained by the natural language processing provided by the present application is an answer involving the field of programming knowledge. For example, answering questions in the field of programming knowledge can be used for popularizing programming knowledge, assisting in providing programming code, etc., which requires high accuracy of answers and reduces the risk of appearing in the answer statement with inconsistent or unverified suspicious information.
[0312] · obtaining a programming requirement sentence, the programming requirement sentence being a character sequence with natural language semantics;
[0313] The programming requirement sentence is a natural language character sequence involving the field of programming requirement knowledge. The input sentence is a sentence input based on human-computer interaction operation, such as a sentence input by a human-computer interaction operation entry of an application program having at least one of the functions of providing auxiliary programming and popularizing programming knowledge.
[0314] · retrieving relevant programming information based on the programming requirement sentence in the programming knowledge base;
[0315] The programming knowledge base includes at least one of programming books, usage instructions of programming software, programming examples, and business databases of open source programming information hosting agencies; the relevant programming information is a retrieval result obtained in the programming knowledge base based on the programming requirement sentence. With the update of programming knowledge and the publication of new programming examples, new programming knowledge / information can be retrieved by updating the programming knowledge base, and then the output sentence can be generated based on the new programming knowledge in the prediction process of the natural language model. Without updating the network parameters of the natural language model, the new programming knowledge can be used in the prediction process.
[0316] · performing natural language processing on the programming requirement sentence to determine a first probability distribution of a first group of candidate words;
[0317] For example, natural language processing is used to obtain a corresponding answer based on the input sentence, and natural language processing is usually implemented based on an artificial neural network model with a question and answer function, such as a large language model (LLM). The first group of candidate words includes at least one candidate word.
[0318] perform natural language processing on the programming requirement sentence and the related programming information to determine a second probability distribution of a second set of candidate words;
[0319] The second set of candidate words includes at least one candidate word. In one example, the candidate word includes a word in the programming knowledge field.
[0320] determine the output word at the current output position according to the first probability distribution and the second probability distribution;
[0321] The current output word belongs to at least one of the first set of candidate words and the second set of candidate words, and the output sentence is an answer sentence to the programming requirement sentence.
[0322] In an optional implementation, the determination of the current output word is as follows:
[0323] obtain a first weight parameter and a second weight parameter, and perform weighting on the first probability distribution and the second probability distribution based on the first weight parameter and the second weight parameter;
[0324] determine the current output word from the first set of candidate words and the second set of candidate words based on the fusion probability value;
[0325] determine a fusion probability value corresponding to each candidate word in the first set of candidate words and the second set of candidate words according to the weighted first probability distribution and the weighted second probability distribution;
[0326] The fusion probability value of the ith candidate word is obtained by subtracting the first probability value in the weighted first probability distribution from the second probability value in the weighted second probability distribution.
[0327] By subtracting the second probability value from the first probability value, the influence of prior knowledge on the output sentence can be eliminated, and by adjusting the first weight parameter and the second weight parameter, the degree of eliminating the influence of prior knowledge in the second probability distribution can be controlled. In the case where the input sentence involves the programming knowledge field, the programming knowledge field is significantly different from general knowledge, and the dependence of the current output word of the output sentence on the retrieved related programming information can be controlled to ensure that the generated current output word has an information source, the words of the output sentence have a clear information source, and in the case where the output sentence includes executable code, the executability of the output sentence can be ensured, and the risk of including unverified suspicious information in the output sentence is reduced.
[0328] Those skilled in the art can understand that the above embodiments can be independently implemented, or the above embodiments can be freely combined to form new embodiments to implement the natural language processing method of the present application.
[0329] Figure 9 A structural block diagram of a natural language processing apparatus provided by an example embodiment of the present application is shown. The apparatus includes:
[0330] An obtaining module 810 is configured to obtain an input sentence, the input sentence being a character sequence having a natural language semantic;
[0331] A retrieving module 820 is configured to retrieve, based on the input sentence, retrieval information from an existing knowledge base, the retrieval information being information associated with a semantic of at least one word in the input sentence in the existing knowledge base;
[0332] A processing module 830 is configured to perform natural language processing on the input sentence, and determine a first probability distribution, the first probability distribution being a probability distribution of a first group of candidate words at a current output position of an output sentence;
[0333] The processing module 830 is further configured to perform the natural language processing on the input sentence and the retrieval information, and determine a second probability distribution, the second probability distribution being a probability distribution of a second group of candidate words at the current output position;
[0334] A determining module 840 is configured to determine an output word at the current output position according to the first probability distribution and the second probability distribution.
[0335] In an optional design of the present application, the determining module 840 is further configured to:
[0336] obtain a first weight parameter and a second weight parameter, perform weighting on the first probability distribution based on the first weight parameter, and perform weighting on the second probability distribution based on the second weight parameter;
[0337] determine, according to the weighted first probability distribution and the weighted second probability distribution, a fusion probability value corresponding to each candidate word in the first group of candidate words and the second group of candidate words;
[0338] determine, based on the fusion probability value, the output word at the current output position from the first group of candidate words and the second group of candidate words.
[0339] In an optional design of the present application, a union of the first group of candidate words and the second group of candidate words includes j candidate words, and any two candidate words in the j candidate words are not repeated.
[0340] The determining module 840 is further configured to:
[0341] subtract the first probability value in the weighted first probability distribution from the second probability value of the i-th candidate word in the weighted second probability distribution to obtain an i-th fusion probability value corresponding to the i-th candidate word;
[0342] update i to i+1, and start to execute the step of subtracting the first probability value in the weighted first probability distribution from the second probability value of the i-th candidate word in the weighted second probability distribution to obtain an i-th fusion probability value corresponding to the i-th candidate word, until j fusion probability values are obtained.
[0343] wherein i is a positive integer, and j is a positive integer greater than or equal to i.
[0344] In an optional design of the present application, the determining module 840 is further configured to:
[0345] obtain a priori knowledge corpus, the a priori knowledge corpus comprising at least one phrase in a general knowledge field;
[0346] in a case where a number of repetitions of the phrase in the a priori knowledge corpus and the input word in the input sentence does not exceed a number threshold, determine the first weight parameter as a first numerical value and the second weight parameter as a second numerical value, the first numerical value and the second numerical value being positive numbers, and the first numerical value being less than the second numerical value;
[0347] in a case where the number of repetitions of the phrase in the a priori knowledge corpus and the input word exceeds the number threshold, determine the first weight parameter as a third numerical value and the second weight parameter as a fourth numerical value, the third numerical value being a negative number and the fourth numerical value being a positive number.
[0348] In an optional design of the present application, in the case where the number of repetitions of the phrase in the a priori knowledge corpus and the input word does not exceed the number threshold, a difference between the second numerical value and the first numerical value is positively correlated with the number of repetitions;
[0349] in the case where the number of repetitions of the phrase in the a priori knowledge corpus and the input word exceeds the number threshold, a difference between the fourth numerical value and the third numerical value is negatively correlated with the number of repetitions.
[0350] In an optional design of the present application, the processing module 830 is further configured to:
[0351] invoke a natural language model to predict the first probability distribution according to the input sentence and the existing word.
[0352] invoke the natural language model to predict the second probability distribution according to the input sentence, the retrieved information, and the existing word;
[0353] The existing word is a word in the output sentence located before the current output position.
[0354] In an optional design of the present application, the determination module 840 is further configured to:
[0355] obtain a training data set of the natural language model;
[0356] construct the prior knowledge base according to feature word groups in the training data set;
[0357] The feature word group has a document frequency greater than a first frequency threshold and an inverse document frequency less than a second frequency threshold in the training data set.
[0358] In an optional design of the present application, the determination module 840 is further configured to:
[0359] perform normalization on the j fusion probability values corresponding to the j candidate words one by one to construct j normalized probability values, and the output word is determined based on the j normalized probability values.
[0360] In an optional design of the present application, the retrieval module 820 is further configured to:
[0361] determine relevance information between the candidate information and the input sentence according to text content of the candidate information in the existing knowledge base and the input sentence;
[0362] determine the candidate information with relevance information exceeding a first relevance threshold as the retrieved information.
[0363] In an optional design of the present application, the retrieval module 820 is further configured to:
[0364] encode a first hidden layer representation of the input sentence and a second hidden layer representation of text content of the candidate information in the existing knowledge base;
[0365] In a case where a similarity between the first hidden layer representation and the second hidden layer representation exceeds a first similarity threshold, determine the candidate information as the retrieved information.
[0366] In an optional design of the present application, the retrieved information includes at least two subparts.
[0367] The determination module 840 is further configured to:
[0368] determine at least one associated subpart from the at least two subparts according to degrees of association between the at least two subparts and the input sentence respectively;
[0369] wherein the second probability distribution is determined by performing the natural language processing on the input sentence and the at least one associated subpart.
[0370] In an optional design of the present application, the determining module 840 is further configured to perform at least one of the following:
[0371] split the search information into the at least two subparts according to a separator in the search information;
[0372] split the search information into the at least two subparts according to a preset string length, wherein at least one overlapping character is included between two adjacent subparts;
[0373] split the search information into the at least two subparts based on a maximum input length of the natural language model.
[0374] In an optional design of the present application, the determining module 840 is further configured to:
[0375] calculate at least two sub-relevance information corresponding to the at least two subparts respectively according to the at least two subparts and the input sentence respectively, the sub-relevance information being used to indicate relevance between a subpart in the search information and the input sentence;
[0376] determine a first subpart corresponding to first sub-relevance information as the associated subpart in a case where the first sub-relevance information exceeds a second relevance threshold.
[0377] In an optional design of the present application, the determining module 840 is further configured to:
[0378] obtain a first hidden layer representation of the input sentence by encoding, and obtain at least two sub-hidden layer representations corresponding to the at least two subparts by encoding respectively;
[0379] determine a first subpart corresponding to a first sub-hidden layer representation as the associated subpart in a case where a similarity between the first hidden layer representation and the first sub-hidden layer representation exceeds a second similarity threshold.
[0380] It should be noted that the apparatus provided by the above embodiments is only exemplified by the above division of various functional modules when realizing its functions, and in actual application, the above functions can be completed by different functional modules according to actual needs, that is, the content structure of the device is divided into different functional modules to complete all or part of the above described functions.
[0381] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments related to the method; the technical effects achieved by the operations performed by various modules are the same as the technical effects in the embodiments related to the method, and will not be described in detail here.
[0382] The embodiments of the present application further provide a computer device, which comprises a processor and a memory, and the memory stores a computer program; the processor is used for executing the computer program in the memory to realize the natural language processing method provided by the above method embodiments.
[0383] Optionally, the computer device is a server. For example, Figure 10 Fig. 8 is a structural block diagram of a server provided by an example embodiment of the present application.
[0384] Generally, the server 2300 comprises a processor 2301 and a memory 2302.
[0385] The processor 2301 can comprise one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 2301 can be realized in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 2301 can also comprise a main processor and a coprocessor, the main processor is a processor for processing data in an awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 2301 can be integrated with a graphics processor (GPU), and the GPU is used to render and draw the content required to be displayed on the display screen. In some embodiments, the processor 2301 can further comprise an artificial intelligence (AI) processor, which is used to process machine learning related computing operations.
[0386] The memory 2302 can include one or more computer-readable storage media. The computer-readable storage media can be non-transitory. The memory 2302 can also include high-speed random access memory and can include non-volatile memory, such as one or more magnetic disk storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 2302 stores the at least one instruction for execution by the processor 2301 to implement the natural language processing method provided by the method embodiments of the present application.
[0387] In some embodiments, the server 2300 can also optionally include an input interface 2303 and an output interface 2304. The processor 2301, the memory 2302, and the input interface 2303 and the output interface 2304 can be connected through a bus or a signal line. Various peripheral devices can be connected to the input interface 2303 and the output interface 2304 through the bus, the signal line, or the circuit board. The input interface 2303 and the output interface 2304 can be used to connect at least one peripheral device related to input / output (I / O) to the processor 2301 and the memory 2302. In some embodiments, the processor 2301, the memory 2302, and the input interface 2303 and the output interface 2304 are integrated on the same chip or circuit board; in some other embodiments, any one or both of the processor 2301, the memory 2302, and the input interface 2303 and the output interface 2304 can be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.
[0388] Those skilled in the art can understand that the structure shown above does not constitute a limitation on the server 2300, and can include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0389] In an exemplary embodiment, a chip is also provided, which includes programmable logic circuitry and / or program instructions, and when the chip is running on a computer device, is used to implement the natural language processing method described in the above aspects.
[0390] In an exemplary embodiment, a computer program product is also provided, which includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor reads and executes the computer instructions from the computer readable storage medium to implement the natural language processing method provided by each of the above method embodiments.
[0391] In an example embodiment, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the natural language processing method provided by each method embodiment.
[0392] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware, and the program can be stored in a computer readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0393] Those skilled in the art should realize that, in one or more examples described above, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware or any combination thereof. When implemented by software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or codes on a computer readable medium. The computer readable medium includes a computer storage medium and a communication medium, and the communication medium includes any medium facilitating the transmission of computer programs from one place to another. The storage medium can be any available medium accessible by a general or special purpose computer.
[0394] The above description is only optional embodiments of the present application, and does not limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A natural language processing method, characterized in that, The method is performed by a computer device, and the method includes: Obtain an input statement, which is a character sequence with natural language semantics; Based on the input statement, retrieval information is obtained from the existing knowledge base. The retrieval information is information in the existing knowledge base that is semantically related to at least one word in the input statement. Natural language processing is performed on the input statement to determine a first probability distribution, which is the probability distribution of the first group of candidate words at the current output position of the output statement; The natural language processing is performed on the input statement and the search information to determine a second probability distribution, which is the probability distribution of the second group of candidate words at the current output position. The output word at the current output position is determined based on the first probability distribution and the second probability distribution.
2. The method according to claim 1, characterized in that, Determining the output word at the current output position based on the first probability distribution and the second probability distribution includes: Obtain a first weight parameter and a second weight parameter; perform weighting on the first probability distribution based on the first weight parameter, and perform weighting on the second probability distribution based on the second weight parameter; Based on the weighted first probability distribution and the weighted second probability distribution, determine the fusion probability value corresponding to each candidate word in the first group of candidate words and the second group of candidate words; Based on the fusion probability value, the output word at the current output position is determined from the first group of candidate words and the second group of candidate words.
3. The method according to claim 2, characterized in that, The union of the first group of candidate words and the second group of candidate words includes j candidate words, and any two candidate words in the j candidate words are unique; The step of determining the fusion probability value corresponding to each candidate word in the first group of candidate words and the second group of candidate words based on the weighted first probability distribution and the weighted second probability distribution includes: The second probability value of the i-th candidate word in the weighted second probability distribution is subtracted from the first probability value in the weighted first probability distribution to obtain the i-th fusion probability value corresponding to the i-th candidate word. Update i to i+1, and start executing the step of subtracting the first probability value of the weighted first probability distribution from the second probability value of the i-th candidate word in the weighted second probability distribution to obtain the i-th fusion probability value of the i-th candidate word again, until j fusion probability values are calculated; Where i is a positive integer, and j is a positive integer greater than or equal to i.
4. The method according to claim 3, characterized in that, The process of obtaining the first weight parameter and the second weight parameter includes: Obtain a prior knowledge lexicon, wherein the prior knowledge lexicon includes at least one phrase from the general knowledge domain; If the number of repetitions between the phrases in the prior knowledge base and the input words in the input statement does not exceed the number threshold, the first weight parameter is determined as a first value and the second weight parameter is determined as a second value. Both the first value and the second value are positive numbers, and the first value is less than the second value. If the number of repetitions between the phrases in the prior knowledge base and the input words exceeds the threshold, the first weight parameter is determined as a third value and the second weight parameter is determined as a fourth value, wherein the third value is negative and the fourth value is positive.
5. The method according to claim 4, characterized in that, If the number of repetitions between the phrases in the prior knowledge base and the input words does not exceed the threshold, the difference between the second value and the first value is positively correlated with the number of repetitions. If the number of repetitions between the phrases in the prior knowledge base and the input words exceeds the threshold, the difference between the fourth value and the third value is negatively correlated with the number of repetitions.
6. The method according to claim 4, characterized in that, The step of performing natural language processing on the input statement to determine the first probability distribution includes: The first probability distribution is predicted by calling a natural language model based on the input statement and existing words. The step of performing natural language processing on the input statement and the search information to determine the second probability distribution includes: The natural language model is invoked to predict the second probability distribution based on the input statement, the search information, and the existing words. The existing words are those words in the output statement that are located before the current output position.
7. The method according to claim 6, characterized in that, The acquisition of the prior knowledge lexicon includes: Obtain the training dataset for the natural language model; The prior knowledge base is constructed based on the feature word groups in the training dataset; Wherein, the document frequency of the feature word group in the training dataset is greater than a first frequency threshold, and the inverse document frequency is less than a second frequency threshold.
8. The method according to claim 3, characterized in that, The method further includes: The j fusion probability values corresponding to the j candidate words are normalized to construct j normalized probability values, and the output words are determined based on the j normalized probability values.
9. The method according to any one of claims 1 to 8, characterized in that, The step of retrieving information from an existing knowledge base based on the input statement includes: Based on the text content of the candidate information in the existing knowledge base and the input statement, determine the relevance information between the candidate information and the input statement; Candidate information whose relevance exceeds a first relevance threshold is identified as the retrieval information.
10. The method according to any one of claims 1 to 8, characterized in that, The step of retrieving information from an existing knowledge base based on the input statement includes: The input statement is encoded to obtain a first hidden layer representation, and the candidate information in the existing knowledge base is encoded to obtain a second hidden layer representation. If the similarity between the first hidden layer representation and the second hidden layer representation exceeds a first similarity threshold, the candidate information is determined as the retrieval information.
11. The method according to any one of claims 1 to 8, characterized in that, The retrieval information includes at least two sub-parts, and the method further includes: Based on the degree of association between the at least two sub-parts and the input statement, at least one associated sub-part is determined from the at least two sub-parts; The second probability distribution is determined by performing the natural language processing on the input statement and the at least one associated sub-part.
12. The method according to claim 11, characterized in that, The method further includes at least one of the following: Based on the delimiters in the search information, the search information is split into at least two sub-parts; Based on a preset string length, the retrieved information is divided into at least two sub-parts; wherein, at least one overlapping character is included between two adjacent sub-parts; Based on the maximum input length of the natural language model, the retrieved information is split into at least two sub-parts.
13. The method according to claim 11, characterized in that, The step of determining at least one related sub-part from the at least two sub-parts based on the degree of association between the at least two sub-parts and the input statement includes: Based on the at least two sub-parts and the input statement respectively, at least two sub-relevance information corresponding one-to-one with the at least two sub-parts are calculated. The sub-relevance information is used to indicate the relevance of a sub-part of the retrieved information to the input statement. If the first sub-correlation information in the at least two sub-correlation information exceeds the second correlation threshold, the first sub-part corresponding to the first sub-correlation information is determined as the associated sub-part.
14. The method according to claim 11, characterized in that, The step of determining at least one related sub-part from the at least two sub-parts based on the degree of association between the at least two sub-parts and the input statement includes: The input statement is encoded to obtain a first hidden layer representation, and the at least two sub-hidden layer representations corresponding one-to-one with the at least two sub-parts are encoded. If the similarity between the first hidden layer representation and the first sub-hidden layer representation among the at least two sub-hidden layer representations exceeds a second similarity threshold, the first sub-part corresponding to the first sub-hidden layer representation is determined as the associated sub-part.
15. A natural language processing device, characterized in that, The device includes: The acquisition module is used to acquire the input statement, which is a character sequence with natural language semantics; The retrieval module is used to retrieve retrieval information from an existing knowledge base based on the input statement. The retrieval information is information in the existing knowledge base that is semantically related to at least one word in the input statement. The processing module is used to perform natural language processing on the input statement to determine a first probability distribution, which is the probability distribution of the first group of candidate words at the current output position of the output statement; The processing module is further configured to perform natural language processing on the input statement and the search information to determine a second probability distribution, wherein the second probability distribution is the probability distribution of the second group of candidate words at the current output position; The determination module is used to determine the output word at the current output position based on the first probability distribution and the second probability distribution.
16. A computer device, characterized in that, The computer device includes: a processor and a memory, wherein the memory stores at least one program; the processor is configured to execute the at least one program in the memory to implement the natural language processing method as described in any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that, The readable storage medium stores executable instructions, which are loaded and executed by a processor to implement the natural language processing method as described in any one of claims 1 to 14.
18. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, and a processor reads from and executes the computer instructions to implement the natural language processing method as described in any one of claims 1 to 14.