Foreign Language Search Method, Device, Computer Equipment and Storage Medium for Buddhist Scriptures
By matching original texts with vernacular foreign languages, the problem of low efficiency in searching for foreign languages in Buddhist scriptures is solved, automated search is realized, and efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202210375773.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-11
AI Technical Summary
In the prior art, the foreign search efficiency of Buddhist scriptures is low, and manual comparison of Buddhist scriptures with English Buddhist scriptures is required, resulting in a decrease in efficiency.
By matching the original text of the Buddhist scriptures, obtain the vernacular text corresponding to the matching target Buddhist scriptures, and match the foreign text according to the target foreign text corresponding to the vernacular text and the vernacular text, the foreign text search of the Buddhist scriptures can be automatically completed.
Without manual comparison, the foreign language search efficiency of Buddhist scriptures is improved, the automated search process is realized, and the accuracy and stability of the search is enhanced.
Smart Images

Figure CN114817496B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and text processing, and particularly to a method, device, computer equipment and storage medium for searching foreign language versions of Buddhist scriptures. Background Art
[0002] Chinese and English are the two major languages in the world and also the two languages that Chinese people encounter the most. Therefore, searching for the corresponding English version of a Buddhist scripture according to the original Buddhist scripture text is a real need for researchers. When users search for the corresponding English version of a Buddhist scripture according to the original Buddhist scripture text, they usually compare the Buddhist scripture text with the English version of the Buddhist scripture manually to obtain the English version of the Buddhist scripture corresponding to the Buddhist scripture text to be searched. This manual comparison method greatly reduces the search efficiency.
[0003] Therefore, how to improve the foreign language search efficiency of Buddhist scriptures has become an urgent problem to be solved. Summary of the Invention
[0004] The present application provides a method, device, computer equipment and storage medium for searching foreign language versions of Buddhist scriptures. By matching the Buddhist scripture text with the original Buddhist scripture, obtaining the vernacular corresponding to the matched original Buddhist scripture, and performing foreign language matching based on the vernacular and the target foreign language paragraph corresponding to the vernacular, it is possible to automatically complete the foreign language search of Buddhist scriptures without manual comparison, thereby improving the foreign language search efficiency of Buddhist scriptures.
[0005] In a first aspect, the present application provides a method for searching foreign language versions of Buddhist scriptures, the method comprising:
[0006] Obtaining the Buddhist scripture text to be searched;
[0007] Based on a Buddhist text matching model, performing matching of the Buddhist scripture text with the original Buddhist scripture to obtain the target original Buddhist scripture corresponding to the matched Buddhist scripture text;
[0008] Obtaining the vernacular corresponding to the target original Buddhist scripture, and determining the target foreign language paragraph corresponding to the vernacular;
[0009] Based on a reading comprehension model, performing foreign language matching according to the vernacular and the target foreign language paragraph to obtain the foreign language search result corresponding to the Buddhist scripture text.
[0010] In a second aspect, the present application further provides a device for searching foreign language versions of Buddhist scriptures, the device comprising:
[0011] A Buddhist scripture text acquisition module, configured to obtain the Buddhist scripture text to be searched;
[0012] A Buddhist scripture original text matching module, configured to perform matching of the Buddhist scripture text with the original Buddhist scripture based on a Buddhist text matching model to obtain the target original Buddhist scripture corresponding to the matched Buddhist scripture text;
[0013] A foreign language paragraph determination module, configured to obtain the vernacular corresponding to the original text of the target Buddhist scripture and determine the target foreign language paragraph corresponding to the vernacular;
[0014] A foreign language matching module, configured to perform foreign language matching based on a reading comprehension model according to the vernacular and the target foreign language paragraph to obtain a foreign language search result corresponding to the Buddhist scripture text.
[0015] In a third aspect, the present application further provides a computer device, which includes a memory and a processor;
[0016] The memory is used to store a computer program;
[0017] The processor is configured to execute the computer program and, when executing the computer program, implement the foreign language search method for Buddhist scriptures as described above.
[0018] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the foreign language search method for Buddhist scriptures as described above.
[0019] The present application discloses a foreign language search method, device, computer device and storage medium for Buddhist scriptures. By obtaining the Buddhist scripture text to be searched and performing Buddhist scripture original text matching on the Buddhist scripture text based on a Buddhist text matching model, a target Buddhist scripture original text corresponding to the Buddhist scripture text can be obtained; by obtaining the vernacular corresponding to the target Buddhist scripture original text and determining the target foreign language paragraph corresponding to the vernacular, it is possible to make full use of the relationship between the Buddhist scripture original text, the vernacular and the foreign language for searching, which is more stable and reliable than directly translating the Buddhist scripture text into a foreign language; by performing foreign language matching based on a reading comprehension model according to the vernacular and the target foreign language paragraph to obtain a foreign language search result corresponding to the Buddhist scripture text, it is possible to automatically complete the foreign language search of Buddhist scriptures without manual comparison, improving the foreign language search efficiency of Buddhist scriptures. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic flowchart of a foreign language search method for Buddhist scriptures provided by an embodiment of the present application;
[0022] Figure 2It is a schematic flowchart of a sub-step for matching Buddhist scripture texts with the original Buddhist scriptures provided by an embodiment of the present application;
[0023] Figure 3 It is a schematic diagram of a foreign language search for Buddhist scriptures provided by an embodiment of the present application;
[0024] Figure 4 It is a schematic diagram of another foreign language search for Buddhist scriptures provided by an embodiment of the present application;
[0025] Figure 5 It is a schematic flowchart of a sub-step for performing foreign language matching provided by an embodiment of the present application;
[0026] Figure 6 It is a schematic diagram of a display interface provided by an embodiment of the present application;
[0027] Figure 7 It is a schematic diagram of another display interface provided by an embodiment of the present application;
[0028] Figure 8 It is a schematic block diagram of a foreign language search device for Buddhist scriptures provided by an embodiment of the present application;
[0029] Figure 9 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0031] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.
[0032] It should be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0033] It should also be understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0034] Embodiments of the present application provide a method, apparatus, computer device, and storage medium for searching for foreign-language versions of Buddhist scriptures. Among them, the method for searching for foreign-language versions of Buddhist scriptures can be applied to a server or a terminal. By matching the Buddhist scripture text with the original Buddhist scripture text, the vernacular corresponding to the matched original Buddhist scripture text is obtained, and foreign-language matching is performed based on the vernacular and the target foreign-language paragraph corresponding to the vernacular, so as to automatically complete the search for foreign-language versions of Buddhist scriptures without manual comparison, improving the efficiency of searching for foreign-language versions of Buddhist scriptures.
[0035] Among them, the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal can be an electronic device such as a smart phone, a tablet computer, a laptop computer, and a desktop computer.
[0036] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0037] As Figure 1 shown, the method for searching for foreign-language versions of Buddhist scriptures includes steps S10 to S40.
[0038] Step S10: Obtain the Buddhist scripture text to be searched.
[0039] It should be noted that the embodiments of the present application can be applied to a cross-language search system, and users can perform foreign-language retrieval of Buddhist scriptures in the cross-language search system. For example, retrieve the corresponding English for a Buddhist scripture. Of course, it is also possible to retrieve Buddhist scripture information in other languages, such as Thai, German, Russian, etc. In the embodiments of the present application, the retrieval of the corresponding English for a Buddhist scripture is used as an example for illustration.
[0040] Exemplarily, the Buddhist scripture text input in the search box in the cross-language search system can be determined as the Buddhist scripture text to be searched. It should be noted that the user can input a single Buddhist scripture text or a paragraph of Buddhist scripture text in the search box in the cross-language search system and click the search button, and then can view the corresponding foreign-language search results on the search interface in the cross-language search system.
[0041] In some embodiments, after obtaining the Buddhist scripture text to be searched, the Buddhist scripture text to be searched can also be stored in a local database or a local disk. When the user subsequently enters the Buddhist scripture text in the search box of the cross-language search system, the Buddhist scripture text being entered can be supplemented and displayed based on the Buddhist scripture text stored in the local database or local disk. Thus, the efficiency of the user entering the Buddhist scripture text can be effectively improved.
[0042] Step S20: Based on the Buddhist text matching model, perform matching of the Buddhist scripture text with the original Buddhist scriptures to obtain the target original Buddhist scriptures corresponding to the matched Buddhist scripture text.
[0043] In the embodiments of the present application, after obtaining the Buddhist scripture text to be searched, the Buddhist scripture text can be input into the Buddhist text matching model for matching with the original Buddhist scriptures to obtain the target original Buddhist scriptures corresponding to the matched Buddhist scripture text.
[0044] By performing matching of the Buddhist scripture text with the original Buddhist scriptures, the target original Buddhist scriptures corresponding to the matched Buddhist scripture text can be obtained. Subsequently, the vernacular Chinese corresponding to the target original Buddhist scriptures can be obtained, and the target foreign language paragraph corresponding to the vernacular Chinese can be determined, avoiding directly translating the Buddhist scripture text into a foreign language.
[0045] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the sub-steps for performing matching of the Buddhist scripture text with the original Buddhist scriptures provided in the embodiments of the present application, and specifically may include the following steps S201 to step S203.
[0046] Step S201: Obtain at least one candidate original Buddhist scripture, and respectively splice the Buddhist scripture text with each candidate original Buddhist scripture to obtain the initial Buddhist scripture set corresponding to each candidate original Buddhist scripture.
[0047] Exemplarily, the original Buddhist scriptures can be stored in a Buddhist scripture database. When obtaining the candidate original Buddhist scriptures, the Buddhist scripture database can be queried, and the Buddhist scripture sentences or Buddhist scripture paragraphs having the same or similar characters as the Buddhist scripture text to be queried are determined as candidate original Buddhist scriptures. Then, the Buddhist scripture text is respectively spliced with each candidate original Buddhist scripture using the [SEP] symbol to obtain the initial Buddhist scripture set corresponding to each candidate original Buddhist scripture.
[0048] Step S202: Input each initial Buddhist scripture set into the Buddhist text matching model to obtain the matching score corresponding to each initial Buddhist scripture set, where the matching score is the score of the matching between the Buddhist scripture text in the initial Buddhist scripture set and the candidate original Buddhist scripture.
[0049] It should be noted that the Buddhist text matching model includes a BERT (Bidirectional Encoder Representations from Transformer) model, a fully connected layer, and a normalization layer. Among them, the BERT model is a vectorization model used to output the first word vector corresponding to the Buddhist text in the initial Buddhist scripture set and the second word vector corresponding to the candidate Buddhist scripture original text. The fully connected layer is used to output the similarity between the first word vector and the second word vector, where the similarity can be used as a matching score. It can be understood that the greater the similarity between the first word vector and the second word vector, the more matching the Buddhist text is with the candidate Buddhist scripture original text. Therefore, the similarity between the first word vector and the second word vector can be used as the matching score between the Buddhist text in the initial Buddhist scripture set and the candidate Buddhist scripture original text. The normalization layer is a softmax regression classifier used to receive an arbitrary real-valued element and compress the real-valued element into a relative probability. The value of the relative probability corresponding to this element is between 0 and 1, and the sum of the relative probabilities corresponding to all real-valued elements is 1. In the embodiments of the present application, the relative probability output by the normalization layer can be used as the matching score corresponding to the initial Buddhist scripture set.
[0050] Exemplarily, the Buddhist text matching model can be a trained model. In the embodiments of the present application, the Buddhist text matching model can be pre-trained with Buddhist scripture data to obtain a trained Buddhist text matching model. Among them, the training process of the Buddhist text matching model is as follows: Obtain a preset number of Buddhist scripture sample data, where the Buddhist scripture sample data includes sample Buddhist texts and sample Buddhist scripture original texts; splice the sample Buddhist texts and the sample Buddhist scripture original texts to determine the training sample data for each round of training, and label category labels for the training sample data for each round; input the training sample data for the current round into the BERT model for vectorization to output the first word vector and the second word vector corresponding to the training sample data for the current round; based on the fully connected layer, perform a fully connected process on the first word vector and the second word vector to obtain the similarity corresponding to the training sample data for the current round; based on the normalization layer, perform a normalization process on the similarity to obtain the matching score corresponding to the training sample data for the current round; based on a preset loss function, calculate the loss function value according to the category label corresponding to the training sample data for the current round; when the loss function value is greater than the loss function threshold, adjust the parameters in the Buddhist text matching model, perform the next round of training and calculate the loss function value for each round; when the calculated loss function threshold is less than the preset loss value or no longer decreases, the training ends, and a trained Buddhist text matching model is obtained.
[0051] Exemplarily, the category labels can include matching and non-matching; the preset loss function can adopt a list-wise loss function. Of course, other types of loss functions can also be adopted. For example, absolute value loss function, logarithmic loss function, square loss function, and exponential loss function, etc.
[0052] Among them, the list-wise loss function is as follows:
[0053]
[0054] In the formula, represents the matching score of the th sentence in the training sample data of the current round; represents the number of sentences with matching category labels in the training sample data of the current round.
[0055] Exemplarily, when adjusting the parameters in the Buddhist text matching model, the parameters of the Buddhist text matching model can be adjusted through the gradient descent algorithm, and the parameters of the Buddhist text matching model can also be adjusted through the backpropagation algorithm, which is not limited herein. The preset loss function threshold can be set according to the actual situation, and the specific value is not limited herein.
[0056] By training the initial Buddhist text matching model based on the Buddhist scripture sample data until convergence, the accuracy of the trained Buddhist scripture matching model for matching the original Buddhist scriptures can be improved. By inputting each initial Buddhist scripture set into the Buddhist text matching model, the corresponding matching score of each initial Buddhist scripture set can be obtained, and then the candidate original Buddhist scriptures that match the Buddhist scripture text can be determined as the target original Buddhist scriptures according to the matching score, ensuring the accuracy of the target original Buddhist scriptures.
[0057] Step S203, determine the target original Buddhist scriptures according to the matching score corresponding to each of the initial Buddhist scripture sets.
[0058] In some embodiments, determining the target original Buddhist scriptures according to the matching score corresponding to each initial Buddhist scripture set may include: determining the initial Buddhist scripture sets with matching scores greater than the preset threshold as the target Buddhist scripture sets; determining the candidate original Buddhist scriptures in the target Buddhist scripture sets as the target original Buddhist scriptures.
[0059] Among them, the preset threshold can be set according to the actual situation, and the specific value is not limited herein. In the embodiments of the present application, the initial Buddhist scripture set with the largest matching score can also be determined as the target Buddhist scripture set.
[0060] Exemplarily, the candidate original Buddhist scriptures in the target Buddhist scripture sets can be determined as the target original Buddhist scriptures. Among them, since there is one or more candidate original Buddhist scriptures, the obtained target original Buddhist scriptures can be one or more, and the target original Buddhist scriptures can be sorted according to the size of the matching score.
[0061] Please refer to Figure 3 , Figure 3 is a schematic diagram of a foreign language search for Buddhist scriptures provided by the embodiments of the present application. As shown in Figure 3As shown, when there is one target Buddhist scripture original text corresponding to the Buddhist scripture text, query the vernacular Chinese corresponding to the target Buddhist scripture original text, and query the target foreign language paragraph corresponding to the vernacular Chinese; then, based on the reading comprehension model, perform foreign language matching according to the vernacular Chinese and the corresponding target foreign language paragraph to obtain the foreign language search result corresponding to the Buddhist scripture text.
[0062] Please refer to Figure 4 , Figure 4 FIG. is a schematic diagram of another foreign language search for Buddhist scriptures provided by an embodiment of the present application. As Figure 4 shown, when there are multiple target Buddhist scripture original texts, the vernacular Chinese corresponding to each target Buddhist scripture original text can be queried, and the target foreign language paragraph corresponding to each vernacular Chinese can be queried; then, based on the reading comprehension model, perform foreign language matching according to each vernacular Chinese and the corresponding target foreign language paragraph to obtain multiple foreign language search results corresponding to the Buddhist scripture text; finally, sort the multiple foreign language search results according to the sorting of the target Buddhist scripture original texts.
[0063] To further ensure the privacy and security of the above-mentioned target Buddhist scripture original texts, the above-mentioned target Buddhist scripture original texts can be stored in a node of a blockchain.
[0064] Step S30: Obtain the vernacular Chinese corresponding to the target Buddhist scripture original text, and determine the target foreign language paragraph corresponding to the vernacular Chinese.
[0065] It should be noted that by obtaining the vernacular Chinese corresponding to the target Buddhist scripture original text and determining the target foreign language paragraph corresponding to the vernacular Chinese, the difficulty of directly retrieving foreign languages according to Buddhist scripture texts is indirectly solved, and the relationship between Buddhist scripture original texts, vernacular Chinese, and foreign languages can be fully utilized for searching, which is more stable and reliable than directly translating Buddhist scripture texts into foreign languages, thereby improving the accuracy of foreign language search for Buddhist scriptures. It can be understood that if the machine translation method is used to directly translate the Buddhist scripture text to be searched into a foreign language, not only the cost is high, but also the translation result is not accurate enough.
[0066] In the embodiment of the present application, for the sake of description, an example where there is one target Buddhist scripture original text is used for illustration.
[0067] Exemplarily, after obtaining the target Buddhist scripture original text corresponding to the Buddhist scripture text, the vernacular Chinese corresponding to the target Buddhist scripture original text can be obtained from the local database or local disk. It should be noted that in the embodiment of the present application, the Buddhist scripture original text and the corresponding vernacular Chinese can be associated and stored in advance.
[0068] In some embodiments, determining the target foreign language paragraph corresponding to the vernacular Chinese may include: based on the preset correspondence between the vernacular Chinese paragraph and the foreign language paragraph, determining the foreign language paragraph corresponding to the vernacular Chinese paragraph where the vernacular Chinese is located as the target foreign language paragraph.
[0069] In the embodiments of the present application, the vernacular paragraphs and foreign language paragraphs can be associated and stored in a local database or a local disk in advance.
[0070] In the embodiments of the present application, by determining the foreign language paragraph corresponding to the vernacular paragraph where the vernacular is located as the target foreign language paragraph based on the correspondence between the preset vernacular paragraph and the foreign language paragraph, the relationship between the vernacular and the foreign language in the existing Buddhist scripture data can be fully utilized to obtain the target foreign language paragraph corresponding to the vernacular, without directly translating the vernacular into a foreign language, which not only saves resources but also makes the obtained target foreign language paragraph more accurate and reliable.
[0071] Step S40: Based on the reading comprehension model, perform foreign language matching according to the vernacular and the target foreign language paragraph to obtain the foreign language search result corresponding to the Buddhist scripture text.
[0072] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a sub-step for performing foreign language matching provided by the embodiments of the present application, and specifically may include the following steps S401 to step S403.
[0073] Step S401: Concatenate the vernacular and the target foreign language paragraph to obtain a sentence set, and the sentence set includes at least two phrases.
[0074] Exemplarily, the vernacular and the target foreign language paragraph can be concatenated through the [SEP] symbol to obtain a sentence set.
[0075] Step S402: Input the sentence set into the reading comprehension model for answer prediction to obtain the answer prediction result corresponding to each phrase in the sentence set.
[0076] It should be noted that the reading comprehension model can be the XLM-Roberta model, where the XLM-Roberta model is a multilingual model for information extraction. For example, the XLM-Roberta model can output the corresponding answer according to the input question and original text, where the answer is a word extracted from the original text. In the embodiments of the present application, the vernacular is used as the question, and the target foreign language paragraph is used as the original text. The XLM-Roberta model is used to perform binary classification prediction on the vernacular and the target foreign language paragraph, predicting which phrases are the starting phrases and which phrases are the ending phrases, so as to extract the target foreign language paragraph according to the starting phrases and the ending phrases to obtain the foreign language search result.
[0077] Exemplarily, a set of statements is input into a reading comprehension model for answer prediction, and an answer prediction result corresponding to each phrase in the set of statements is obtained, where the answer prediction result includes a first prediction probability corresponding to the phrase as the starting phrase and a second prediction probability corresponding to the phrase as the ending phrase.
[0078] Exemplarily, the reading comprehension model can be a trained model. In the embodiments of the present application, the reading comprehension model can be pre-trained with Buddhist scripture data to obtain a trained reading comprehension model.
[0079] Among them, the training process of the reading comprehension model is as follows: obtain a preset number of Buddhist scripture data, where the Buddhist scripture data includes sample vernacular Chinese and sample foreign language paragraphs; splice the sample vernacular Chinese and the sample foreign language paragraphs to determine the training sample data for each round of training, and label a category label for each phrase in the training sample data for each round of training; input the training sample data for the current round into the initial reading comprehension model for classification prediction, and output the classification prediction result corresponding to the training sample data for the current round; based on a preset loss function, calculate the loss function value according to the classification prediction result corresponding to the training sample data for the current round and the category label; when the loss function value is greater than the loss function threshold, adjust the parameters in the reading comprehension model, perform the next round of training and calculate the loss function value for each round; when the calculated loss function threshold is less than the preset loss value or no longer decreases, the training ends, and a trained reading comprehension model is obtained.
[0080] Among them, the classification prediction result includes the prediction probability corresponding to the phrase as the starting phrase or the ending phrase; the category label can include the starting phrase and the ending phrase;
[0081] Exemplarily, the preset loss function can adopt the cross-entropy cost function. Of course, other types of loss functions can also be adopted, which are not limited herein.
[0082] Among them, the cross-entropy cost function is:
[0083]
[0084] In the formula, ; represents the number of phrases ; represents the true category label corresponding to the phrase , where represents the starting phrase label, represents the ending phrase label; ; is the weight value of the fully connected layer; is the prediction probability of the phrase for the true category label ; The vector corresponding to the th phrase of the hidden layer output of XML-Roberta.
[0085] Exemplarily, when adjusting the parameters in the reading comprehension model, the parameters of the reading comprehension model can be adjusted by the gradient descent algorithm, and the parameters of the reading comprehension model can also be adjusted by the backpropagation algorithm, which is not limited herein. The loss function threshold can be set according to the actual situation, and the specific value is not limited herein.
[0086] By training the initial reading comprehension model according to the sample vernacular and the sample foreign language paragraph until convergence, the accuracy of the foreign language matching of the trained reading comprehension model can be improved.
[0087] Step S403: Determine the foreign language search result according to the answer prediction result and the target foreign language paragraph.
[0088] In some embodiments, determining the foreign language search result according to the answer prediction result and the target foreign language paragraph may include: determining the target start phrase and the target end phrase according to the first prediction probability and the second prediction probability corresponding to each phrase; determining the target sentence in the target foreign language paragraph according to the target start phrase and the target end phrase, and determining the target sentence as the foreign language search result.
[0089] Among them, determining the target start phrase and the target end phrase according to the first prediction probability and the second prediction probability corresponding to each phrase includes: determining the third prediction probability corresponding to each phrase, where the third prediction probability is the difference between the first prediction probability and the second prediction probability of each phrase; determining the fourth prediction probability corresponding to each phrase, where the fourth prediction probability is the difference between the second prediction probability and the first prediction probability of each phrase; determining the phrase corresponding to the largest third prediction probability as the target start phrase, and determining the phrase corresponding to the largest fourth prediction probability as the target end phrase.
[0090] Exemplarily, the difference between the first prediction probability and the second prediction probability of each phrase can be determined as the third prediction probability of each phrase; the difference between the second prediction probability and the first prediction probability of each phrase can be determined as the fourth prediction probability of each phrase.
[0091] Exemplarily, for phrases A, B, C, and D, if the phrase corresponding to the largest third prediction probability is phrase A, then phrase A is determined as the target starting phrase; if the phrase corresponding to the largest fourth prediction probability is phrase D, then phrase D is determined as the target starting phrase. Then, in the target foreign language paragraph, the sentences between phrase A and phrase D are determined as the target sentences. It should be noted that the target starting phrase and the target ending phrase are usually generated in the target foreign language paragraph. If the target starting phrase is generated in vernacular Chinese, then according to the target starting phrase and the target ending phrase, the target sentences are extracted from the sentence set; then, the vernacular Chinese part in the target sentences is deleted. Thus, the obtained target sentences are in foreign language.
[0092] In some embodiments, when it is detected that the user enters the Buddhist scripture text "demoted to the south of the Five Ridges" in the search box of the cross-language search system, the Buddhist scripture text "demoted to the south of the Five Ridges" can be matched with the original Buddhist scripture based on the Buddhist scripture matching model to obtain the target original Buddhist scripture corresponding to the Buddhist scripture text "demoted to the south of the Five Ridges". For example, the target original Buddhist scripture is "Huineng's strict father was originally from Fanyang. He was demoted and exiled to the south of the Five Ridges and became a commoner in Xinzhou". Then, the vernacular Chinese corresponding to the target original Buddhist scripture is obtained, and the target foreign language paragraph corresponding to the vernacular Chinese is determined. For example, the vernacular Chinese is "Huineng's father was originally from Fanyang. He was demoted and exiled to the south of the Five Ridges and became a commoner in Xinzhou". Finally, the above vernacular Chinese and the target foreign language paragraph corresponding to the vernacular Chinese are input into the reading comprehension model for foreign language matching to obtain the foreign language search result corresponding to the Buddhist scripture text "demoted to the south of the Five Ridges". For example, the foreign language search result is
[0093] "HuiNeng's stern father was originally from Fanyang. He was banished to Xinzhou in Lingnan, where he became a commoner".
[0094] Exemplarily, after obtaining the foreign language search result corresponding to the Buddhist scripture text, the foreign language search result can be output. For example, the foreign language search result is displayed on the display interface of the cross-language search system.
[0095] Please refer to Figure 6 , Figure 6 which is a schematic diagram of a display interface provided by an embodiment of the present application. As Figure 6 shown, the foreign language search result can be displayed on the display interface.
[0096] Please refer to Figure 7 , Figure 7 which is another schematic diagram of a display interface provided by an embodiment of the present application. As Figure 7 shown, the target original Buddhist scripture, the vernacular Chinese, and the foreign language search result can be displayed on the display interface.
[0097] By using a reading comprehension model to perform foreign language matching based on vernacular Chinese and target foreign language paragraphs, foreign language search results corresponding to Buddhist scripture texts can be obtained, enabling automatic foreign language search of Buddhist scriptures without manual comparison, thereby improving the efficiency of foreign language search of Buddhist scriptures.
[0098] The foreign language search method for Buddhist scriptures provided in the above embodiments can obtain the target Buddhist scripture original text corresponding to the Buddhist scripture text through matching the Buddhist scripture text with the original Buddhist scripture text. Subsequently, the vernacular Chinese corresponding to the target Buddhist scripture original text can be obtained, and the target foreign language paragraph corresponding to the vernacular Chinese can be determined, avoiding directly translating the Buddhist scripture text into a foreign language; by inputting each initial Buddhist scripture set into the Buddhist text matching model, the matching score corresponding to each initial Buddhist scripture set can be obtained, and then the candidate Buddhist scripture original text matching the Buddhist scripture text can be determined as the target Buddhist scripture original text according to the matching score, ensuring the accuracy of the target Buddhist scripture original text; by obtaining the vernacular Chinese corresponding to the target Buddhist scripture original text and determining the target foreign language paragraph corresponding to the vernacular Chinese, the difficulty of directly retrieving a foreign language based on the Buddhist scripture text is indirectly solved, enabling full utilization of the relationship between the original Buddhist scripture text, vernacular Chinese, and foreign language for search, which is more stable and reliable than directly translating the Buddhist scripture text into a foreign language, thereby improving the accuracy of foreign language search of Buddhist scriptures; by using a reading comprehension model to perform foreign language matching based on vernacular Chinese and target foreign language paragraphs, foreign language search results corresponding to Buddhist scripture texts can be obtained, enabling automatic foreign language search of Buddhist scriptures without manual comparison, thereby improving the efficiency of foreign language search of Buddhist scriptures.
[0099] Please refer to Figure 8 , Figure 8 FIG. is a schematic block diagram of a foreign language search device 1000 for Buddhist scriptures provided by an embodiment of the present application. The foreign language search device for Buddhist scriptures is used to execute the foregoing foreign language search method for Buddhist scriptures. Among them, the foreign language search device for Buddhist scriptures can be configured in a server or a terminal.
[0100] As Figure 8 shown, the foreign language search device 1000 for Buddhist scriptures includes: a Buddhist scripture text acquisition module 1001, a Buddhist scripture original text matching module 1002, a foreign language paragraph determination module 1003, and a foreign language matching module 1004.
[0101] The Buddhist scripture text acquisition module 1001 is used to acquire the Buddhist scripture text to be searched.
[0102] The Buddhist scripture original text matching module 1002 is used to perform matching of the Buddhist scripture text with the original Buddhist scripture text based on the Buddhist text matching model to obtain the target Buddhist scripture original text corresponding to the Buddhist scripture text.
[0103] The foreign language paragraph determination module 1003 is used to acquire the vernacular Chinese corresponding to the target Buddhist scripture original text and determine the target foreign language paragraph corresponding to the vernacular Chinese.
[0104] The foreign language matching module 1004 is used to perform foreign language matching on the vernacular and the target foreign language paragraph based on the reading comprehension model, so as to obtain the foreign language search result corresponding to the Buddhist scripture text.
[0105] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0106] The above device can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 9 shown.
[0107] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.
[0108] Please refer to Figure 9 , the computer device includes a processor and a memory connected through a system bus. Among them, the memory can include a storage medium and an internal memory. The storage medium can be a non-volatile storage medium or a volatile storage medium.
[0109] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0110] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any foreign language search method for Buddhist scriptures.
[0111] It should be understood that the processor can be a central processing unit (CPU), and this processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.
[0112] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:
[0113] Obtain the Buddhist scripture text to be searched; based on the Buddhist text matching model, perform matching of the original Buddhist scripture text on the Buddhist scripture text to obtain the target original Buddhist scripture text corresponding to the Buddhist scripture text; obtain the vernacular corresponding to the target original Buddhist scripture text, and determine the target foreign language paragraph corresponding to the vernacular; based on the reading comprehension model, perform foreign language matching according to the vernacular and the target foreign language paragraph to obtain the foreign language search result corresponding to the Buddhist scripture text.
[0114] In one embodiment, when the processor implements performing matching of the original Buddhist scripture text on the Buddhist scripture text based on the Buddhist text matching model to obtain the target original Buddhist scripture text corresponding to the Buddhist scripture text, it is used to implement:
[0115] Obtain at least one candidate original Buddhist scripture text, respectively splice the Buddhist scripture text with each candidate original Buddhist scripture text to obtain an initial Buddhist scripture set corresponding to each candidate original Buddhist scripture text; input each initial Buddhist scripture set into the Buddhist text matching model to obtain the matching score corresponding to each initial Buddhist scripture set, where the matching score is the score of the Buddhist scripture text in the initial Buddhist scripture set matching the candidate original Buddhist scripture text; determine the target original Buddhist scripture text according to the matching score corresponding to each initial Buddhist scripture set.
[0116] In one embodiment, when the processor implements determining the target original Buddhist scripture text according to the matching score corresponding to each initial Buddhist scripture set, it is used to implement:
[0117] Determine the initial Buddhist scripture set with a matching score greater than the preset threshold as the target Buddhist scripture set; determine the candidate original Buddhist scripture text in the target Buddhist scripture set as the target original Buddhist scripture text.
[0118] In one embodiment, when the processor implements determining the target foreign language paragraph corresponding to the vernacular, it is used to implement:
[0119] Based on the corresponding relationship between the preset vernacular paragraph and the foreign language paragraph, determine the foreign language paragraph corresponding to the vernacular paragraph where the vernacular is located as the target foreign language paragraph.
[0120] In one embodiment, when the processor implements performing foreign language matching based on the reading comprehension model according to the vernacular and the target foreign language paragraph to obtain the foreign language search result corresponding to the Buddhist scripture text, it is used to implement:
[0121] Splice the vernacular and the target foreign language paragraph to obtain a sentence set, where the sentence set includes at least two phrases; input the sentence set into the reading comprehension model for answer prediction to obtain the answer prediction result corresponding to each phrase in the sentence set; determine the foreign language search result according to the answer prediction result and the target foreign language paragraph.
[0122] In one embodiment, the answer prediction result includes a first prediction probability corresponding to a phrase as a starting phrase and a second prediction probability corresponding to a phrase as an ending phrase; when the processor implements determining the foreign language search result according to the answer prediction result and the target foreign language paragraph, it is used to implement:
[0123] Determine a target starting phrase and a target ending phrase according to the first prediction probability and the second prediction probability corresponding to each phrase; determine a target sentence in the target foreign language paragraph according to the target starting phrase and the target ending phrase, and determine the target sentence as the foreign language search result.
[0124] In one embodiment, when the processor implements determining a target starting phrase and a target ending phrase according to the first prediction probability and the second prediction probability corresponding to each phrase, it is used to implement:
[0125] Determine a third prediction probability corresponding to each phrase, where the third prediction probability is the difference between the first prediction probability and the second prediction probability of each phrase; determine a fourth prediction probability corresponding to each phrase, where the fourth prediction probability is the difference between the second prediction probability and the first prediction probability of each phrase; determine the phrase corresponding to the largest third prediction probability as the target starting phrase, and determine the phrase corresponding to the largest fourth prediction probability as the target ending phrase.
[0126] An embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, the computer program includes program instructions, and the processor executes the program instructions to implement any foreign language search method for Buddhist scriptures provided in the embodiments of the present application.
[0127] For example, when the program is loaded by the processor, the following steps can be executed:
[0128] Obtain the Buddhist scripture text to be searched; based on the Buddhist text matching model, perform matching of the Buddhist scripture text with the original Buddhist scripture text to obtain the target original Buddhist scripture text corresponding to the Buddhist scripture text; obtain the vernacular corresponding to the target original Buddhist scripture text, and determine the target foreign language paragraph corresponding to the vernacular; based on the reading comprehension model, perform foreign language matching according to the vernacular and the target foreign language paragraph to obtain the foreign language search result corresponding to the Buddhist scripture text.
[0129] Among them, the computer-readable storage medium may be the internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital Card (SD Card), a Flash Card, etc. equipped on the computer device.
[0130] Further, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.
[0131] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. A blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0132] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for searching foreign languages of Buddhist scriptures, characterized in that, it includes: Obtain the Buddhist scripture text to be searched; Based on the Buddhist text matching model, perform matching of the original Buddhist scripture for the Buddhist scripture text, and obtain the target original Buddhist scripture corresponding to the matching of the Buddhist scripture text; Obtain the vernacular corresponding to the target original Buddhist scripture, and determine the target foreign language paragraph corresponding to the vernacular; Based on the reading comprehension model, perform foreign language matching according to the vernacular and the target foreign language paragraph, and obtain the foreign language search result corresponding to the Buddhist scripture text; The determination of the target foreign language paragraph corresponding to the vernacular includes: based on the preset correspondence between the vernacular paragraph and the foreign language paragraph, determine the foreign language paragraph corresponding to the vernacular paragraph where the vernacular is located as the target foreign language paragraph; The performing foreign language matching according to the vernacular and the target foreign language paragraph based on the reading comprehension model to obtain the foreign language search result corresponding to the Buddhist scripture text includes: splicing the vernacular and the target foreign language paragraph to obtain a sentence set, the sentence set including at least two phrases; inputting the sentence set into the reading comprehension model for answer prediction, and obtaining the answer prediction result corresponding to each phrase in the sentence set; determining the foreign language search result according to the answer prediction result and the target foreign language paragraph.
2. The method for searching foreign languages of Buddhist scriptures according to claim 1, characterized in that, the performing matching of the original Buddhist scripture for the Buddhist scripture text based on the Buddhist text matching model to obtain the target original Buddhist scripture corresponding to the matching of the Buddhist scripture text includes: Obtain at least one candidate original Buddhist scripture, and splice the Buddhist scripture text with each candidate original Buddhist scripture respectively to obtain an initial Buddhist scripture set corresponding to each candidate original Buddhist scripture; Input each initial Buddhist scripture set into the Buddhist text matching model to obtain the matching score corresponding to each initial Buddhist scripture set, the matching score being the score of the matching of the Buddhist scripture text in the initial Buddhist scripture set with the candidate original Buddhist scripture; Determine the target original Buddhist scripture according to the matching score corresponding to each initial Buddhist scripture set.
3. The method for searching foreign languages of Buddhist scriptures according to claim 2, characterized in that, the determining the target original Buddhist scripture according to the matching score corresponding to each initial Buddhist scripture set includes: Determine the initial Buddhist scripture set with a matching score greater than a preset threshold as the target Buddhist scripture set; Determine the candidate original Buddhist scripture in the target Buddhist scripture set as the target original Buddhist scripture.
4. The method for searching foreign languages of Buddhist scriptures according to claim 1, characterized in that, the answer prediction result includes the first prediction probability corresponding to the phrase as the starting phrase and the second prediction probability corresponding to the phrase as the ending phrase; the determining the foreign language search result according to the answer prediction result and the target foreign language paragraph includes: Determine the target starting phrase and the target ending phrase according to the first prediction probability and the second prediction probability corresponding to each phrase; Determine the target sentence in the target foreign language paragraph according to the target starting phrase and the target ending phrase, and determine the target sentence as the foreign language search result.
5. The foreign language search method for Buddhist scriptures according to claim 4, wherein, the determination of the target start phrase and the target end phrase according to the first prediction probability and the second prediction probability corresponding to each phrase includes: determining a third prediction probability corresponding to each of the phrases, where the third prediction probability is the difference between the first prediction probability and the second prediction probability of each of the phrases; determining a fourth prediction probability corresponding to each of the phrases, where the fourth prediction probability is the difference between the second prediction probability and the first prediction probability of each of the phrases; determining the phrase corresponding to the maximum third prediction probability as the target start phrase, and determining the phrase corresponding to the maximum fourth prediction probability as the target end phrase.
6. A foreign language search device for Buddhist scriptures, wherein, for performing the foreign language search method for Buddhist scriptures according to any one of claims 1 to 5, the foreign language search device includes: a Buddhist scripture text acquisition module for acquiring the Buddhist scripture text to be searched; a Buddhist scripture original text matching module for performing matching of the Buddhist scripture text with the original Buddhist scripture text based on a Buddhist text matching model to obtain the target original Buddhist scripture text corresponding to the Buddhist scripture text; a foreign language paragraph determination module for obtaining the vernacular corresponding to the target original Buddhist scripture text and determining the target foreign language paragraph corresponding to the vernacular; a foreign language matching module for performing foreign language matching based on a reading comprehension model according to the vernacular and the target foreign language paragraph to obtain the foreign language search result corresponding to the Buddhist scripture text.
7. A computer device, wherein, the computer device includes a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program and implementing the foreign language search method for Buddhist scriptures according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the foreign language search method for Buddhist scriptures according to any one of claims 1 to 5.
Citation Information
Patent Citations
Answer text obtaining method and device, computer equipment and storage medium
CN113821609A