Lexical database construction method and apparatus

Through word-cutting processing and probability calculation, a vocabulary for indicating the occurrence of predefined events is constructed, which solves the problem of difficulty in effectively building the vocabulary in the prior art and achieves efficient and accurate event prediction.

CN114548077BActive Publication Date: 2025-05-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210193854.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2025-05-27
Estimated Expiration
2042-03-01

AI Technical Summary

Technical Problem

It is difficult to effectively construct a lexicon for indicating the occurrence of predefined events, especially when dealing with big data and historical samples.

Method used

By obtaining a collection of historical samples, the word cut process obtains words, calculates the probability information of the words based on the label information, and constructs a vocabulary for indicating the occurrence of predefined events.

Benefits of technology

It realizes efficient thesaurus construction based on historical samples, improves the accuracy of the thesaurus and the efficiency of event prediction, and saves resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548077B_ABST
    Figure CN114548077B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for constructing a thesaurus, relating to the field of computer technologies, and specifically to the fields of big data and data processing technologies. The specific implementation solution is as follows: First, obtain a set of historical samples, where the set of historical samples includes a plurality of first historical samples for constructing the thesaurus. The first historical samples include first identification information and first tag information for indicating whether a predefined event has occurred. Then, perform word segmentation processing on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information. After that, based on the first tag information corresponding to the first identification information, calculate the probability information corresponding to the plurality of first words. Finally, based on the plurality of first words and the probability information corresponding to the plurality of first words, construct a thesaurus for indicating the occurrence of a predefined event, which can determine a thesaurus for indicating the occurrence of a predefined event based on the first historical information of the first identification information in the set of historical samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, specifically to the fields of big data and data processing technologies, and particularly to a method and apparatus for building a thesaurus. Background Art

[0002] With the continuous development and progress of science and technology, some service providers can often predict the behavior of users based on the information filled in by users or the supporting materials provided, and obtain the probability of the occurrence of predefined events. For example, financial institutions will determine the probability of default of users after providing loans based on the forms filled in by users, and insurance companies will determine the probability of claim settlement after providing insurance based on the information provided by users, and so on. Summary of the Invention

[0003] The present disclosure provides a method and apparatus for building a thesaurus, an electronic device, a storage medium, and a computer program product.

[0004] According to one aspect of the present disclosure, there is provided a method for building a thesaurus, the method including: obtaining a set of historical samples, where the set of historical samples includes a plurality of first historical samples for building the thesaurus, the first historical sample includes first identification information and first label information for indicating whether a predefined event occurs, and the first identification information corresponds to the first label information; performing word segmentation processing on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information; calculating probability information corresponding to the plurality of first words based on the label information corresponding to the first identification information, where the probability information is used to indicate the probability of the occurrence of the predefined event; and building a thesaurus for indicating the occurrence of the predefined event based on the plurality of first words and the probability information corresponding to the plurality of first words.

[0005] According to another aspect of the present disclosure, there is provided a thesaurus building apparatus, the apparatus including: an obtaining module configured to obtain a set of historical samples, where the set of historical samples includes a plurality of first historical samples for building the thesaurus, the first historical sample includes first identification information and first label information for indicating whether a predefined event occurs, and the first identification information corresponds to the first label information; a word segmentation module configured to perform word segmentation processing on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information; a calculation module configured to calculate probability information corresponding to the plurality of first words based on the label information corresponding to the first identification information, where the probability information is used to indicate the probability of the occurrence of the predefined event; and a building module configured to build a thesaurus for indicating the occurrence of the predefined event based on the plurality of first words and the probability information corresponding to the plurality of first words.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, which includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned thesaurus construction method.

[0007] According to another aspect of the present disclosure, there is provided a computer-readable medium, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the above-mentioned thesaurus construction method.

[0008] According to another aspect of the present disclosure, an embodiment of the present application provides a computer program product, which includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the above-mentioned thesaurus construction method is implemented.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a flowchart of an embodiment of the thesaurus construction method according to the present disclosure;

[0012] Figure 2 is a flowchart of another embodiment of the thesaurus construction method according to the present disclosure;

[0013] Figure 3 is a flowchart of an embodiment of performing word segmentation processing on the first historical information corresponding to the first identification information according to the present disclosure;

[0014] Figure 4 is a flowchart of an embodiment of updating the thesaurus according to the present disclosure;

[0015] Figure 5 is a flowchart of an embodiment of testing the thesaurus according to the present disclosure;

[0016] Figure 6 is a schematic structural diagram of an embodiment of the thesaurus construction device according to the present disclosure;

[0017] Figure 7 is a block diagram of an electronic device for implementing the thesaurus construction method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0019] Reference Figure 1 , Figure 1 , FIG. 100 is a schematic flowchart showing an embodiment of a thesaurus construction method that can be applied to the present disclosure. The thesaurus construction method includes the following steps:

[0020] Step 110, obtaining a historical sample set. The first historical sample includes first identification information and first tag information for indicating whether a predefined event has occurred, and the first identification information corresponds to the first tag information.

[0021] In this embodiment, the execution subject (such as a server) of the thesaurus construction method may send an authorization request to the storage device of the historical samples, and the authorization request may be used to obtain an authorization permission to read the historical samples from the storage device of the historical samples. After the execution subject receives the authorization permission from the storage device, it may obtain a plurality of historical samples from the storage device, thereby obtaining the historical sample set.

[0022] Among them, the historical sample set may include a plurality of first historical samples for constructing a thesaurus. The first historical sample may be an order class sample with a completed status and may be an order record generated based on historical user requests. For example, if the historical user request is a request for the user to apply for insurance, the first historical sample may be the sample information of the insurance application order, or if the historical user request is a request for the user to apply for a loan, the first historical sample may be the sample information of the loan order.

[0023] The first historical sample may include first identification information and first label information for indicating whether a predefined event has occurred. The first identification information may be the identification information of the historical user corresponding to the first historical sample, and may be information such as the encrypted user account, mobile phone number, email, etc. that can be used to identify the user. The first label information may be a label for indicating whether a predefined event has occurred, and may include the label "yes" indicating that the predefined event has occurred or the label "no" indicating that the predefined event has not occurred, that is, it includes the label "yes" or the label "no", and the first identification information corresponds to the first label information. The predefined event may be a prediction event associated with the historical sample. For example, if the historical sample is the sample information of an insurance application order, the predefined event may be that the current insurance application order has an accident, that is, the user does not meet the insurance conditions or will trigger an accident of the current insurance policy; or, if the historical sample is the sample information of a loan order, the predefined event may be a default event, that is, the user fails to repay the loan within the agreed time period.

[0024] Preferably, each of the multiple first historical samples may include first identification information and first label information for indicating whether a predefined event has occurred. Each first historical sample includes only one label information, so that each first identification information corresponds to one type of first label information. As an example, the historical sample set includes the first historical sample A, the first historical sample B, and the first historical sample C. Among them, the first historical sample A includes the first identification information a and the first label information "yes" for indicating that the predefined event has occurred, the first historical sample B includes the first identification information b and the first label information "yes" for indicating that the predefined event has occurred, and the first historical sample C includes the first identification information c and the first label information "no" for indicating that the predefined event has occurred.

[0025] Step 120: Perform word segmentation on the first historical information corresponding to the first identification information to obtain multiple first words corresponding to the first identification information.

[0026] In this embodiment, after the above-mentioned execution entity obtains multiple first historical samples, it can obtain the first historical information corresponding to the first identification information in the first historical sample. The first historical information may be relevant information input by the user with the user's authorization and permission and associated with the first identification information. Then, the above-mentioned execution entity can receive the first historical information input by the user corresponding to the first identification information, and the first historical information may include multiple text information and other contents.

[0027] The above-mentioned execution entity can perform word segmentation on the first historical information corresponding to the first identification information, that is, it can perform word segmentation on multiple text information corresponding to the first identification information respectively, so as to obtain multiple first words corresponding to the first identification information. The first word is obtained by performing word segmentation on the first historical information corresponding to the first identification information.

[0028] Preferably, after the above-mentioned execution entity obtains a plurality of first historical samples, for each first historical sample, the first historical information corresponding to the first identification information in the first historical sample can be obtained. The first historical information can be relevant information input by the user with user authorization and permission and associated with the first identification information. Then, the above-mentioned execution entity can receive the first historical information input by the user corresponding to each first identification information. The first historical information can include multiple text information and other contents. The above-mentioned execution entity can perform word segmentation processing on the first historical information corresponding to each first identification information respectively, that is, perform word segmentation processing on the multiple text information corresponding to each first identification information respectively, so as to obtain multiple first words corresponding to each first identification information.

[0029] As an example, the historical sample set includes a first historical sample A, a first historical sample B, and a first historical sample C. Among them, the first historical sample A includes the first identification information a and the first label information "yes" for indicating the occurrence of a predefined event. The first historical sample B includes the first identification information b and the first label information "yes" for indicating the occurrence of a predefined event. The first historical sample C includes the first identification information c and the first label information "no" for indicating the occurrence of a predefined event. The above-mentioned execution entity can obtain the first historical information A1 corresponding to the first identification information a after obtaining user authorization, the first historical information B1 corresponding to the first identification information b, and the first historical information C1 corresponding to the first identification information c. Then, the above-mentioned execution entity can perform word segmentation processing on the first historical information A1 respectively to obtain multiple first words corresponding to the first identification information a, perform word segmentation processing on the first historical information B1 to obtain multiple first words corresponding to the first identification information b, and perform word segmentation processing on the first historical information C1 to obtain multiple first words corresponding to the first identification information c. Thus, the above-mentioned execution entity can obtain multiple first words corresponding to each first identification information.

[0030] Step 130, calculate the probability information corresponding to the multiple first words based on the first label information corresponding to the first identification information.

[0031] In this embodiment, after the above-mentioned execution entity obtains multiple first words corresponding to the first identification information, since the same first word can exist in different first historical information at the same time, the same first word can correspond to multiple first identification information. The above-mentioned execution entity can obtain different first identification information corresponding to the multiple first words, and determine different first label information corresponding to the multiple first words based on the first label information corresponding to the first identification information, that is, the multiple first words can correspond to the first label information corresponding to multiple different first identification information.

[0032] The above-mentioned execution entity may calculate the proportion of each type of first label information corresponding to multiple first words based on the different first label information corresponding to the determined multiple first words, that is, calculate the proportion of the label "yes" and the proportion of the label "no" in the multiple first label information corresponding to the multiple first words, and then use the proportion of the label "yes" corresponding to the multiple first words as the probability information corresponding to the multiple first words. This probability information is used to indicate the probability of the occurrence of a predefined event, and this probability may be the proportion of the above-mentioned label "yes".

[0033] Preferably, after the above-mentioned execution entity obtains multiple first words corresponding to each first identification information, since the same first word may exist in different first historical information at the same time, the same first word may correspond to multiple first identification information. The above-mentioned execution entity may obtain different first identification information corresponding to each first word, and based on the first label information corresponding to each first identification information, determine different first label information corresponding to each first word, that is, each first word may correspond to the first label information corresponding to multiple different first identification information.

[0034] The above-mentioned execution entity may calculate the proportion of each type of first label information corresponding to each first word based on the different first label information corresponding to each first word determined, that is, calculate the proportion of the label "yes" and the proportion of the label "no" in the multiple first label information corresponding to each first word, and then use the proportion of the label "yes" corresponding to each first word as the probability information corresponding to each first word. This probability information is used to indicate the probability of the occurrence of a predefined event, and this probability may be the proportion of the above-mentioned label "yes".

[0035] As an example, the historical sample set includes the first historical sample A, the first historical sample B, and the first historical sample C. Among them, the first historical sample A includes the first identification information a and the first label information "yes" indicating the occurrence of a predefined event, the first historical sample B includes the first identification information b and the first label information "yes" indicating the occurrence of a predefined event, and the first historical sample C includes the first identification information c and the first label information "no" indicating the occurrence of a predefined event. The above-mentioned execution entity may obtain the first historical information A1 corresponding to the first identification information a after obtaining user authorization, the first historical information B1 corresponding to the first identification information b, and the first historical information C1 corresponding to the first identification information c. Then, the above-mentioned execution entity may perform word segmentation processing on the first historical information A1 respectively to obtain multiple first words corresponding to the first identification information a, perform word segmentation processing on the first historical information B1 to obtain multiple first words corresponding to the first identification information b, and perform word segmentation processing on the first historical information C1 to obtain multiple first words corresponding to the first identification information c.

[0036] If multiple first words corresponding to the above first identification information a include word 1 and word 2, multiple first words corresponding to the first identification information b include word 1, word 2, and word 3, and multiple first words corresponding to the first identification information c include word 1 and word 4. The above execution entity may determine that the first label information corresponding to word 1 may include "yes" corresponding to the first identification information a, "yes" corresponding to the first identification information b, and "no" corresponding to the first identification information c. Calculate that the proportion value of the label "yes" is 66.7%, and the proportion value of the label "no" is 33.3%; the first label information corresponding to word 2 may include "yes" corresponding to the first identification information a and "yes" corresponding to the first identification information b. Calculate that the proportion value of the label "yes" is 100%; the first label information corresponding to word 3 may include "yes" corresponding to the first identification information b. Calculate that the proportion value of the label "yes" is 100%; the first label information corresponding to word 4 may include "no" corresponding to the first identification information c. Calculate that the proportion value of the label "yes" is 0%, and the proportion value of the label "no" is 100%. Then the above execution entity may determine that the probability information corresponding to word 1 is 66.7%, the probability information corresponding to word 2 is 100%, the probability information corresponding to word 3 is 100%, and the probability information corresponding to word 4 is 0%, thereby obtaining the probability information corresponding to each first word.

[0037] The above execution entity may also calculate the probability information corresponding to each first word according to other related technical means in the related art. The present disclosure does not make specific limitations on this.

[0038] Step 140, based on multiple first words and the probability information corresponding to the multiple first words, construct a word library for indicating the occurrence of a predefined event.

[0039] In this embodiment, after the above execution entity obtains the probability information corresponding to multiple first words, it may sort the probability information corresponding to the multiple first words from large to small or from small to large, select multiple probability information that meets the preset probability conditions according to the probability information of the multiple first words, and construct the first words corresponding to the multiple probability information into a word library for indicating the occurrence of a predefined event.

[0040] The above execution entity may sort the probability information corresponding to multiple first words from large to small or from small to large, select a preset number of probability information from the sorting, determine the first words corresponding to the preset number of probability information from the multiple first words, and construct these first words into a word library for indicating the occurrence of a predefined event.

[0041] Specifically, if the above probability information is sorted from large to small, the above execution entity may select a preset number of probability information ranked at the front, determine the first words corresponding to the preset number of probability information from multiple first words, and construct a word library for indicating the occurrence of a predefined event with these first words.

[0042] If the above probability information is sorted from small to large, the above execution entity may select a preset number of probability information ranked at the back, determine the first words corresponding to the preset number of probability information from multiple first words, and construct a word library for indicating the occurrence of a predefined event with these first words.

[0043] Alternatively, the above execution entity may sort the probability information corresponding to multiple first words from large to small or from small to large, select multiple probability information greater than a preset threshold from the sorting, and construct a word library for indicating the occurrence of a predefined event with the first words corresponding to the multiple probability information. The preset threshold may be set according to experience, and the present disclosure does not make specific limitations thereon.

[0044] The word library construction method provided by the embodiments of the present disclosure obtains a historical sample set, where the historical sample set includes multiple first historical samples for constructing the word library. The first historical sample includes first identification information and first label information for indicating whether a predefined event occurs, and the first identification information corresponds to the first label information. Then, word segmentation processing is performed on the first historical information corresponding to the first identification information to obtain multiple first words corresponding to the first identification information. After that, based on the first label information corresponding to the first identification information, probability information corresponding to the multiple first words is calculated, and the probability information is used to indicate the probability of the predefined event occurring. Finally, based on the multiple first words and the probability information corresponding to the multiple first words, a word library for indicating the occurrence of the predefined event is constructed. It can determine a word library for indicating the occurrence of a predefined event based on the first historical information of the first identification information in the historical sample set, improve the efficiency of word library construction, make the constructed word library closer to the historical information corresponding to the first identification information, and the obtained word library can more accurately indicate the occurrence of the predefined event. Thus, it is not necessary to construct a prediction model to predict the predefined event, and only the above word library is needed to determine the probability information of the predefined event occurring, saving resources and improving the efficiency of event prediction.

[0045] See Figure 2 , Figure 2 shows a flowchart of another embodiment of the word library construction method. The word library construction method may include the following steps:

[0046] Step 210, obtain a historical sample set. The first historical sample includes first identification information and first label information for indicating whether a predefined event occurs, and the first identification information corresponds to the first label information.

[0047] Step 210 of this embodiment may be executed in a manner similar to that of step 110 in the Figure 1 illustrated embodiment, and details are not described herein.

[0048] Step 220: Perform word segmentation on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information.

[0049] Step 220 of this embodiment may be executed in a manner similar to that of step 120 in the Figure 1 illustrated embodiment, and details are not described herein.

[0050] Step 230: Calculate probability information corresponding to the plurality of first words based on the tag information corresponding to the first identification information.

[0051] Step 230 of this embodiment may be executed in a manner similar to that of step 130 in the Figure 1 illustrated embodiment, and details are not described herein.

[0052] Step 240: Obtain the labeled words used to indicate the occurrence of a predefined event.

[0053] In this embodiment, the above-mentioned execution entity may receive the labeled words input by the labeling personnel through a network. The labeled words may be words labeled by the labeling personnel according to the preset conditions in the historical samples and the occurrence conditions of the predefined events. For example, if the historical sample is the sample information of an insurance application order, the labeled words may be words related to the insurance conditions in the insurance application order or words that can trigger the occurrence of an insurance claim in the insurance application order; or, if the historical sample is the sample information of a loan order, the labeled words may be words related to the loan conditions in the loan order, etc.

[0054] Step 250: Determine a set of words that meet the preset conditions from the plurality of first words based on the probability information corresponding to the plurality of first words.

[0055] In this embodiment, after the above-mentioned execution entity obtains the probability information corresponding to the plurality of first words, it may perform probability screening on the plurality of first words according to the probability information of the plurality of first words, and may select the probability information greater than the preset threshold. The plurality of first words corresponding to these probability information greater than the preset threshold form a set of words that meet the preset conditions. Then, the preset condition is that the probability information of the first words in the set of words is greater than the preset threshold. The preset threshold may be set according to experience, and the present disclosure does not make specific limitations thereto.

[0056] Step 260: Construct a word library for indicating the occurrence of a predefined event based on the labeled words and the set of words.

[0057] In this embodiment, after the above-mentioned execution entity obtains the labeled words and the word set, it can add the labeled words to the word set to construct a word library for indicating the occurrence of a predefined event.

[0058] Alternatively, the above-mentioned execution entity can compare the labeled words with the first words in the word set to determine whether there are labeled words in the word set, and add the non-existent labeled words to the word set to construct a word library for indicating the occurrence of a predefined event.

[0059] In this embodiment, by further obtaining the labeled words for indicating the occurrence of a predefined event, a word library can be constructed based on the labeled words and multiple first words, making the obtained word library more comprehensive and accurate, thereby improving the accuracy of event prediction based on the word library.

[0060] Reference Figure 3 , Figure 3 shows a flowchart of an embodiment of performing word segmentation processing on the first historical information corresponding to the first identification information, that is, the above-mentioned step 120. Performing word segmentation processing on the first historical information corresponding to the first identification information to obtain multiple first words corresponding to the first identification information may include the following steps:

[0061] Step 310, obtain the first historical information corresponding to the first identification information within a preset time period.

[0062] In this step, after the above-mentioned execution entity obtains multiple first historical samples, it can obtain the first historical information corresponding to the first identification information in the first historical samples. The first historical information may be relevant information input by the user with user authorization and permission and associated with the first identification information. The first historical information may be the historical information within a preset time period corresponding to the first identification information. The first historical information may include multiple text information and other contents. The preset time period may be the historical time period corresponding to the first identification information, such as the past six months, etc. The present disclosure does not make specific limitations on this.

[0063] Preferably, after the above-mentioned execution entity obtains multiple first historical samples, for each first historical sample, it can separately obtain the first historical information corresponding to the first identification information in each first historical sample.

[0064] Step 320, perform word segmentation processing on the first historical information corresponding to the first identification information to obtain multiple initial words corresponding to the first identification information.

[0065] In this step, the above-mentioned execution entity can perform word segmentation on the first historical information corresponding to the first identification information, that is, it can perform word segmentation on multiple text information corresponding to the first identification information, so as to obtain multiple first words corresponding to the first identification information, and the first word is obtained by performing word segmentation on the first historical information corresponding to the first identification information.

[0066] Preferably, the above-mentioned execution entity can perform word segmentation on the first historical information corresponding to each first identification information respectively, that is, it can perform word segmentation on multiple text information corresponding to each first identification information respectively, so as to obtain multiple first words corresponding to each first identification information, and the first word is obtained by performing word segmentation on the first historical information corresponding to the first identification information.

[0067] As an example, the above-mentioned execution entity can obtain the first historical information A1 corresponding to the first identification information a after obtaining user authorization, the first historical information B1 corresponding to the first identification information b, and the first historical information C1 corresponding to the first identification information c. Then, the above-mentioned execution entity can perform word segmentation on the first historical information A1 respectively to obtain multiple initial words corresponding to the first identification information a, perform word segmentation on the first historical information B1 to obtain multiple initial words corresponding to the first identification information b, and perform word segmentation on the first historical information C1 to obtain multiple initial words corresponding to the first identification information c. Thus, the above-mentioned execution entity can obtain multiple initial words corresponding to each first identification information.

[0068] Step 330: Screen the multiple initial words corresponding to the first identification information based on the part of speech of the words to obtain multiple first words corresponding to the first identification information.

[0069] In this step, after the above-mentioned execution entity obtains the multiple initial words corresponding to the first identification information, it can perform part-of-speech analysis on the multiple initial words, delete the common stop words in the multiple initial words, and obtain the multiple first words corresponding to the screened first identification information.

[0070] Preferably, after the above-mentioned execution entity obtains the multiple initial words corresponding to each first identification information respectively, it can perform part-of-speech analysis on each initial word, delete the common stop words in the multiple initial words, and obtain the multiple first words corresponding to each screened first identification information.

[0071] In this embodiment, by screening the initial words corresponding to the first identification information, the obtained first words are more accurate, can better display the characteristic information corresponding to the first identification information, and improve the accuracy of the first words.

[0072] Reference Figure 4 , Figure 4The flowchart shows an embodiment of updating a thesaurus, including the following steps:

[0073] Step 410, obtain new historical samples. The new historical samples include new identification information and new label information for indicating whether a predefined event has occurred. The new identification information corresponds to the new label information.

[0074] In this step, the above-mentioned execution entity can also obtain new historical samples from a storage device based on the authorization permission of the storage device. Among them, the new historical samples can include new identification information and new label information for indicating whether a predefined event has occurred. The new identification information can be the identification information of the historical user corresponding to the new historical sample, and can be information such as the encrypted user account, mobile phone number, email, etc. that can be used to identify the user. The new label information can be a label for indicating whether a predefined event has occurred, and can include a label "yes" indicating that the predefined event has occurred or a label "no" indicating that the predefined event has not occurred. The new identification information corresponds to the new label information. The predefined event can be a prediction event associated with the historical sample. For example, if the historical sample is the sample information of an insurance application order, the predefined event can be that the current insurance application order has an accident, that is, the user does not meet the insurance conditions or will trigger an accident of the current insurance policy; or, if the historical sample is the sample information of a loan order, the predefined event can be a default event, that is, the user has not repaid the loan within the agreed time period.

[0075] Step 420, perform word segmentation on the new historical information corresponding to the new identification information to obtain multiple new words corresponding to the new identification information.

[0076] In this step, after the above-mentioned execution entity obtains the new historical samples, it can obtain the new historical information corresponding to the new identification information. The new historical information can be relevant information input by the user with the authorization permission of the user and associated with the new identification information. Then, the above-mentioned execution entity can receive the new historical information input by the user corresponding to the new identification information. The new historical information can include multiple text information and other contents.

[0077] The above-mentioned execution entity can perform word segmentation on the new historical information corresponding to the new identification information, that is, it can perform word segmentation on multiple text information corresponding to the new identification information, so as to obtain multiple new words corresponding to the new identification information. The new words are obtained by performing word segmentation on the new historical information corresponding to the new identification information.

[0078] Step 430, calculate the probability information corresponding to multiple new words and multiple first words based on the new label information corresponding to the new identification information and the label information corresponding to the first identification information.

[0079] In this step, after the above-mentioned execution entity obtains multiple new words corresponding to the new identification information, since the same word can exist in different historical information at the same time, the same word can correspond to multiple identification information. The above-mentioned execution entity can determine different first label information and / or new label information corresponding to the multiple new words, and different first label information and / or new label information corresponding to the multiple first words.

[0080] The above-mentioned execution entity can calculate the proportion of the label "yes" and the proportion of the label "no" corresponding to each of the multiple new words and each first word according to the different first label information and / or new label information corresponding to the multiple new words, and the different first label information and / or new label information corresponding to the multiple first words. Then, the proportion of the label "yes" corresponding to the multiple new words and each first word is used as the probability information corresponding to each first word. This probability information is used to indicate the probability of the occurrence of a predefined event, and this probability can be the proportion of the above-mentioned label "yes".

[0081] Step 440, update the word library based on the probability information corresponding to the multiple new words and the multiple first words to obtain an updated word library.

[0082] In this step, after the above-mentioned execution entity obtains the probability information corresponding to the multiple new words and the multiple first words, it can sort the probability information corresponding to the multiple new words and the multiple first words from large to small or from small to large, select multiple probability information that meets the preset probability conditions according to the probability information corresponding to the multiple new words and the multiple first words, and update the word library with the first words and / or new words corresponding to these multiple probability information to obtain an updated word library.

[0083] In this embodiment, by updating the word library based on the new historical samples, the words in the word library used to indicate the occurrence of a predefined event can be updated in real time, so that the word library can be updated in real time, improving the accuracy of the word library.

[0084] Reference Figure 5 , Figure 5 shows a flowchart of an embodiment of a test word library, which may include the following steps:

[0085] Step 510, in response to obtaining a word library used to indicate the occurrence of a predefined event, perform word segmentation processing on the second historical information corresponding to the second identification information to obtain multiple second words corresponding to the second identification information.

[0086] The above-mentioned historical sample set also includes second historical samples for testing the above-mentioned word library. The second historical samples may include second identification information and second label information used to indicate whether a predefined event occurs, and the second identification information corresponds to the second label information.

[0087] In this step, after the above-mentioned execution entity obtains the thesaurus indicating the occurrence of a predefined event, it can obtain the second historical information corresponding to the second identification information in the second historical sample. The second historical information may be relevant information input by the user with the user's authorization and associated with the second identification information. Then, the above-mentioned execution entity can receive the second historical information input by the user corresponding to the second identification information, and the second historical information may include multiple text information and other contents. The above-mentioned execution entity can perform word segmentation processing on the second historical information corresponding to the second identification information, that is, it can perform word segmentation processing on the multiple text information corresponding to the second identification information, so as to obtain multiple second words corresponding to the second identification information. The second word is obtained by performing word segmentation on the second historical information corresponding to the second identification information.

[0088] Preferably, after the above-mentioned execution entity obtains the thesaurus indicating the occurrence of a predefined event, for each second identification information among the multiple second identification information, it can obtain the second historical information corresponding to the second identification information in the second historical sample. The second historical information may be relevant information input by the user with the user's authorization and associated with the second identification information. Then, the above-mentioned execution entity can receive the second historical information input by the user corresponding to each second identification information, and the second historical information may include multiple text information and other contents.

[0089] For each second identification information, the above-mentioned execution entity can perform word segmentation processing on the second historical information corresponding to the second identification information, that is, it can perform word segmentation processing on the multiple text information corresponding to the second identification information, so as to obtain multiple second words corresponding to the second identification information. The second word is obtained by performing word segmentation on the second historical information corresponding to the second identification information.

[0090] As an example, the historical sample set includes a second historical sample D, a second historical sample E, and a second historical sample F. Among them, the second historical sample D includes second identification information d and second label information "yes" indicating the occurrence of a predefined event, the second historical sample E includes second identification information e and second label information "yes" indicating the occurrence of a predefined event, and the second historical sample F includes second identification information f and second label information "no" indicating the occurrence of a predefined event. The above-mentioned execution entity can obtain the second historical information D1 corresponding to the second identification information d after obtaining user authorization, the second historical information E1 corresponding to the second identification information e, and the second historical information F1 corresponding to the second identification information f. Then, the above-mentioned execution entity can perform word segmentation processing on the second historical information D1 respectively to obtain multiple second words corresponding to the second identification information d, perform word segmentation processing on the second historical information E1 to obtain multiple second words corresponding to the second identification information e, and perform word segmentation processing on the second historical information F1 to obtain multiple second words corresponding to the second identification information f. Thus, the above-mentioned execution entity can obtain multiple second words corresponding to each second identification information.

[0091] Step 520: Based on the multiple second words corresponding to the second identification information and the thesaurus, determine the generated label information corresponding to the second identification information.

[0092] In this step, after the above-mentioned execution entity obtains the multiple second words corresponding to the second identification information, it can compare the multiple second words corresponding to the second identification information with the words in the thesaurus, determine the number of words in the multiple second words that include the words in the thesaurus, and calculate the proportion value of these words in the multiple second words. The generated label information corresponding to the second identification information can be determined according to this proportion value. The generated label information can include the label "yes" indicating the occurrence of a predefined event or the label "no" indicating the non-occurrence of a predefined event.

[0093] Specifically, if the proportion value exceeds the preset threshold, it is determined that the generated label information corresponding to the second identification information is the label "yes"; if the proportion value does not exceed the preset threshold, it is determined that the generated label information corresponding to the second identification information is the label "no".

[0094] Therefore, the above-mentioned execution entity can use the technical means of data statistics in related technologies to determine the generated label information corresponding to the second identification information based on the multiple second words corresponding to the second identification information and the thesaurus.

[0095] Step 530: Based on the generated label information corresponding to the second identification information and the second label information corresponding to the second identification information, generate a test result for characterizing the effectiveness of the thesaurus.

[0096] In this step, after the above-mentioned execution entity obtains the generated tag information corresponding to the second identification information according to the above-mentioned technical means, it can compare the generated tag information corresponding to the second identification information with its corresponding second tag information respectively to determine whether they are consistent, so as to determine whether the prediction result of the predefined event based on the thesaurus is accurate, and further can be used to represent the test result of the effectiveness of the thesaurus.

[0097] The above-mentioned execution entity can count the number of the same generated tag information and second tag information, calculate the proportion value it occupies. If the proportion value exceeds the preset threshold, it can generate a test result for representing that the thesaurus can effectively predict the predefined event; if the proportion value does not exceed the preset threshold, it can generate a test result for representing that the thesaurus cannot effectively predict the predefined event, and the number of the first historical samples can be increased to update and improve the thesaurus.

[0098] Therefore, the above-mentioned execution entity can adopt the technical means of data statistics in the related technology to generate a test result for representing the effectiveness of the thesaurus based on the generated tag information corresponding to the second identification information and the second tag information corresponding to the second identification information.

[0099] In this embodiment, testing the thesaurus with the second historical sample can test the effectiveness of the thesaurus, make the words in the thesaurus more accurate, and improve the accuracy of the thesaurus and the accuracy of predicting predefined events.

[0100] Reference Figure 6 , as an implementation of the methods shown in the above-mentioned figures, the present disclosure provides an embodiment of a thesaurus construction device. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0101] As Figure 6 shown, the thesaurus construction device 600 in this embodiment includes: an acquisition module 610, a word segmentation module 620, a calculation module 630, and a construction module 640.

[0102] Among them, the acquisition module 610 is configured to acquire a historical sample set, where the historical sample set includes a plurality of first historical samples for constructing the thesaurus. The first historical sample includes first identification information and first tag information for indicating whether a predefined event occurs, and the first identification information corresponds to the first tag information;

[0103] The word segmentation module 620 is configured to perform word segmentation processing on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information;

[0104] A calculation module 630, configured to calculate probability information corresponding to multiple first words based on first tag information corresponding to first identification information, where the probability information is used to indicate the probability of a predefined event occurring;

[0105] A construction module 640, configured to construct a word library for indicating the occurrence of a predefined event based on multiple first words and probability information corresponding to the multiple first words.

[0106] In some optional implementations of this embodiment, the acquisition module 610 is further configured to: acquire labeled words for indicating the occurrence of a predefined event; and the construction module 640 is further configured to: determine a set of words that meet a preset condition from the multiple first words based on the probability information corresponding to the multiple first words; construct a word library for indicating the occurrence of a predefined event based on the labeled words and the set of words.

[0107] In some optional implementations of this embodiment, the word segmentation module 620 is further configured to: acquire first historical information corresponding to the first identification information within a preset time period; perform word segmentation processing on the first historical information corresponding to the first identification information to obtain multiple initial words corresponding to the first identification information; screen the multiple initial words corresponding to the first identification information based on the part of speech of the words to obtain multiple first words corresponding to the first identification information.

[0108] In some optional implementations of this embodiment, the apparatus further includes an update module; the acquisition module 610 is further configured to: acquire new historical samples, where the new historical samples include new identification information and new tag information for indicating whether a predefined event occurs, and the new identification information corresponds to the new tag information; the word segmentation module 620 is further configured to: perform word segmentation processing on the new historical information corresponding to the new identification information to obtain multiple new words corresponding to the new identification information; the calculation module 630 is further configured to: calculate event probability information corresponding to the multiple new words and the multiple first words based on the new tag information corresponding to the new identification information and the tag information corresponding to the first identification information; the update module is configured to: update the word library based on the event probability information corresponding to the multiple new words and the multiple first words to obtain an updated word library.

[0109] In some alternative embodiments of the present embodiment, the historical sample set further includes a second historical sample for testing the thesaurus. The second historical sample includes second identification information and second tag information for indicating whether a predefined event has occurred. The second identification information corresponds to the second tag information. Further, the apparatus further includes a determination module and a generation module. The word segmentation module 620 is further configured to: in response to obtaining a thesaurus for indicating the occurrence of a predefined event, perform word segmentation processing on the second historical information corresponding to the second identification information to obtain a plurality of second words corresponding to the second identification information; the determination module is configured to: based on the plurality of second words corresponding to the second identification information and the thesaurus, determine the generated tag information corresponding to the second identification information, where the generated tag information is used to indicate whether a predefined event has occurred; the generation module is configured to: based on the generated tag information corresponding to the second identification information and the second tag information corresponding to the second identification information, generate a test result for characterizing the effectiveness of the thesaurus.

[0110] The thesaurus construction apparatus provided by the embodiments of the present disclosure obtains a historical sample set, where the historical sample set includes a plurality of first historical samples for constructing the thesaurus. The first historical sample includes first identification information and first tag information for indicating whether a predefined event has occurred. The first identification information corresponds to the first tag information. Then, word segmentation processing is performed on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information. After that, based on the first tag information corresponding to the first identification information, probability information corresponding to the plurality of first words is calculated. The probability information is used to indicate the probability of the occurrence of the predefined event. Finally, based on the plurality of first words and the probability information corresponding to the plurality of first words, a thesaurus for indicating the occurrence of the predefined event is constructed. It can determine a thesaurus for indicating the occurrence of a predefined event based on the first historical information of the first identification information in the historical sample set, improve the efficiency of thesaurus construction, make the constructed thesaurus closer to the historical information corresponding to the first identification information, and enable the obtained thesaurus to more accurately indicate the occurrence of the predefined event. Thus, there is no need to construct a prediction model to predict the predefined event, and only the above-mentioned thesaurus is needed to determine the probability information of the occurrence of the predefined event, saving resources and improving the efficiency of event prediction.

[0111] Those skilled in the art can understand that the above-mentioned apparatus further includes some other well-known structures, such as a processor, a memory, etc. In order not to unnecessarily obscure the embodiments of the present disclosure, these well-known structures are not shown in Figure 6 it.

[0112] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0113] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0114] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0115] As Figure 7 shown, the electronic device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 805 is also connected to the bus 704.

[0116] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0117] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the thesaurus construction method. For example, in some embodiments, the thesaurus construction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the thesaurus construction method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the thesaurus construction method by any other suitable means (e.g., by means of firmware).

[0118] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0119] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0121] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0122] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0123] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0124] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0125] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for constructing a thesaurus, including: Obtaining a set of historical samples, wherein the set of historical samples includes a plurality of first historical samples for constructing the thesaurus, the first historical samples include first identification information and first label information for indicating whether a predefined event has occurred, and the first identification information corresponds to the first label information; Performing word segmentation on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information; Calculating probability information corresponding to the plurality of first words based on the first label information corresponding to the first identification information, where the probability information is used to indicate the probability of occurrence of a predefined event; Constructing a thesaurus for indicating the occurrence of a predefined event based on the plurality of first words and the probability information corresponding to the plurality of first words.

2. The method according to claim 1, wherein, The method further includes: obtaining labeled words for indicating the occurrence of a predefined event; and The constructing a thesaurus for indicating the occurrence of a predefined event based on the plurality of first words and the probability information corresponding to the plurality of first words includes: Determining a set of words that meet preset conditions from the plurality of first words based on the probability information corresponding to the plurality of first words; Constructing a thesaurus for indicating the occurrence of a predefined event based on the labeled words and the set of words.

3. The method according to claim 1, wherein, The performing word segmentation on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information includes: Obtaining the first historical information corresponding to the first identification information within a preset time period; Performing word segmentation on the first historical information corresponding to the first identification information to obtain a plurality of initial words corresponding to the first identification information; Screening the plurality of initial words corresponding to the first identification information based on the part of speech of the words to obtain a plurality of first words corresponding to the first identification information.

4. The method according to claim 1, wherein, The method further includes: Obtaining new historical samples, wherein the new historical samples include new identification information and new label information for indicating whether a predefined event has occurred, and the new identification information corresponds to the new label information; Performing word segmentation on the new historical information corresponding to the new identification information to obtain a plurality of new words corresponding to the new identification information; Calculating probability information corresponding to the plurality of new words and the plurality of first words based on the new label information corresponding to the new identification information and the label information corresponding to the first identification information; Updating the thesaurus based on the probability information corresponding to the plurality of new words and the plurality of first words to obtain an updated thesaurus.

5. The method according to any one of claims 1-4, wherein, The set of historical samples further includes second historical samples for testing the thesaurus, the second historical samples include second identification information and second label information for indicating whether a predefined event has occurred, and the second identification information corresponds to the second label information; and, the method further includes: In response to obtaining a thesaurus for indicating the occurrence of a predefined event, perform word segmentation processing on the second historical information corresponding to the second identification information to obtain a plurality of second words corresponding to the second identification information; Based on the plurality of second words corresponding to the second identification information and the thesaurus, determine the generated label information corresponding to the second identification information, where the generated label information is used to indicate whether the predefined event occurs; Based on the generated label information corresponding to the second identification information and the second label information corresponding to the second identification information, generate a test result for characterizing the effectiveness of the thesaurus.

6. A thesaurus construction device, including: An acquisition module, configured to acquire a historical sample set, where the historical sample set includes a plurality of first historical samples for constructing a thesaurus, and the first historical sample includes first identification information and first label information for indicating whether a predefined event occurs, and the first identification information corresponds to the first label information; A word segmentation module, configured to perform word segmentation processing on the first historical information corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information; A calculation module, configured to calculate probability information corresponding to the plurality of first words based on the first label information corresponding to the first identification information, where the probability information is used to indicate the probability of occurrence of a predefined event; A construction module, configured to construct a thesaurus for indicating the occurrence of a predefined event based on the plurality of first words and the probability information corresponding to the plurality of first words.

7. The device according to claim 6, wherein, the acquisition module is further configured to: acquire labeled words for indicating the occurrence of a predefined event; and the construction module is further configured to: Based on the probability information corresponding to the plurality of first words, determine a set of words that meet preset conditions from the plurality of first words; Construct a thesaurus for indicating the occurrence of a predefined event based on the labeled words and the set of words.

8. The device according to claim 6, wherein, the word segmentation module is further configured to: Acquire the first historical information corresponding to the first identification information within a preset time period; Perform word segmentation processing on the first historical information corresponding to the first identification information to obtain a plurality of initial words corresponding to the first identification information; Based on the part of speech of the words, screen the plurality of initial words corresponding to the first identification information to obtain a plurality of first words corresponding to the first identification information.

9. The device according to claim 6, wherein, the device further includes an update module; the acquisition module is further configured to: acquire new historical samples, where the new historical samples include new identification information and new label information for indicating whether a predefined event occurs, and the new identification information corresponds to the new label information; the word segmentation module is further configured to: perform word segmentation processing on the new historical information corresponding to the new identification information to obtain a plurality of new words corresponding to the new identification information; The computing module is further configured to: calculate event probability information corresponding to the multiple new words and the multiple first words based on the new tag information corresponding to the new identification information and the tag information corresponding to the first identification information; The updating module is configured to: update the word library based on the event probability information corresponding to the multiple new words and the multiple first words to obtain an updated word library.

10. The apparatus according to claims 6-9, wherein, The historical sample set further includes a second historical sample for testing the word library, the second historical sample includes second identification information and second tag information for indicating whether a predefined event occurs, and the second identification information corresponds to the second tag information; and, the apparatus further includes a determination module and a generation module; The word segmentation module is further configured to: in response to obtaining a word library for indicating the occurrence of a predefined event, perform word segmentation processing on the second historical information corresponding to the second identification information to obtain multiple second words corresponding to the second identification information; The determination module is configured to: determine generation tag information corresponding to the second identification information based on the multiple second words corresponding to the second identification information and the word library, where the generation tag information is used to indicate whether a predefined event occurs; The generation module is configured to: generate a test result for characterizing the effectiveness of the word library based on the generation tag information corresponding to the second identification information and the second tag information corresponding to the second identification information.

11. An electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.

13. A computer program product comprising computer programs / instructions, characterized in that, When the computer programs / instructions are executed by a processor, the steps of the method according to any one of claim 1 are implemented.

Citation Information

Patent Citations

  • System and method of searching key word for determining event happening

    CN101261637A

  • Event trigger word recognition method and device

    CN104598510A