Methods, Apparatus, Electronic Devices, and Storage Media for Text Processing
Through the preset vocabulary and knowledge base, the root database is expanded, the filtering and word segmentation combination is combined, and the business scenario association sequence is spliced and scored, the problem of time-consuming and inaccurate labeling of business entities in the existing technology is solved, and efficient and accurate business entity recognition is achieved.
Patent Information
- Application Number
- CN202111041964.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-09-07
AI Technical Summary
In the prior art, it is difficult to unify standards for manual labeling of business entities, which consumes time and reduces the accuracy of business identification labeling, thereby affecting the identification accuracy of business entities in the text.
By obtaining the preset lexicon, sequential annotation of the text in the knowledge base, the root database is calculated, and word segmentation is combined based on the root database filtering statement, combined with the business scenario association sequence splicing word segmentation combinations, and calling the calculation model to calculate the score to determine the business entity.
It improves the recognition efficiency and accuracy of business entities, reduces the time and errors of manual annotation, and enhances the automation and accuracy of text processing.
Smart Images

Figure CN113743115B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, and a storage medium for text processing. Background Art
[0002] In the e-commerce field, the after-sales link plays a great role in user experience and user retention. At present, in the after-sales consultation scenarios of users, a large number of processing processes require the assistance of artificial intelligence and usually involve text processing, such as identifying business entities in text. Since each business entity is usually related to a business scenario and sometimes lacks generality, in order to accurately identify business entities in text, in the prior art, it is necessary for humans to segment the corpus in related fields and label each business entity according to business understanding and the definition of business entities, and then extract the business entities in the text based on the labeled business entities. However, this manual labeling method is difficult to unify standards, not only takes a long time, but also reduces the accuracy of business recognition labeling, thereby reducing the accuracy of business entity recognition in text. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, an apparatus, an electronic device, and a storage medium for text processing, which can solve the problems that it is difficult to unify standards in the manual labeling method, not only takes a long time, but also reduces the accuracy of business recognition labeling and the accuracy of business entity recognition in text.
[0004] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for text processing is provided.
[0005] A method for text processing according to an embodiment of the present invention includes: obtaining a preset word library to perform sequence labeling on each text in a knowledge base, inputting the labeled text into a preset recognition model, and calculating a root word library; for each root word in the root word library, screening sentences including the root word from the knowledge base, performing word segmentation on each sentence, and combining the word segmentation results to determine a word segmentation combination corresponding to each sentence, where the word segmentation combination includes the root word; based on the business scenario to which the word segmentation combination corresponds and a preset business scenario association order, splicing the word segmentation combinations including the same root word to obtain a splicing result; calling a preset calculation model to calculate the effective scores of each splicing result, and then determining business entities from the splicing results based on the effective scores to obtain a business entity library; and identifying business entities in the received text to be processed based on the business entity library.
[0006] In one embodiment, performing word segmentation on each sentence and combining the word segmentation results to determine a word segmentation combination corresponding to each sentence includes:
[0007] For each statement, tokenize the statement, use the root word as the starting word, and sequentially use the tokens in the statement that are located after the root word as the ending words. Based on the starting word and the ending words, determine token combinations from the statement.
[0008] In yet another embodiment, based on the business scenario to which the statement corresponding to the token combination belongs and the preset business scenario association order, splice the token combinations including the same root word to obtain a splicing result, including:
[0009] For the token combination corresponding to each statement, sort based on the order of the ending words in the corresponding statement to generate a token array corresponding to each statement;
[0010] Based on the business scenario to which the token array corresponds, determine the business scenario corresponding to the token array;
[0011] Filter the target token arrays corresponding to the same root word, and splice the token combinations in the same position in the target token arrays corresponding to each business scenario based on the business scenario association order to obtain a splicing result.
[0012] In yet another embodiment, calling a preset calculation model to calculate the effective score of each splicing result, including:
[0013] For each splicing result, calculate the feature vector of the splicing result, call a preset calculation model, extract the features of the splicing result based on the feature vector, and then calculate the floating-point value corresponding to the feature, and determine the floating-point value as the effective score of the splicing result.
[0014] In yet another embodiment, based on the effective score, determine business entities from the splicing results and update them to the business entity library, including:
[0015] Judge whether the effective score of the splicing result is greater than a preset score threshold;
[0016] If so, determine the token combination with the last splicing order in the splicing result as the business entity and update it to the business entity library; if not, do not process the splicing result.
[0017] In yet another embodiment, before screening the statements including the root word from the knowledge base, it further includes:
[0018] Obtain the historical search terms of the knowledge base, and add the historical keywords and the preset thesaurus to the root word library.
[0019] In yet another embodiment, the text to be processed includes the consultation text sent by the user;
[0020] Identifying business entities in the received text to be processed based on the business entity library includes:
[0021] Receiving consultation information sent by a user, and determining the consultation text corresponding to the consultation information;
[0022] Matching the business entity library with the consultation text, identifying the business entities included in the consultation text, querying the corresponding response file based on the business entities included in the consultation text, and then replying to the consultation information.
[0023] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a text processing device.
[0024] A text processing device according to an embodiment of the present invention includes: a calculation unit, configured to obtain a preset word library to perform sequence annotation on each text in the knowledge base, input the annotated text into a preset recognition model, and calculate a root word library; a determination unit, configured to, for each root word in the root word library, screen out the sentences including the root word from the knowledge base, perform word segmentation on each sentence, and combine the word segmentation results to determine the word segmentation combination corresponding to each sentence, where the word segmentation combination includes the root word; a splicing unit, configured to splice the word segmentation combinations including the same root word based on the business scenario to which the word segmentation combination corresponds and a preset business scenario association order to obtain a splicing result; the calculation unit is further configured to call a preset calculation model to calculate the effective scores of the splicing results, and then determine business entities from the splicing results based on the effective scores to update the business entity library; a processing unit, configured to identify business entities in the received text to be processed based on the business entity library.
[0025] In one embodiment, the calculation unit is specifically configured to:
[0026] For each sentence, perform word segmentation on the sentence, use the root word as the starting word, and sequentially use the words after the root word in the sentence as the ending words, and determine the word segmentation combination from the sentence based on the starting word and the ending words.
[0027] In yet another embodiment, the splicing unit is specifically configured to:
[0028] Sort the word segmentation combinations corresponding to each sentence based on the order of the ending words in the corresponding sentence to generate a word segmentation array corresponding to each sentence;
[0029] Determine the business scenario corresponding to the word segmentation array based on the business scenario to which the word segmentation array corresponds;
[0030] Screen the target word segmentation arrays corresponding to the same root word, and splice the word segmentations at the same positions in the target word segmentation arrays corresponding to each business scenario based on the associated order of the business scenarios to obtain a splicing result.
[0031] In another embodiment, the calculation unit is specifically configured to:
[0032] For each splicing result, calculate the feature vector of the splicing result, call a preset calculation model, extract the features of the splicing result based on the feature vector, and then calculate the floating-point value corresponding to the feature, and determine the floating-point value as the effective score of the splicing result.
[0033] In another embodiment, the processing unit is specifically configured to:
[0034] Judge whether the effective score of the splicing result is greater than a preset score threshold;
[0035] If so, determine the word segment combination with the last splicing order in the splicing result as the business entity and update it to the business entity library; if not, do not process the splicing result.
[0036] In another embodiment, the device further includes:
[0037] An adding unit, configured to obtain the historical search terms of the knowledge base and add the historical keywords and the preset thesaurus to the root word library.
[0038] In another embodiment, the text to be processed includes the consultation text sent by the user; the processing unit is specifically configured to:
[0039] Receive the consultation information sent by the user and determine the consultation text corresponding to the consultation information;
[0040] Match the business entity library with the consultation text, identify the business entities included in the consultation text, query the corresponding response file based on the business entities included in the consultation text, and then reply to the consultation information.
[0041] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided an electronic device.
[0042] An electronic device according to an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the text processing method provided by the embodiments of the present invention.
[0043] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a computer-readable medium.
[0044] A computer-readable medium according to an embodiment of the present invention stores a computer program thereon, and when the program is executed by a processor, it implements the text processing method provided by the embodiment of the present invention.
[0045] One embodiment of the above invention has the following advantages or beneficial effects: In the embodiment of the present invention, based on a preset word library, sequence labeling can be performed on the text in the knowledge base to calculate a root word library through an identification model. The words in the root word library represent the root words of business entities; based on the root word library, sentences including the root words can be screened from the knowledge base, and then each sentence can be segmented and the segmentation results can be combined to determine the segmentation combination corresponding to each sentence that includes the root word; then, based on the business scenario to which the segmentation combination corresponds and the preset business scenario association order, the segmentation combinations including the same root word can be spliced to obtain a splicing result. Furthermore, through a calculation model, the scores of each splicing result can be calculated, and based on the valid scores, business entities can be determined from the splicing results to obtain a business entity library, so as to identify business entities in the received text to be processed based on the business entity library. In the embodiment of the present invention, after expanding the root word library based on the preset word library and the knowledge base, sentences including the root words in the root word library can be screened out and the segmentation combinations corresponding to each sentence can be determined. This is equivalent to expanding the root words based on the text in the knowledge base to obtain the segmentation combinations corresponding to each scenario. Then, based on the business scenario association order and the business scenario to which each segmentation combination corresponds, the segmentation combinations including the same root word can be spliced to obtain a splicing result. This is equivalent to splicing the segmentation combinations in combination with the association between scenarios, so that the reasoning relationship between business scenarios is reflected in the splicing result. Then, based on the valid scores of the splicing results, business entities are determined. In this way, first, the root word library is expanded from the preset word library, then the segmentation combinations of each business scenario are expanded based on the root word library, and then the business entities are determined in combination with the reasoning relationship between business scenarios to obtain a business entity library for identifying business entities in the received text to be processed. This not only improves the efficiency of determining business entities, but also improves the accuracy and comprehensiveness of determining business entities, and further increases the accuracy of business entity recognition in the text.
[0046] The further effects of the above non-conventional optional methods will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:
[0048] Figure 1 is a schematic diagram of a main process of the text processing method according to an embodiment of the present invention;
[0049] Figure 2 is a schematic diagram of a process for calculating a root word library according to an embodiment of the present invention;
[0050] Figure 3 It is a schematic diagram of another main process of the text processing method according to an embodiment of the present invention;
[0051] Figure 4 It is a schematic diagram of a process for determining a business entity according to an embodiment of the present invention;
[0052] Figure 5 It is a schematic diagram of the main units of the text processing device according to an embodiment of the present invention;
[0053] Figure 6 It is another exemplary system architecture diagram to which the embodiment of the present invention can be applied;
[0054] Figure 7 It is a schematic structural diagram of a computer system suitable for implementing the embodiment of the present invention. Detailed implementation manners
[0055] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding. It should be considered that they are merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0056] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0057] The embodiment of the present invention provides a text processing method, which can be executed by a text processing device, such as Figure 1 As shown, the method includes:
[0058] S101: Obtain a preset word library to perform sequence labeling on each text in the knowledge base, and input the labeled text into a preset recognition model to calculate a root word library.
[0059] Among them, the word library is pre-set. Since business entities need to be determined in the embodiments of the present invention, which are not exactly the same as general words and usually correspond to a specific business scenario, the initial word library can be manually annotated. For example, customer service staff can manually annotate some basic words frequently used in each business scenario. For example, words such as "AAPLUS member", "AA Allianz", "AA Home", and "AA Home APP" often appear in some businesses and are all extended from AA. Therefore, AA can be annotated, that is, used as a word in the preset word library. In the embodiments of the present invention, the above-mentioned preset word library can also be implemented by an extended method of a weakly supervised algorithm. To avoid including general words in the word library, the word library can be de-duplicated based on the general word library first, and then the de-duplicated word library can be processed subsequently. The general word library can be pre-set or provided by a third party.
[0060] The knowledge base is pre-established and can be a corpus used in each business scenario. For example, if business entities need to be determined in the customer sales scenario in the embodiments of the present invention, the knowledge base can be a corpus used in the customer sales scenario, which can include various text files, etc. Each text can include one or more sentences, and each text is associated with keywords and the corresponding business scenario, that is, each sentence is also associated with keywords and the corresponding business scenario. The knowledge base is the basis for determining business entities.
[0061] Since each text and other corpus will be associated with corresponding keywords when being input into the knowledge base, and these keywords are usually words associated with the business scenario, the keyword library associated with each text and other corpus in the knowledge base can be added to the preset word library in this step to improve the comprehensiveness of the word library.
[0062] Based on the word library, sequence annotation can be performed on each corpus in the knowledge base, that is, generate the sequence annotation corresponding to the word library. For example, it can be BIO annotation. Then the sequence annotation is input into the preset recognition model, and the root word library can be calculated. The recognition model can include one or more, for example, it can include a supervised Albert model and an unsupervised combined entropy model (Unper). Based on the word library calculated by the recognition model, to avoid including general words, it can be de-duplicated based on the general word library first, and then the root word library can be obtained.
[0063] Since the knowledge base can be used for corpus retrieval, retrieval words usually need to be input during retrieval, and these retrieval words are usually also words related to the business. Therefore, to improve the comprehensiveness of the root words, the historical retrieval words of the knowledge base can be added to the root word library. At the same time, to avoid losing the words in the word library during the calculation process, the preset word library can also be added to the root word library.
[0064] As Figure 2 shown, it is a schematic flowchart of calculating the root word library in the embodiments of the present invention. As Figure 2As shown in the figure, first execute step1: After removing duplicates from the pre-annotated thesaurus CD through the common thesaurus, add the keyword thesaurus kefu associated with each corpus in the knowledge base to obtain the thesaurus Kd-1; then execute step2: Use Kd-1 to perform sequence annotation on the text K&A in the knowledge base, and then calculate the corresponding thesauruses through the Albert model and the Unper model respectively, namely Sup and Unsup. After removing duplicates from Sup and Unsup through the common thesaurus and then superimposing them, Kd-2’ can be obtained. At this time, Kd-2’ can be proofread by manual screening to obtain the thesaurus Kd-2; then execute step3: Proofread the historical search thesaurus search of the knowledge base through manual screening to obtain the Kd-afs thesaurus, superimpose Kd-1 and Kd-2 to obtain the Kd-bfs thesaurus, and superimpose the two thesauruses Kd-afs and Kd-bfs to obtain the final root thesaurus Kd-3.
[0065] S102: For each root in the root thesaurus, screen the sentences including the root from the knowledge base, perform word segmentation on each sentence, and combine the word segmentation results to determine the word segmentation combination corresponding to each sentence.
[0066] Among them, the word segmentation combination includes the said root. Any root in the root thesaurus is included in the screened sentences. After obtaining the root thesaurus, the root can be extended based on the text in the knowledge base. Therefore, for each root, the sentences including the root can be screened from the knowledge base. Word segmentation can be performed on each sentence to obtain the word segmentation result of each sentence, and then the word segmentation results can be combined to obtain the word segmentation combination including the root.
[0067] It should be noted that, in order to determine the comprehensiveness of business entities, in the embodiments of the present invention, each root in the root thesaurus can be processed in sequence to obtain the business entity corresponding to the root, and then a business entity library can be formed. Based on the business entity library, the business entities in the subsequent received text to be processed can be recognized.
[0068] For each sentence, word segmentation can be performed first. Since the embodiments of the present invention are based on root extension to obtain business entities, word segmentation combination can be performed based on the roots in the sentence, that is, the word segmentation combination needs to include the root. Specifically, generating the word segmentation combination including the root can be executed as follows: For each sentence, perform word segmentation on the sentence. Taking the root as the starting word, and taking the words in the sentence that are sequentially located after the root as the ending words in turn. Based on the starting word and the ending words, determine the word segmentation combination from the sentence. For example, if the sentence is "AA Logistics Transfer Details", after word segmentation, the words obtained can be: AA, Logistics, Transfer, Details. Taking the root as the starting word and taking the words in the sentence that are sequentially located after the root as the ending words in turn, the word segmentation combinations obtained can be: AA Logistics, AA Logistics Transfer, AA Logistics Transfer Details.
[0069] S103: Based on the business scenario to which the word segmentation combination corresponds and the preset business scenario association order, splice the word segmentation combinations including the same root word to obtain a splicing result.
[0070] When business entities in different scenarios are involved in a statement, the inference relationship between the scenarios is usually reflected. Therefore, the inference relationship between business scenarios can be considered when determining business entities. In the embodiments of the present invention, the association order between business scenarios can be set based on the inference relationship between business scenarios to indicate that the business scenario located behind in the association order is affected by the business scenario located in front. For example, the business scenarios include UAD (After-sales Work Order System), JDL (Logistics), and IM (Online Chatbot), and the preset association order can be: UAD-JDL-IM. IM is affected by UAD and JDL, JDL is affected by UAD, and UAD is not affected by other business scenarios.
[0071] In the embodiments of the present invention, for each root word, each business scenario can correspond to multiple statements including the root word, and then multiple word segmentation combinations will be obtained. In order to determine comprehensive business entities, it is necessary to splice the word segmentation combinations corresponding to each statement in each business scenario.
[0072] In the embodiments of the present invention, for each root word, one statement can be sequentially selected from the statements corresponding to each business scenario as the target statement first, and then based on the business scenario association order and the business scenarios to which the target statements belong, the word segmentation combinations corresponding to the target statements are spliced. If a certain business scenario does not correspond to a statement including the root word, it means that this business scenario is not included in the expansion of the root word, and then this business scenario in the preset business association order can be deleted. For example, the business scenarios include UAD, JDL, and IM, and the preset association order is: UAD-JDL-IM. When processing the root word AA, if there is no statement including AA in the statements corresponding to JDL, then JDL is deleted from the business scenario association order, so the business scenario association order is updated to: UAD–IM.
[0073] Specifically, the method of obtaining the splicing result in this step can be executed as follows: for the word segmentation combination corresponding to each statement, sort based on the order of the end word in the corresponding statement to generate a word segmentation array corresponding to each statement; determine the business scenario corresponding to the word segmentation array based on the business scenario to which the word segmentation array belongs; screen the target word segmentation arrays including the same root word, and splice the word segmentation combinations at the same position in the target word segmentation data corresponding to each business scenario based on the business scenario association order to obtain the splicing result.
[0074] The segmentation combination permutation for each statement can generate a corresponding segmentation array, and the permutation order of the segmentation combinations in the segmentation array is the permutation order of the end words in the statement within the segmentation combination. For example, if the segmentation combinations are: AA Logistics, AA Logistics Transfer, AA Logistics Transfer Details, then the resulting segmentation array can be (AA Logistics, AA Logistics Transfer, AA Logistics Transfer Details), that is, the order of each segmentation combination in the segmentation array is the permutation order of the end words in the corresponding statement within the segmentation combination. Then, based on the business scenario to which the statement belongs corresponding to the segmentation array, the business scenario corresponding to the segmentation data can also be determined. In this way, the target segmentation arrays with the same root can be selected from all the segmentation arrays, and based on the business scenario association order, the segmentation combinations in the same position in the target segmentation arrays corresponding to each business scenario are concatenated to obtain the concatenation result.
[0075] It should be noted that the business scenarios later in the business scenario association order are affected by the business scenarios earlier in the order. Therefore, the segmentation arrays corresponding to the business scenarios later in the business scenario association order need to be concatenated with the segmentation arrays of the business scenarios earlier in the order one by one based on the business scenario association order. Therefore, in this step, based on the business scenario association order, concatenating the segmentation combinations in the same position in the obtained segmentation arrays to obtain the concatenation result can be executed as follows: for the segmentation array corresponding to each business scenario, based on the business scenario association order, concatenate it with the segmentation combinations in the same position in the segmentation arrays corresponding to the business scenarios earlier in the business scenario association order than this business scenario to obtain the concatenation result.
[0076] For example, the business scenarios include: UAD, JDL, and IM. The preset association order is: UAD - JDL - IM. The root word is AA. A statement including AA corresponding to UAD is: AAPLUS membership classic annual card. A statement including AA corresponding to JDL is: AA logistics transfer details. A statement including AA corresponding to IM is: AA Allianz Insurance. Then, a word segmentation array corresponding to UAD can be obtained as (AAPLUS, AAPLUS membership, AAPLUS membership classic, AAPLUS membership classic annual card). A word segmentation array corresponding to JDL is (AA logistics, AA logistics transfer, AA logistics transfer details). A word segmentation array corresponding to IM is (AA Allianz, AA Allianz Insurance). From the UAD - JDL - IM association order, it can be seen that IM is affected by UAD and JDL, JDL is affected by UAD, and UAD is not affected by other business scenarios. At this time, since UAD is not affected by other business scenarios, the word segmentation combinations in its word segmentation array can be directly used as the splicing results. JDL is affected by UAD, so the word segmentation combinations with the same position in the word segmentation arrays corresponding to JDL and UAD are spliced to obtain the splicing results: AAPLUS - AA logistics, AAPLUS membership - AA logistics transfer. IM is affected by UAD and JDL, then the obtained splicing results are: AAPLUS - AA logistics - AA Allianz, AAPLUS membership - AA logistics transfer - AA Allianz Insurance. So the final obtained splicing results are: AAPLUS, AAPLUS membership, AAPLUS membership classic, AAPLUS membership classic annual card, AAPLUS - AA logistics, AAPLUS membership - AA logistics transfer, AAPLUS - AA logistics - AA Allianz, AAPLUS membership - AA logistics transfer - AA Allianz Insurance.
[0077] It should be noted that during the splicing process, if the number of word segmentation combinations included in the word segmentation array corresponding to the business scenario is insufficient, it can first be determined whether the business scenario is at the last position in the business scenario association order. If so, the splicing of the word segmentation combinations of this business scenario can be stopped. Otherwise, the last word segmentation combination in the word segmentation array can be used to replace it for subsequent splicing. For example, if the word segmentation array corresponding to JDL is (AA logistics), then AAPLUS membership - AA logistics transfer - AA Allianz Insurance in the above splicing results should be: AAPLUS membership - AA logistics - AA Allianz Insurance.
[0078] S104: Call the preset calculation model to calculate the effective scores of each splicing result, and then based on the effective scores, determine the business entities from the splicing results to obtain the business entity library.
[0079] The calculation model is pre-trained. In this step, after calculating the effective score for the splicing result, it can be determined whether the effective score of the splicing result is greater than a preset score threshold; if so, the word segmentation combination with the last splicing order in the splicing result is determined as the business entity and updated to the business entity library; if not, the splicing result is not processed. For example, for AAPLUS-AA Logistics, if its effective score is greater than the score threshold, AA Logistics can be determined as the business entity.
[0080] S105: Based on the business entity library, identify the business entities in the received text to be processed.
[0081] Among them, in order to ensure the comprehensiveness of the business entities, in the embodiments of the present invention, each root word in the root word library can be processed step by step according to S102-S104 to obtain the business entity corresponding to the root word, and then a business entity library can be formed. Based on the business entity library, the business entities in the subsequent received text to be processed can be identified. Specifically, the received text to be processed can be the consultation text determined based on the consultation information sent by the user. At this time, the business entity library can be matched with the consultation text to identify the business entities in the consultation text, and then the corresponding response file can be queried from materials such as the knowledge base based on the identified business entities to reply to the user's consultation information, thereby improving the accuracy of the user consultation reply.
[0082] It should be noted that after the business entity library is determined, the business entities can be edited manually, mainly for manually correcting and editing the business entity results; based on the business scenarios corresponding to the business entities, a tree structure can also be established to effectively summarize the inference relationships between the business entities; and in the embodiments of the present invention, the business entities can also be updated based on manual or real-time calculations; at the same time, business entity queries and new business scenarios can also be realized.
[0083] After the business entity is determined, the effective rate of the business entity can be evaluated and determined through the fitting calculation module. Since the usage methods of business entities are different in different scenarios, the fitting can be carried out in combination with corresponding business indicators. For example, when applying the business entity recognition to the knowledge base search, the search effectiveness of the recognized business entity can be fitted and calculated with the historical search results or real-time search results. If the fitting degree decreases, the correction method can be activated to correct the business entity. Correcting the business entity can include correction judgment and dynamic optimization. Since the effective score is calculated in the embodiments of the present invention, the splicing results with higher effective scores but not greater than the preset score threshold can be screened for manual re-screening, so as to directly perform the correction judgment of the business entity. Dynamic optimization is to use the results with high fitting degree and low fitting degree as positive and negative samples respectively into the calculation model for self-updating of the model according to the fitting results.
[0084] After expanding the root word library based on the preset word library and knowledge base, sentences including the root words in the root word library can be screened out and the corresponding word segmentation combinations of each sentence can be determined. This is equivalent to expanding the root words based on the text in the knowledge base to obtain the word segmentation combinations corresponding to each scenario. Then, based on the association order of each business scenario and the business scenarios to which the word segmentation combinations of each sentence belong, the word segmentation combinations including the same root words can be spliced to obtain the splicing result. This is equivalent to splicing the word segmentation combinations in combination with the associations between scenarios, so that the inference relationship between business scenarios is reflected in the splicing result. Then, based on the effective score of the splicing result, business entities can be determined. In this way, first expand the root word library from the preset word library, then expand the word segmentation combinations of each business scenario based on the root word library, and then determine the business entities in combination with the inference relationship between business scenarios to obtain the business entity library for identifying the business entities in the received text to be processed, which not only improves the efficiency of determining business entities, but also improves the accuracy and comprehensiveness of determining business entities, and further increases the accuracy of business entity recognition in the text.
[0085] It should be noted that steps S102, S103, and S104 in the embodiments of the present invention can be executed simultaneously. For example, for a root word in the root word library, sentences including the root word can be screened out and determined as target sentences, and then the processing of steps S102, S103, and S104 can be performed on the target sentences. In the embodiments of the present invention, taking the preset association order as: UAD-JDL-IM, the root word as AA, a target sentence of UAD as: AAPLUS membership classic annual card, a target sentence of JDL as: AA logistics transfer details, and a target sentence of IM as: AA Allianz Insurance as an example, the method for determining business entities will be described as follows. Figure 3 As shown, the method includes:
[0086] S301: Segment each target sentence, with the target root word as the starting word and the adjacent word after the root word as the ending word.
[0087] As Figure 4 shown is the logical schematic diagram for determining business scenarios, and AA is the root word. Therefore, the ending words of each target sentence in this step are PLUS, logistics, and Allianz in sequence.
[0088] S302: Based on the starting word and the ending word, determine the corresponding word segmentation combinations of each target sentence.
[0089] Based on the processing results of step S301, the word segmentation combinations can be obtained as: AAPLUS, AA logistics, and AA Allianz respectively.
[0090] S303: Based on the business scenario association order, splice the word segmentation combinations corresponding to each target sentence.
[0091] As Figure 4As shown, Concat represents a splicing algorithm that can splice the pre- and post-word segmentations together to obtain the splicing results: AAPLUS, AAPLUS-AA Logistics, and AAPLUS-AA Logistics-AA Allianz.
[0092] S304: Calculate the feature vectors of the splicing results, call the preset calculation model to extract the features of the splicing results based on the feature vectors, and then calculate the floating-point values corresponding to the features, and determine the floating-point values as the effective scores of the splicing results.
[0093] As Figure 4 shown, Inverse Decoding represents a vector calculation algorithm, and Conv1d represents a vector extraction and floating-point calculation algorithm. The vector extraction algorithm can be specifically implemented through a convolutional layer.
[0094] From Figure 4 as shown, the calculation model can calculate the effective scores of AAPLUS, AAPLUS-AA Logistics, and AAPLUS-AA Logistics-AA Allianz in sequence.
[0095] S305: If the effective score of the splicing result is greater than the preset score threshold, then determine the word segmentation combination with the last splicing order in the splicing result as the business entity and update it to the business entity library; if the effective score of the splicing result is not greater than the preset score threshold, then do not process the splicing result.
[0096] For each splicing result, the business entity can be determined by comparing the effective score with the score threshold. The score threshold can be determined during the model training process. As Figure 4 shown, the left side of the figure represents the score thresholds corresponding to the splicing results of each layer.
[0097] S306: Determine whether there are no word segmentations after the end word in the target sentence of each business scenario. If so, end the process; if not, execute step S307.
[0098] S307: Determine whether the business scenario corresponding to the target sentence without word segmentations after the end word is the last one in the business scenario association order. If so, delete the last business scenario in the business scenario association order to update the business scenario association order and execute this step; if not, execute step S308.
[0099] When the business scenario corresponding to the target sentence without word segmentations after the end word is the last one in the business scenario association order, since it will not affect other business scenarios, the last business scenario in the business scenario association order can be deleted to update the business scenario association order.
[0100] S308: For the target sentence that includes word segmentation after the end word, update the word segmentation after the end word to the end word, determine the new word segmentation combinations corresponding to each target sentence based on the start word and the end word, and then splice the new word segmentation combinations corresponding to each target sentence in the order associated with the updated business scenario to obtain a splicing result, and execute step S304.
[0101] For the target sentence that includes word segmentation after the end word, update the word segmentation after the end word to the end word to update the word segmentation combination, while for the target sentence that does not include word segmentation after the end word, do not update the word segmentation combination.
[0102] It should be noted that a target sentence usually includes at most one business entity. Therefore, after determining a business entity in the target sentence, in order to simplify the calculation process, the word segmentation combination after this business entity can no longer be processed. Therefore, for the target sentence with a determined business entity, the end word can also no longer be updated to stop updating the word segmentation combination.
[0103] After expanding the root word library based on the preset word library and knowledge base, sentences including the root words in the root word library can be screened out and the word segmentation combinations corresponding to each sentence can be determined. This is equivalent to expanding the root words based on the text in the knowledge base to obtain the word segmentation combinations corresponding to each scenario. Then, based on the association order of each business scenario and the business scenario to which each word segmentation combination corresponds, the word segmentation combinations including the same root words can be spliced to obtain a splicing result. This is equivalent to splicing the word segmentation combinations in combination with the associations between scenarios, so that the inference relationship between business scenarios is reflected in the splicing result. Then, based on the effective score of the splicing result, the business entity is determined. In this way, first expand the root word library from the preset word library, then expand the word segmentation combinations of each business scenario based on the root word library, and then determine the business entity in combination with the inference relationship between business scenarios to obtain the business entity library, which not only improves the efficiency of determining the business entity, but also improves the accuracy and comprehensiveness of determining the business entity, and further increases the accuracy of business entity recognition in the text.
[0104] To solve the problems existing in the prior art, an embodiment of the present invention provides a text processing device 500, as Figure 5 shown. The device 500 includes:
[0105] A calculation unit 501, configured to obtain a preset word library to perform sequence annotation on each text in the knowledge base, and input the annotated text into a preset recognition model to calculate a root word library;
[0106] A determination unit 502, configured to, for each root word in the root word library, screen out sentences including the root word from the knowledge base, perform word segmentation on each sentence, and combine the word segmentation results to determine the word segmentation combination corresponding to each sentence, where the word segmentation combination includes the root word;
[0107] A screening unit 503, configured to splice word segment combinations including the same root word based on the business scenario to which the word segment combination corresponding statement belongs and a preset business scenario association order, and obtain a splicing result;
[0108] The calculation unit 501 is further configured to call a preset calculation model, calculate the effective scores of each splicing result, and then determine business entities from the splicing results based on the effective scores, so as to update them to the business entity library;
[0109] A processing unit 504, configured to identify business entities in the received text to be processed based on the business entity library.
[0110] It should be understood that the implementation manner of the embodiments of the present invention is the same as that of the embodiments Figure 1 shown, and will not be described in detail here.
[0111] In an implementation manner of the embodiments of the present invention, the calculation unit 501 is specifically configured to:
[0112] For each statement, segment the statement, use the root word as the starting word, and sequentially use the word segments located after the root word in the statement as the ending words. Based on the starting word and the ending words, determine word segment combinations from the statement.
[0113] In an implementation manner of the embodiments of the present invention, the splicing unit 503 is specifically configured to:
[0114] For the word segment combination corresponding to each statement, sort them based on the order of the ending words in the corresponding statement to generate a word segment array corresponding to each statement;
[0115] Based on the business scenario to which the word segment array corresponds, determine the business scenario corresponding to the word segment array;
[0116] Screen target word segment arrays corresponding to the same root word, and splice the word segment combinations in the same position in the target word segment arrays corresponding to each business scenario based on the business scenario association order, so as to obtain a splicing result.
[0117] In an implementation manner of the embodiments of the present invention, the calculation unit 501 is specifically configured to:
[0118] For each splicing result, calculate the feature vector of the splicing result, call a preset calculation model, extract the features of the splicing result based on the feature vector, and then calculate the floating-point value corresponding to the feature, and determine the floating-point value as the effective score of the splicing result.
[0119] In an implementation manner of the embodiments of the present invention, the processing unit 504 is specifically configured to:
[0120] Determine whether the effective score of the splicing result is greater than a preset score threshold;
[0121] If so, determine the word segmentation combination with the last splicing order in the splicing result as the business entity and update it to the business entity library; if not, do not process the splicing result.
[0122] In an implementation manner of the embodiment of the present invention, the device 500 further includes:
[0123] An adding unit, configured to obtain the historical search terms of the knowledge base and add the historical keywords and the preset thesaurus to the root word library.
[0124] In an implementation manner of the embodiment of the present invention, the text to be processed includes the consultation text sent by the user; the processing unit 504 is specifically configured to:
[0125] Receive the consultation information sent by the user and determine the consultation text corresponding to the consultation information;
[0126] Match the business entity library with the consultation text, identify the business entities included in the consultation text, and query the corresponding response file based on the business entities included in the consultation text, so as to reply to the consultation information.
[0127] It should be understood that the implementation manner of the embodiment of the present invention is the same as that of the embodiment shown in Figure 1 2 or 3, and will not be described in detail here.
[0128] In the embodiment of the present invention, after expanding the root word library based on the preset thesaurus and the knowledge base, sentences including the root words in the root word library can be screened out and the corresponding word segmentation combinations of each sentence can be determined. It is equivalent to expanding the root words based on the text in the knowledge base to obtain the word segmentation combinations corresponding to each scenario. Then, based on the association order between business scenarios and the business scenarios to which the sentences corresponding to each word segmentation combination belong, the word segmentation combinations including the same root words can be spliced to obtain a splicing result. It is equivalent to splicing the word segmentation combinations in combination with the association between scenarios, so that the reasoning relationship between business scenarios is reflected in the splicing result. Then, based on the effective score of the splicing result, the business entity is determined. In this way, the root word library is first expanded from the preset thesaurus, and then the word segmentation combinations of each business scenario are expanded based on the root word library. Furthermore, the business entity is determined in combination with the reasoning relationship between business scenarios to obtain the business entity library for identifying the business entities in the received text to be processed, which not only improves the efficiency of determining the business entity, but also improves the accuracy and comprehensiveness of determining the business entity, and further increases the accuracy of identifying the business entities in the text.
[0129] According to the embodiment of the present invention, the embodiment of the present invention also provides an electronic device and a readable storage medium.
[0130] The electronic device of an embodiment of the present invention comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the one processor, and the instructions are executed by the at least one processor so that the at least one processor executes the text processing method provided by the embodiment of the present invention.
[0131] Figure 6 An exemplary system architecture 600 is shown to which the text processing method or text processing apparatus according to the embodiment of the present invention can be applied.
[0132] like Figure 6 As shown, system architecture 600 may include terminal devices 601, 602, 603, network 604 and server 605. Network 604 is used to provide a medium for communication links between terminal devices 601, 602, 603 and server 605. Network 604 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0133] The user can use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various client applications can be installed on the terminal devices 601, 602, 603.
[0134] The terminal devices 601 , 602 , and 603 may be, but are not limited to, smart phones, tablet computers, laptop computers, and desktop computers, etc.
[0135] The server 605 may be a server that provides various services. The server may analyze and process the received data such as the product information query request, and feed back the processing result (such as product information - just an example) to the terminal device.
[0136] It should be noted that the text processing method provided in the embodiment of the present invention is generally executed by the server 605 , and accordingly, the text processing device is generally set in the server 605 .
[0137] It should be understood that Figure 6 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0138] Reference below Figure 7 , which shows a schematic diagram of the structure of a computer system 700 suitable for implementing an embodiment of the present invention. Figure 7 The computer system shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0139] like Figure 7As shown, computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0140] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom can be installed into the storage section 708 as needed.
[0141] Specifically, according to an embodiment disclosed by the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment disclosed by the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above functions defined in the system of the present invention are executed.
[0142] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a unit, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in a block can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0144] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a calculation unit, a screening unit, and a processing unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the calculation unit can also be described as "a unit for the function of the calculation unit".
[0145] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device is caused to execute the text processing method provided by the present invention.
[0146] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A text processing method, characterized in that, it includes: Obtain a preset thesaurus to perform sequence annotation on each text in the knowledge base, input the annotated text into a preset recognition model, and calculate a root word library; For each root word in the root word library, screen the sentences including the root word from the knowledge base, perform word segmentation on each sentence, and combine the word segmentation results to determine the word segmentation combination corresponding to each sentence, where the word segmentation combination includes the root word; Based on the business scenario to which the word segmentation combination corresponds and the preset business scenario association order, splice the word segmentation combinations including the same root word to obtain a splicing result; Call a preset calculation model to calculate the effective score of each splicing result, and then based on the effective score, determine business entities from the splicing results to obtain a business entity library; Based on the business entity library, identify business entities in the received text to be processed; The calling of the preset calculation model to calculate the effective score of each splicing result includes: For each splicing result, calculate the feature vector of the splicing result, call a preset calculation model to extract the features of the splicing result based on the feature vector, and then calculate the floating-point value corresponding to the feature, and determine the floating-point value as the effective score of the splicing result.
2. The method according to claim 1, characterized in that, Performing word segmentation on each sentence and combining the word segmentation results to determine the word segmentation combination corresponding to each sentence includes: For each sentence, perform word segmentation on the sentence, use the root word as the starting word, and sequentially use the word segments in the sentence that are located after the root word as the ending words, and based on the starting word and the ending words, determine the word segmentation combination from the sentence.
3. The method according to claim 2, characterized in that, The splicing of the word segmentation combinations including the same root word based on the business scenario to which the word segmentation combination corresponds and the preset business scenario association order to obtain a splicing result includes: For the word segmentation combination corresponding to each sentence, sort based on the order of the ending words in the corresponding sentence to generate a word segmentation array corresponding to each sentence; Based on the business scenario to which the word segmentation array corresponds, determine the business scenario corresponding to the word segmentation array; Screen the target word segmentation arrays corresponding to the same root word, and based on the business scenario association order, splice the word segmentation combinations in the same position in the target word segmentation arrays corresponding to each business scenario to obtain a splicing result.
4. The method according to claim 1, characterized in that, Determining business entities from the splicing results based on the effective score and updating them to the business entity library includes: Judging whether the effective score of the splicing result is greater than a preset score threshold; If so, determine the word segmentation combination with the last splicing order in the splicing result as the business entity and update it to the business entity library; if not, do not process the splicing result.
5. The method according to claim 1, characterized in that, Before screening the sentences including the root word from the knowledge base, it further includes: Obtain the historical search terms of the knowledge base, and add the historical search terms and the preset thesaurus to the root word library.
6. The method according to claim 1, wherein, the text to be processed includes the consultation text sent by the user; identifying the business entities in the received text to be processed based on the business entity library includes: receiving the consultation information sent by the user, and determining the consultation text corresponding to the consultation information; matching the business entity library with the consultation text, identifying the business entities included in the consultation text, querying the corresponding response file based on the business entities included in the consultation text, and then replying to the consultation information.
7. A text processing device, wherein, it includes: a calculation unit, configured to obtain a preset word library, perform sequence annotation on each text in the knowledge base, input the annotated text into a preset recognition model, and calculate a root word library; a determination unit, configured to, for each root word in the root word library, screen out the sentences including the root word from the knowledge base, perform word segmentation on each sentence, and combine the word segmentation results to determine the word segmentation combination corresponding to each sentence, wherein the word segmentation combination includes the root word; a splicing unit, configured to splice the word segmentation combinations including the same root word based on the business scenario to which the word segmentation combination belongs and the preset business scenario association order, and obtain a splicing result; the calculation unit is further configured to call a preset calculation model, calculate the effective score of each splicing result, and then determine the business entity from the splicing results based on the effective score, so as to update the business entity library; a processing unit, configured to identify the business entities in the received text to be processed based on the business entity library; the calculation unit is specifically configured to: for each splicing result, calculate the feature vector of the splicing result, call a preset calculation model, extract the feature of the splicing result based on the feature vector, and then calculate the floating-point value corresponding to the feature, and determine the floating-point value as the effective score of the splicing result.
8. An electronic device, wherein, it includes: one or more processors; a storage device, configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium, on which a computer program is stored, wherein, when the program is executed by a processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Classification method and device for text information
CN109002443A
Electric power data analysis method and system based on data grading model
CN112257425A