Method, device and electronic equipment for processing corpus data

By dividing the corpus into groups based on the hit rate and establishing relationships within the bank's customer service system, and dynamically adjusting the corpus classification, the problem of poor user experience caused by manual annotation was solved, and the user experience and response accuracy of the corpus were improved.

CN115391539BActive Publication Date: 2026-01-09BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211052774.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2026-01-09
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

The response data in the bank's customer service system requires manual annotation, resulting in a poor user experience.

Method used

By dividing the corpus into groups in the system database and classifying the corpus into first and second corpus sets based on the corpus hit rate, the corpus in the first corpus set is used to add or delete data from the second corpus set, establishing associations, optimizing the similarity and keyword repetition between corpus sets, and dynamically adjusting the corpus classification.

Benefits of technology

It improved the user experience and response accuracy of the corpus in the bank's customer service system, reduced the reliance on manual annotation, and enhanced the system's automation and corpus search efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391539B_ABST
    Figure CN115391539B_ABST
Patent Text Reader

Abstract

The application discloses a corpus data processing method and device and electronic equipment, applied to the field of big data, the method comprises: obtaining the system database corresponding to the application system, the application system can obtain the corpus corresponding to the query request in the system database for the query request; the system database contains at least a plurality of corpus groups; each corpus group contains at least one target corpus; for each corpus group, the target corpus is divided into a first corpus set and a second corpus set according to the corpus hit rate of the target corpus; the first corpus set contains at least one first corpus, and the second corpus set contains at least one second corpus, and the corpus hit rate of the first corpus is greater than that of the second corpus; using the corpus in the second corpus set corresponding to the first corpus group, the corpus in the second corpus set corresponding to the second corpus group is added or deleted, and the first corpus group and the second corpus group have an association relationship.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a corpus data processing method and device and electronic equipment. BACKGROUND

[0002] In a bank customer service system, after receiving a question raised by a customer, a suitable response corpus is screened out through corpus labeling, thereby providing consulting services for the customer.

[0003] However, these response corpora need to be manually labeled, resulting in poor user experience of the bank customer service system. SUMMARY

[0004] Therefore, the present application provides a corpus data processing method and device and electronic equipment to solve the technical problem of poor user experience of the application system corresponding to the system database.

[0005] A corpus data processing method, the method comprising:

[0006] obtaining a system database corresponding to an application system, the application system being capable of obtaining a corpus corresponding to a query request in the system database for the query request; the system database comprising at least a plurality of corpus groups, each corpus group corresponding to a business type; and each corpus group comprising at least one target corpus;

[0007] for each corpus group, dividing the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set comprising at least one first corpus, and the second corpus set comprising at least one second corpus, the corpus hit rate of the first corpus being greater than the corpus hit rate of the second corpus;

[0008] using the corpus in the second corpus set corresponding to a first corpus group to add or delete the corpus in the second corpus set corresponding to a second corpus group, the first corpus group and the second corpus group having an association relationship.

[0009] The above method, preferably, the first corpus group and the second corpus group have an association relationship, comprising:

[0010] the corpus group similarity between the first corpus group and the second corpus group is greater than the corpus group similarity between the first corpus group and other corpus groups in the plurality of corpus groups;

[0011] and the corpus group similarity between the first corpus group and the second corpus group is greater than the corpus group similarity between the second corpus group and other corpus groups in the plurality of corpus groups.

[0012] The method, preferably, the corpus similarity between the first corpus and the second corpus is:

[0013] The first set similarity and the second set similarity are weighted and averaged to obtain an overall similarity using respective weights;

[0014] The first set similarity is the similarity between the first corpus set in the first corpus and the first corpus set in the second corpus; and the second set similarity is the similarity between the second corpus set in the first corpus and the second corpus set in the second corpus.

[0015] The method, preferably, the corpus in the second corpus set corresponding to the first corpus is used to add the corpus in the second corpus set corresponding to the second corpus, comprising:

[0016] In the second corpus contained in the second corpus set corresponding to the first corpus, a third corpus with a corpus hit rate greater than or equal to a first threshold is obtained; and the third corpus is added to the second corpus set corresponding to the second corpus.

[0017] The method, preferably, the corpus in the second corpus set corresponding to the first corpus is used to delete the corpus in the second corpus set corresponding to the second corpus, comprising:

[0018] In the second corpus set corresponding to the second corpus, a fourth corpus is deleted, the fourth corpus being a corpus added to the second corpus from the second corpus set corresponding to the first corpus under the condition that the corpus hit rate is greater than or equal to a first threshold, and the fourth corpus having a corpus hit rate in the second corpus set corresponding to the first corpus decreasing from greater than or equal to the first threshold to less than the first threshold.

[0019] The method, preferably, the method further comprises:

[0020] In the second corpus set corresponding to the first corpus, a first associated corpus is obtained;

[0021] In the second corpus set corresponding to the second corpus, a second associated corpus is obtained; the second associated corpus and the first associated corpus being derived from a target source document, the target source document having a number of target corpora generated in the plurality of source documents corresponding to the system database satisfying a target screening condition;

[0022] The first associated corpus is moved to the second corpus set corresponding to the second corpus;

[0023] The second associated corpus is moved to the second corpus set corresponding to the first corpus.

[0024] The method, preferably, further comprises:

[0025] obtaining at least one new corpus;

[0026] performing word segmentation on the new corpus to obtain corpus keywords of each of the new corpus;

[0027] obtaining keyword repetition degrees of the corpus keywords in a first corpus set and a second corpus set corresponding to a corpus group to which the new corpus corresponds, respectively;

[0028] adding the new corpus to the first corpus set or the second corpus set corresponding to the corpus group to which the new corpus corresponds, according to the keyword repetition degrees.

[0029] The method, preferably, further comprises:

[0030] obtaining a target query request, the target query request containing at least a query keyword;

[0031] performing corpus query on the first corpus set and the second corpus set corresponding to a target corpus group using the query keyword, to obtain a first query result and a second query result; the target corpus group is a corpus group corresponding to a business type of the query keyword;

[0032] sorting the corpus in the first query result and the corpus in the second query result according to corpus similarity, to obtain a sorting result;

[0033] outputting the corpus in the first query result and the corpus in the second query result according to the sorting result.

[0034] A corpus data processing apparatus, comprising:

[0035] a data obtaining unit configured to obtain a system database corresponding to an application system, the application system being capable of obtaining corpus corresponding to a query request in the system database according to the query request; the system database containing at least a plurality of corpus groups, each of the corpus groups corresponding to a business type; each of the corpus groups containing at least one target corpus;

[0036] a corpus dividing unit configured to divide, for each of the corpus groups, the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set containing at least one first corpus, the second corpus set containing at least one second corpus, the corpus hit rate of the first corpus being greater than that of the second corpus;

[0037] The corpus processing unit is configured to add or delete the corpus in the second corpus set corresponding to the second corpus group by using the corpus in the second corpus set corresponding to the first corpus group.

[0038] An electronic device comprises:

[0039] A memory is configured to store a computer program and data generated by running of the computer program.

[0040] A processor is configured to execute the computer program to implement the following steps: obtaining a system database corresponding to an application system, the application system being capable of obtaining corpus corresponding to a query request in the system database according to the query request; the system database comprising at least a plurality of corpus groups, each corpus group corresponding to a business type; each corpus group comprising at least one target corpus; for each corpus group, dividing the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set comprising at least one first corpus, the second corpus set comprising at least one second corpus, the corpus hit rate of the first corpus being greater than the corpus hit rate of the second corpus; adding or deleting the corpus in the second corpus set corresponding to a second corpus group by using the corpus in the second corpus set corresponding to a first corpus group, the first corpus group and the second corpus group having an association relationship.

[0041] As can be seen from the above solution, in the corpus data processing method, apparatus and electronic device provided by the present application, the target corpus in the system database is first divided into two corpus groups according to the business type and the corpus hit rate, and then the corpus in the corpus group corresponding to the associated business type is added or deleted by using the corpus in one of the corpus groups, so as to adjust the corpus in the corpus group corresponding to each business type in the system database, and further improve the use experience of the application system corresponding to the system database. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0043] Figure 1 A flowchart of a corpus data processing method provided by Embodiment One of the present application;

[0044] Figure 2 、 Figure 3 and Figure 4Part of flow charts in a corpus data processing method provided by Embodiment One of the present application respectively;

[0045] Figure 5 Structure schematic diagram of a corpus data processing apparatus provided by Embodiment Two of the present application;

[0046] Figure 6 and Figure 7 Another structure schematic diagram of a corpus data processing apparatus provided by Embodiment Two of the present application respectively;

[0047] Figure 8 Structure schematic diagram of an electronic device provided by Embodiment Three of the present application. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0049] Reference Figure 1 As shown in the figure, the implementation flow chart of a corpus data processing method provided by Embodiment One of the present application. The method can be applied to electronic devices capable of data processing, such as computers or servers, etc. The technical solution in the present embodiment is mainly used to improve the use experience of the application system corresponding to the system database.

[0050] Specifically, the method in the present embodiment can include the following steps:

[0051] Step 101: Obtain the system database corresponding to the application system.

[0052] The application system can obtain the corpus corresponding to the query request in the system database for the query request. For example, the keyword in the query request is used to query the corpus satisfying the query condition in the system data for output.

[0053] Specifically, the system database contains at least a plurality of corpus groups, each of which corresponds to a business type, such as mobile banking type, loan type, personal online banking type, and account management type. Each corpus group contains at least one target corpus.

[0054] Step 102: For each corpus group, divide the target corpus into the first corpus set and the second corpus set according to the corpus hit rate of the target corpus.

[0055] In the embodiment, the corpus hit rate of each target corpus can be counted in the history use of the system database. The corpus hit rate refers to the number of query requests in which the keywords meet the query condition of the target corpus, i.e., the number of times the target corpus is queried to meet the query condition. Based on this, the first corpus set and the second corpus set corresponding to each corpus group are created according to the corpus hit rate of each target corpus. Thus, the first corpus set contains at least one first corpus, and the second corpus set contains at least one second corpus. The corpus hit rate of the first corpus is greater than that of the second corpus.

[0056] Specifically, in the embodiment, for each corpus group, the target corpus with a corpus hit rate greater than or equal to the hit threshold is divided into the first corpus set, and the target corpus with a corpus hit rate less than the hit threshold is divided into the second corpus set.

[0057] Step 103: Adding or deleting the corpus in the second corpus set corresponding to the second corpus group using the corpus in the second corpus set corresponding to the first corpus group.

[0058] The first corpus group and the second corpus group have an association relationship. That is, in the embodiment, the addition or deletion of the corpus is performed between the second corpus sets corresponding to the corpus groups having the association relationship, so as to enrich the corpus in the corpus set.

[0059] As can be seen from the above solution, the method for processing corpus data provided by the first embodiment of the present application first divides the target corpus in the system database into two corpus groups according to the corpus hit rate by business type, and then adds or deletes the corpus in the corpus group corresponding to the associated business type using the corpus in one of the corpus groups. Thus, the corpus in the corpus group corresponding to each business type in the system database is adjusted, and the use experience of the application system corresponding to the system database is improved.

[0060] In an implementation manner, the first corpus group and the second corpus group have an association relationship, which can be specifically as follows:

[0061] The corpus group similarity between the first corpus group and the second corpus group is greater than the corpus group similarity between the first corpus group and other corpus groups in the plurality of corpus groups;

[0062] The corpus group similarity between the first corpus group and the second corpus group is greater than the corpus group similarity between the second corpus group and other corpus groups in the plurality of corpus groups.

[0063] Specifically, in the embodiment, the corpus similarity between any two corpora in all corpora is counted, and then the corpora are sorted according to the corpus similarity, and the corpora whose corpus similarity satisfies a similarity condition, such as being greater than or equal to a similarity threshold, are associated.

[0064] The corpus similarity between the first corpus and the second corpus can be an overall similarity obtained by weighting and averaging the first set similarity and the second set similarity using respective corresponding weights.

[0065] The first set similarity is the similarity between the first corpus set in the first corpus and the first corpus set in the second corpus, and the second set similarity is the similarity between the second corpus set in the first corpus and the second corpus set in the second corpus.

[0066] It should be noted that the weight corresponding to the first corpus set and the weight corresponding to the second corpus set can be set according to requirements.

[0067] In an implementation manner, in step 103, when the corpus in the second corpus set corresponding to the first corpus is used to add the corpus in the second corpus set corresponding to the second corpus, the following manner can be used:

[0068] In the second corpus included in the second corpus set corresponding to the first corpus, third corpora with a corpus hit rate greater than or equal to a first threshold are obtained, and the third corpora are added to the second corpus set corresponding to the second corpus.

[0069] That is, in the embodiment, the corpus hit rate of the second corpus in the second corpus set corresponding to each corpus is counted, and then the third corpus with a corpus hit rate greater than or equal to the first threshold is selected, and then the third corpus is added to the second corpus set corresponding to the corpus with the association relationship, thereby enriching the corpus.

[0070] It should be noted that in the embodiment, the third corpus added to the second corpus set corresponding to the second corpus can be set with a corpus mark corresponding to the first corpus to indicate that the third corpus originates from the first corpus, so as to be deleted from the second corpus set corresponding to the second corpus when the corpus hit rate of the third corpus in the second corpus set corresponding to the first corpus decreases to less than the first threshold.

[0071] In an implementation manner, in step 103, when the corpus in the second corpus set corresponding to the first corpus is used to delete the corpus in the second corpus set corresponding to the second corpus, the following manner can be used:

[0072] In the second corpus set corresponding to the second corpus group, a fourth corpus is deleted, the fourth corpus is a corpus added to the second corpus group from the second corpus set corresponding to the first corpus group in a case that the corpus hit rate is greater than or equal to the first threshold, and the corpus hit rate of the fourth corpus in the second corpus set corresponding to the first corpus group decreases from greater than or equal to the first threshold to less than the first threshold.

[0073] That is, in the embodiment, after the fourth corpus is added to the second corpus group in a case that the corpus hit rate of the fourth corpus in the second corpus set corresponding to the first corpus group is greater than or equal to the first threshold, the corpus hit rate of the fourth corpus in the second corpus set corresponding to the first corpus group is continuously counted, and if the corpus hit rate of the fourth corpus decreases to less than the first threshold, the fourth corpus in the second corpus set corresponding to the second corpus group can be deleted.

[0074] In an implementation manner, the method in the embodiment can further include the following steps, as shown in Figure 2

[0075] Step 201: In the second corpus set corresponding to the first corpus group, a first associated corpus is obtained, and in the second corpus set corresponding to the second corpus group, a second associated corpus is obtained.

[0076] The second associated corpus and the first associated corpus are derived from a target source document, and the number of target corpora generated in the target source document in the plurality of source documents corresponding to the system database satisfies a target screening condition.

[0077] The target screening condition can be that the number of target corpora generated in the target source document in the plurality of source documents corresponding to the system database is the largest, or the target screening condition can be that the number of target corpora generated in the target source document in the plurality of source documents corresponding to the system database is sorted in the top Q from large to small, and Q is a positive integer greater than or equal to 2.

[0078] Step 202: Moving the first associated corpus to the second corpus set corresponding to the second corpus group, and moving the second associated corpus to the second corpus set corresponding to the first corpus group.

[0079] That is, in the embodiment, the second corpora in each second corpus set and the second corpora in the second corpus set having an associated relationship are compared one by one, and then the second corpora from the same source document are determined, the source documents are counted, and the target source document generating the most second corpora in the second corpus set is counted, thereby exchanging the corpora corresponding to the target source document in the second corpus set having an associated relationship.

[0080] In an implementation manner, the method in the embodiment can further include the following steps, as shown in​Figure 3 As shown in

[0081] Step 301: Obtain at least one new corpus.

[0082] The new corpus can be a newly obtained corpus, or a corpus that is out of the preset collection space of the first corpus set and the second corpus set.

[0083] Step 302: Tokenize the new corpus to obtain corpus keywords of each new corpus.

[0084] Specifically, in the embodiment, the new corpus can be tokenized by using a tokenization algorithm to obtain the corpus keywords.

[0085] Step 303: Obtain the keyword repetition degree of the corpus keywords in the first corpus set and the second corpus set corresponding to the corpus group corresponding to the new corpus.

[0086] Specifically, in the embodiment, the keyword repetition degree of the corpus keywords in the first corpus set corresponding to the business type to which the corpus keywords belong is counted, for example, the keyword similarity between the corpus keywords and the first corpus in the first corpus set corresponding to the business type to which the corpus keywords belong is compared, and the number of words with a keyword similarity greater than or equal to a corresponding threshold is taken as the keyword repetition degree.

[0087] In addition, in the embodiment, the keyword repetition degree of the corpus keywords in the second corpus set corresponding to the business type to which the corpus keywords belong is counted, for example, the keyword similarity between the corpus keywords and the second corpus in the second corpus set corresponding to the business type to which the corpus keywords belong is compared, and the number of words with a keyword similarity greater than or equal to a corresponding threshold is taken as the keyword repetition degree. The keyword repetition degree can be understood as the frequency of occurrence of the keyword.

[0088] Step 304: According to the keyword repetition degree, add the new corpus to the first corpus set or the second corpus set corresponding to the corpus group corresponding to the new corpus.

[0089] For example, the new corpus with a keyword repetition degree greater than or equal to a repetition degree threshold is added to the first corpus set corresponding to the business type to which the new corpus belongs, and the new corpus with a keyword repetition degree less than the repetition degree threshold is added to the second corpus set corresponding to the business type to which the new corpus belongs.

[0090] In an implementation manner, the method in the embodiment can further include the following steps, as shown in Figure 4

[0091] Step 401: Obtain a target query request, the target query request containing at least a query keyword.

[0092] ​Step 402: respectively performing corpus query in the first corpus set and the second corpus set corresponding to the target corpus group using the query keyword to obtain the first query result and the second query result.

[0093] The target corpus group is a corpus group corresponding to the business type and the query keyword. For example, the first query result is obtained by using the query keyword to query the corpus with a similarity greater than or equal to a corresponding threshold in the first corpus set corresponding to the target corpus group corresponding to the corresponding business type, and the second query result is obtained by using the query keyword to query the corpus with a similarity greater than or equal to a corresponding threshold in the second corpus set corresponding to the target corpus group corresponding to the corresponding business type.

[0094] Step 403: sorting the corpus in the first query result and the corpus in the second query result according to the corpus similarity to obtain a sorting result.

[0095] Specifically, in the embodiment, the corpus in the first query result can be sorted before the corpus in the second query result, the corpus in the first query result is sorted according to the keyword similarity between the query keyword from large to small, the corpus in the second query result is sorted according to the keyword similarity between the query keyword from large to small, then taking the corpus in the first query result as a reference, the sorting position of the corpus in the second query result with a corpus similarity greater than or equal to a corresponding threshold with the corpus in the first query result is adjusted to the adjacent position of the corresponding corpus in the first query result, that is, the sorting position of the corpus in the second query result with a higher similarity with the corpus in the first query result is adjusted forward.

[0096] Step 404: outputting the corpus in the first query result and the corpus in the second query result according to the sorting result.

[0097] For example, the corpus in the first query result and the corpus in the second query result are outputted according to the order from front to back in the sorting result for use.

[0098] Reference Figure 5 A structure diagram of a corpus data processing device provided by the second embodiment of the application, which can be configured in an electronic device capable of data processing, such as a computer or a server. The technical solution in the embodiment is mainly used to improve the use experience of the application system corresponding to the system database.

[0099] Specifically, the device in the embodiment can include the following units:

[0100] The data obtaining unit 501 is configured to obtain a system database corresponding to an application system, the application system being capable of obtaining corpus corresponding to a query request in the system database according to the query request; the system database comprising at least a plurality of corpus groups, each of the corpus groups corresponding to a business type; and each of the corpus groups comprising at least one target corpus.

[0101] The corpus dividing unit 502 is configured to divide, for each of the corpus groups, the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set comprising at least one first corpus, the second corpus set comprising at least one second corpus, and the corpus hit rate of the first corpus being greater than the corpus hit rate of the second corpus.

[0102] The corpus processing unit 503 is configured to add or delete corpus in a second corpus set corresponding to a second corpus group by using corpus in a second corpus set corresponding to a first corpus group, the first corpus group and the second corpus group having an association relationship.

[0103] According to the above scheme, the embodiment two of the present application provides a corpus data processing device, which first divides target corpus in a system database into two corpus groups according to corpus hit rate by business type, and then adds or deletes corpus in a corpus group corresponding to an associated business type by using corpus in another corpus group, thereby adjusting corpus in the corpus group corresponding to each business type in the system database, and further improving the use experience of an application system corresponding to the system database.

[0104] In an implementation manner, the association relationship between the first corpus group and the second corpus group comprises that a corpus group similarity between the first corpus group and the second corpus group is greater than a corpus group similarity between the first corpus group and other corpus groups in the plurality of corpus groups, and the corpus group similarity between the first corpus group and the second corpus group is greater than a corpus group similarity between the second corpus group and other corpus groups in the plurality of corpus groups.

[0105] In an implementation manner, the corpus group similarity between the first corpus group and the second corpus group is an overall similarity obtained by weighting and averaging a first set similarity and a second set similarity by using respective corresponding weights; the first set similarity is a similarity between a first corpus set in the first corpus group and a first corpus set in the second corpus group; and the second set similarity is a similarity between a second corpus set in the first corpus group and a second corpus set in the second corpus group.

[0106] In an implementation manner, the corpus processing unit 503 is specifically configured to: in the second corpus set corresponding to the first corpus group, obtain third corpus with a corpus hit rate greater than or equal to a first threshold; and add the third corpus to the second corpus set corresponding to the second corpus group, when adding corpus in the second corpus set corresponding to the second corpus group using corpus in the second corpus set corresponding to the first corpus group.

[0107] In an implementation manner, the corpus processing unit 503 is specifically configured to: in the second corpus set corresponding to the second corpus group, delete fourth corpus, which is corpus added to the second corpus group from the second corpus set corresponding to the first corpus group with a corpus hit rate greater than or equal to a first threshold, and the corpus hit rate of the fourth corpus in the second corpus set corresponding to the first corpus group decreases from greater than or equal to the first threshold to less than the first threshold, when deleting corpus in the second corpus set corresponding to the second corpus group using corpus in the second corpus set corresponding to the first corpus group.

[0108] In an implementation manner, the apparatus in the embodiment can further include the following units, as shown in Figure 6

[0109] The corpus moving unit 504 is configured to: obtain first associated corpus in the second corpus set corresponding to the first corpus group; obtain second associated corpus in the second corpus set corresponding to the second corpus group; the second associated corpus and the first associated corpus are derived from a target source document, and the number of target corpus generated in the target source document from the plurality of source documents corresponding to the system database satisfies a target screening condition; move the first associated corpus to the second corpus set corresponding to the second corpus group; and move the second associated corpus to the second corpus set corresponding to the first corpus group.

[0110] In an implementation manner, the corpus dividing unit 502 is further configured to: obtain at least one new corpus; perform word segmentation on the new corpus to obtain corpus keywords of each new corpus; obtain keyword repetition degrees of the corpus keywords in the first corpus set and the second corpus set corresponding to the corpus group to which the new corpus corresponds; and add the new corpus to the first corpus set or the second corpus set corresponding to the corpus group to which the new corpus corresponds according to the keyword repetition degrees.

[0111] In an implementation manner, the apparatus in the embodiment can further include the following units, as shown in Figure 7

[0112] ​​The query processing unit 505 is configured to obtain a target query request, wherein the target query request at least contains a query keyword; perform corpus query in a first corpus set and a second corpus set corresponding to a target corpus group respectively by using the query keyword, to obtain a first query result and a second query result; the target corpus group is a corpus group corresponding to a business type and the query keyword; sort the corpus in the first query result and the corpus in the second query result according to corpus similarity, to obtain a sorting result; and output the corpus in the first query result and the corpus in the second query result according to the sorting result.

[0113] It should be noted that the specific implementation of each unit in this embodiment can refer to the corresponding content in the foregoing, which will not be described in detail here.

[0114] Reference Figure 8 A structure schematic diagram of an electronic device provided in Embodiment Three of the present application can include:

[0115] The memory 801 is configured to store a computer program and data generated by running of the computer program.

[0116] The processor 802 is configured to execute the computer program to implement the following: obtaining a system database corresponding to an application system, wherein the application system can obtain corpus corresponding to a query request in the system database for the query request; the system database at least contains a plurality of corpus groups, and each corpus group corresponds to a business type; each corpus group at least contains a target corpus; for each corpus group, dividing the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set at least contains a first corpus, the second corpus set at least contains a second corpus, and the corpus hit rate of the first corpus is greater than the corpus hit rate of the second corpus; and adding or deleting corpus in a second corpus set of a second corpus group by using corpus in a second corpus set of a first corpus group, wherein the first corpus group and the second corpus group have an association relationship.

[0117] As can be known from the above solution, in the electronic device provided in Embodiment Three of the present application, first, the target corpus in the system database is divided into two corpus groups according to the business type and the corpus hit rate, and then the corpus in the corpus group corresponding to the associated business type is added or deleted by using the corpus in one of the corpus groups, so as to adjust the corpus in the corpus group corresponding to each business type in the system database, and further improve the use experience of the application system corresponding to the system database.

[0118] Taking the customer service module of the mobile bank as an example, in order to facilitate the selection of customers and improve the accuracy of the question and answer, multiple business modules are built in, such as mobile bank, loan, personal online bank, account management, etc., each business tag has different questions and answer corpus, but the corpus tag is basically manually annotated, and the corpus tag cannot be adjusted in time according to the use, which affects the accuracy of the customer service module and the corpus use experience to some extent.

[0119] Therefore, a mobile bank fuzzy boundary classification corpus adjustment scheme is established in the present application, which realizes the classification adjustment of the corpus by storing the fuzzy area of specific corpus and associating the data of multiple fuzzy areas, dynamically replicates or classifies the corpus according to the use of the customer, improves the search accuracy of the corpus, promotes the automation of corpus classification, and improves the use experience of the mobile bank customer service. Specifically, the present application mainly includes the following three parts:

[0120] 1. Fuzzy area filling: sort the hit frequency of the corpus based on the database corpus tag, and select specific corpus for fuzzy area filling.

[0121] 2. Fuzzy area association: associate multiple fuzzy areas corresponding to the corpus under multiple classification tags, and ensure that the corpus propagates between multiple areas.

[0122] 3. Corpus optimization adjustment: monitor the corpus of multiple classifications, and adjust the corpus across classifications and corpus regression.

[0123] The specific scheme is as follows:

[0124] Firstly, in the service process of the mobile bank customer service module, the corpus data in the system database is filtered according to the classification tag (business type), the corpus data obtained is divided into intervals according to the hit rate data of the mobile bank front end, the corpus data greater than the system threshold a (hit threshold) is placed in the regular data area (first corpus set, also called regular area), and the corpus data less than the system threshold a is stored in the fuzzy data area (second corpus set, also called fuzzy area).

[0125] For each classification area corpus data, set the fuzzy area size according to the data amount of the current classification by a certain proportion, for the data exceeding the fuzzy area, perform word segmentation on the corpus, calculate the repetition degree of the keywords of the fuzzy area corpus, move the data with low repetition degree to the regular data area, and the remaining data is filled into the fuzzy area.

[0126] Secondly, the correlation between each fuzzy area is established, the text similarity of the regular area data of multiple classifications and the text similarity of the fuzzy area data are calculated, the weighted average of the two is obtained to get the overall similarity of each classification corpus, and the overall similarity is sorted, and the fuzzy area correlation of adjacent classification corpus is established.

[0127] Data exchange is performed on the related fuzzy area, that is, fuzzy area M and fuzzy area N are associated, at this time, part of the data in fuzzy area M is exchanged to fuzzy area N, and the specific exchange method is that the corpus in the fuzzy area is compared with the corpus of the associated classification one by one, the containing relationship of the current corpus is calculated, and the data of the associated source document in the fuzzy area is exchanged.

[0128] Finally, when the mobile bank customer performs classification data retrieval application, the corpus is matched according to the following steps:

[0129] 1. The customer question and answer retrieval is performed on the regular data area of the classification corpus;

[0130] 2. The question and answer retrieval is performed on the corpus of the fuzzy area;

[0131] 3. The corpus in steps 1 and 2 is retrieved simultaneously, the similarity addition is performed, and the result after secondary sorting is returned;

[0132] 4. The customer retrieval use is recorded, the corpus in the regular data area less than the system threshold b is warned to be offline, the use frequency of the hit corpus data in the fuzzy area in multiple categories is recorded, if the use frequency of multiple categories is greater than the system threshold, the data in the fuzzy area is copied to two related classification areas, meanwhile, the primary and backup relationship of two corpora is established in the fuzzy area, for subsequent use that does not meet the conditions of the corpus in the fuzzy area, the primary and backup relationship is adjusted, and the corpus is returned.

[0133] It can be seen that the mobile bank fuzzy boundary classification corpus adjustment scheme provided by the application can set fuzzy areas for multiple classification corpora, establish association between multiple fuzzy areas, and monitor the use of the corpus during the use of the corpus to adjust the classification of the corpus.

[0134] The corpus data processing method, device and electronic equipment provided by the application can be used in big data or other fields, for example, can be used in corpus retrieval scenarios in the field of big data. Other fields are any fields except the financial field, for example, distributed field, cloud computing field, artificial intelligence field and Internet of Things field. The above are only examples, and do not limit the application field of the application name provided by the application.

[0135] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other. For the device disclosed by the embodiment, since it corresponds to the method disclosed by the embodiment, the description is relatively simple, and the related parts can be referred to the method part.

[0136] Those skilled in the art will further appreciate that the units and algorithms described in connection with the examples disclosed herein can be embodied directly in hardware, in software, or in a combination of the two. For the sake of brevity, descriptions of a method or an algorithm described in the preceding description will not be repeated in the following description of the examples. For the same reason, not all components and algorithms described in the examples will be repeated. It is to be understood that the above description is intended to be illustrative and not restrictive. Many other embodiments will be apparent to those of skill in the art upon reading the above description. The scope of the application should, therefore, be determined not with reference to the above description, but should instead be determined with reference to the appended claims, along with their full scope of equivalents.

[0137] The steps of a method or algorithm described in connection with the examples disclosed herein can be embodied directly in hardware, in software, or in a combination of the two. A software module can reside in Random Access Memory (RAM), non-volatile memory (e.g., Flash memory, ROM, EEPROM, EPROM, programmable ROM, etc.), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an application-specific integrated circuit (ASIC).

[0138] The above description is intended to be illustrative and not restrictive. Many other embodiments will be apparent to those of skill in the art upon reading the above description. The scope of the application should, therefore, be determined not with reference to the above description, but should instead be determined with reference to the appended claims, along with their full scope of equivalents. The disclosure of all articles and references are incorporated by reference in their entirety.

Claims

1. A method of processing corpus data, characterized by, The method comprises: obtaining a system database corresponding to an application system, the application system being capable of obtaining corpus corresponding to a query request in the system database for the query request; the system database comprising at least a plurality of corpus groups, each of the corpus groups corresponding to a business type; each of the corpus groups comprising at least one target corpus; for each of the corpus groups, dividing the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set comprising at least one first corpus, the second corpus set comprising at least one second corpus, the corpus hit rate of the first corpus being greater than the corpus hit rate of the second corpus; using corpus in the second corpus set corresponding to the first corpus group to add or delete corpus in the second corpus set corresponding to the second corpus group, the first corpus group and the second corpus group having an association relationship; wherein the method further comprises: obtaining a first associated corpus in the second corpus set corresponding to the first corpus group; obtaining a second associated corpus in the second corpus set corresponding to the second corpus group; the second associated corpus and the first associated corpus being derived from a target source document, the number of target corpus in the target source document among a plurality of source documents corresponding to the system database satisfying a target screening condition; moving the first associated corpus to the second corpus set corresponding to the second corpus group; moving the second associated corpus to the second corpus set corresponding to the first corpus group; wherein further comprising: obtaining at least one new corpus; performing word segmentation on the new corpus to obtain corpus keywords of each of the new corpus; obtaining keyword repetition degrees of the corpus keywords in the first corpus set and the second corpus set corresponding to the corpus group corresponding to the new corpus; according to the keyword repetition degrees, adding the new corpus to the first corpus set or the second corpus set corresponding to the corpus group corresponding to the new corpus.

2. The method of claim 1, wherein, The association relationship between the first corpus group and the second corpus group comprises: the corpus group similarity between the first corpus group and the second corpus group is greater than the corpus group similarity between the first corpus group and other corpus groups in the plurality of corpus groups; and, the corpus group similarity between the first corpus group and the second corpus group is greater than the corpus group similarity between the second corpus group and other corpus groups in the plurality of corpus groups.

3. The method of claim 2, wherein, The corpus group similarity between the first corpus group and the second corpus group is: the first set similarity and the second set similarity are weighted and averaged using respective corresponding weights to obtain an overall similarity; wherein the first set similarity is the similarity between the first corpus set in the first corpus group and the first corpus set in the second corpus group; and the second set similarity is the similarity between the second corpus set in the first corpus group and the second corpus set in the second corpus group.

4. The method according to claim 1 or 2, characterized in that, using corpus in the second corpus set corresponding to the first corpus group to add corpus in the second corpus set corresponding to the second corpus group, comprises: In the second corpus set corresponding to the first corpus group, third corpus with a corpus hit rate greater than or equal to a first threshold is obtained; and the third corpus is added to the second corpus set corresponding to the second corpus group.

5. The method according to claim 1 or 2, characterized in that, The corpus in the second corpus set corresponding to the first corpus group is used to delete the corpus in the second corpus set corresponding to the second corpus group, including: In the second corpus set corresponding to the second corpus group, the fourth corpus is deleted, the fourth corpus is the corpus added to the second corpus group from the second corpus set corresponding to the first corpus group under the condition that the corpus hit rate is greater than or equal to the first threshold, and the fourth corpus hit rate in the second corpus set corresponding to the first corpus group decreases from greater than or equal to the first threshold to less than the first threshold.

6. The method of claim 1 or 2, wherein, Also including: Obtaining a target query request, the target query request containing at least a query keyword; Using the query keyword to respectively query the first corpus set and the second corpus set corresponding to the target corpus group to obtain the first query result and the second query result; the target corpus group is a corpus group corresponding to the business type and the query keyword; The corpus in the first query result and the corpus in the second query result are sorted according to the corpus similarity to obtain a sorting result; According to the sorting result, the corpus in the first query result and the corpus in the second query result are output.

7. A processing apparatus of corpus data, characterized by, Including: A data obtaining unit is configured to obtain a system database corresponding to an application system, the application system being capable of obtaining corpus corresponding to a query request in the system database for the query request; The system database contains at least a plurality of corpus groups, each of the corpus groups corresponding to a business type; each of the corpus groups contains at least one target corpus; A corpus division unit is configured to divide, for each of the corpus groups, the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set contains at least one first corpus, the second corpus set contains at least one second corpus, and the corpus hit rate of the first corpus is greater than the corpus hit rate of the second corpus; A corpus processing unit is configured to use the corpus in the second corpus set corresponding to the first corpus group to add or delete the corpus in the second corpus set corresponding to the second corpus group, the first corpus group and the second corpus group having an association relationship; A corpus moving unit is configured to obtain a first associated corpus in the second corpus set corresponding to the first corpus group, and obtain a second associated corpus in the second corpus set corresponding to the second corpus group; the second associated corpus and the first associated corpus are derived from a target source document, and the number of target corpus in the target source document in a plurality of source documents corresponding to the system database satisfies a target screening condition; The first associated corpus is moved to the second corpus set corresponding to the second corpus group; The second associated corpus is moved to the second corpus set corresponding to the first corpus group; The corpus division unit is further configured to: obtain at least one new corpus; perform word segmentation on the new corpus to obtain corpus keywords of each new corpus; obtain keyword repetition degrees of the corpus keywords in the first corpus set and the second corpus set corresponding to the corpus group to which the new corpus corresponds; and add the new corpus to the first corpus set or the second corpus set corresponding to the corpus group to which the new corpus corresponds according to the keyword repetition degrees.

8. An electronic device, comprising: The computer program comprises: a memory configured to store a computer program and data generated by running of the computer program; a processor configured to execute the computer program to implement: obtaining a system database corresponding to an application system, the application system being capable of obtaining corpus corresponding to a query request in the system database according to the query request; the system database comprising at least a plurality of corpus groups, each corpus group corresponding to a business type; each corpus group comprising at least one target corpus; for each corpus group, dividing the target corpus into a first corpus set and a second corpus set according to a corpus hit rate of the target corpus; the first corpus set comprising at least one first corpus, the second corpus set comprising at least one second corpus, the corpus hit rate of the first corpus being greater than the corpus hit rate of the second corpus; and using the corpus in the second corpus set corresponding to a first corpus group to add or delete corpus in the second corpus set corresponding to a second corpus group, the first corpus group and the second corpus group having an association relationship; The processor is further configured to: obtain a first associated corpus in the second corpus set corresponding to the first corpus group, and obtain a second associated corpus in the second corpus set corresponding to the second corpus group; the second associated corpus and the first associated corpus being derived from a target source document, the target source document comprising target corpus satisfying a target screening condition among a plurality of source documents corresponding to the system database; move the first associated corpus to the second corpus set corresponding to the second corpus group; and move the second associated corpus to the second corpus set corresponding to the first corpus group. The processor is further configured to: obtain at least one new corpus; perform word segmentation on the new corpus to obtain corpus keywords of each new corpus; obtain keyword repetition degrees of the corpus keywords in the first corpus set and the second corpus set corresponding to the corpus group to which the new corpus corresponds; and add the new corpus to the first corpus set or the second corpus set corresponding to the corpus group to which the new corpus corresponds according to the keyword repetition degrees.

Citation Information

Patent Citations

  • Method and device for processing data and knowledge graph

    CN105893551A

  • Corpus establishing method and device

    CN110222192A