An update method, category determination method and device for a corpus processing model
Adjusting the corpus processing model parameters through grouping and correlation calculations, the problem of slow model update speed in the prior art is solved, and faster model updates and more efficient category determination are achieved.
Patent Information
- Application Number
- CN202011363647.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-11-27
AI Technical Summary
In the prior art, the corpus processing model has a slow update speed and a long update cycle, making it difficult to quickly adapt to changes in business categories.
By obtaining the current batch sample set, grouping according to the class object annotation information carried by the sample corpus, calculating the representation vector of the sample corpus and its internal and external correlation, adjusting the corpus processing model parameters to meet the model convergence conditions, and achieving rapid update of the model.
It improves the update efficiency of the corpus processing model, can respond to business category changes more quickly, and shortens the model update cycle.
Smart Images

Figure CN114564557B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a method and apparatus for updating a corpus processing model and a method for determining a category. Background Art
[0002] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that combines linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language that people use daily, so it has a close connection with the research of linguistics.
[0003] To determine the corresponding category for the corpus to be processed, the category can, to a certain extent, reflect the portrait description of the corpus to be processed. In the related art, a corpus processing model is often used when determining the corresponding category for the corpus to be processed. The corpus processing model uses the sample corpus with class annotation information as the modeling unit and focuses on the relationship between the sample corpus itself and its corresponding class annotation information. As time goes by and the relevant business categories change (such as adding new categories and deleting old categories), it is necessary to update the model based on the sample corpus with the changed class annotation information, which will result in a slow model update speed and a long model update cycle. Summary of the Invention
[0004] The present disclosure provides a method and apparatus for updating a corpus processing model and a method for determining a category, so as to at least solve the problems of slow model update speed and long update cycle in the related art. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a method for updating a corpus processing model is provided, including:
[0006] Obtain the current batch of sample sets;
[0007] Group according to the class annotation information carried by the sample corpus in the current batch of sample sets, so that the sample corpus carrying the same class annotation information is located in the same sample corpus group;
[0008] Obtain the representation vectors of the sample corpus based on the current corpus processing model;
[0009] Calculate the correlation degree between the representation vector of the sample corpus and the representation vectors of the sample corpus in the same group, and obtain the first correlation degree of the sample corpus, where the sample corpus in the same group is other sample corpus located in the same sample corpus group as the sample corpus;
[0010] Calculate the correlation between the representation vector of the sample corpus and the representation vector of the out-group sample corpus to obtain the second correlation of the sample corpus, where the out-group sample corpus is other sample corpora that are in different sample corpus groups from the sample corpus;
[0011] According to the first correlation and the second correlation, adjust the parameters of the current corpus processing model until the model convergence condition is met, and use the current corpus processing model that meets the model convergence condition as the target corpus processing model.
[0012] In an exemplary embodiment, the step of adjusting the parameters of the current corpus processing model according to the first correlation and the second correlation until the model convergence condition is met includes:
[0013] Obtain the current batch actual global correlation of the sample corpus according to the first correlation and the second correlation;
[0014] Obtain the current batch expected global correlation of the sample corpus;
[0015] Calculate the loss function value according to the current batch actual global correlation and the current batch expected global correlation;
[0016] Adjust the parameters of the current corpus processing model based on the loss function value until the model convergence condition is met.
[0017] In an exemplary embodiment, the step of obtaining the current batch actual global correlation of the sample corpus according to the first correlation and the second correlation includes:
[0018] Perform normalization processing on the first correlation and the second correlation to obtain the current batch actual global correlation.
[0019] In an exemplary embodiment, the step of obtaining the current batch sample set includes:
[0020] Receive a category expansion instruction, where the category expansion instruction includes a target category and a sample corpus carrying target category annotation information;
[0021] Construct the current batch sample set based on the sample corpus carrying target category annotation information.
[0022] In an exemplary embodiment, the step of obtaining the representation vector of the sample corpus based on the current corpus processing model includes:
[0023] Performing word segmentation on the sample corpus using the first corpus processing structure of the current corpus processing model to obtain at least two sample corpus segments, performing vector transformation on each of the sample corpus segments, and obtaining a matrix representing the sample corpus based on the vectors corresponding to each of the sample corpus segments;
[0024] Performing encoding processing on the matrix using the second corpus processing structure of the current corpus processing model to obtain a representation vector of the sample corpus.
[0025] According to a second aspect of the embodiments of the present disclosure, there is provided a category determination method, including:
[0026] Obtaining a to-be-processed corpus indicating a target object;
[0027] Using the target corpus processing model described in the first aspect with the to-be-processed corpus as input to obtain a representation vector of the to-be-processed corpus;
[0028] Based on the similarity between the representation vector of the to-be-processed corpus and multiple standard representation vectors, determining a standard representation vector that matches the representation vector of the to-be-processed corpus, each of the standard representation vectors carrying its corresponding class annotation information;
[0029] Determining the category of the target object based on the class annotation information corresponding to the matched standard representation vector.
[0030] In an exemplary implementation manner, before the step of determining a standard representation vector that matches the representation vector of the to-be-processed corpus based on the similarity between the representation vector of the to-be-processed corpus and multiple standard representation vectors, the method further includes a step of determining the multiple standard representation vectors;
[0031] The step of determining the multiple standard representation vectors includes:
[0032] Obtaining a standard corpus, where the standard corpus records standard corpora and the representation vectors of the standard corpora, each of the standard corpora carrying its corresponding class annotation information, and the representation vectors of the standard corpora are obtained using the target corpus processing model;
[0033] Performing word segmentation on the to-be-processed corpus to obtain at least two corpus segments;
[0034] Querying the standard corpus based on each of the corpus segments to obtain a standard corpus set corresponding to each of the corpus segments, and the standard corpora in the standard corpus set corresponding to the corpus segment all contain the corpus segment;
[0035] Obtaining a standard corpus collection according to the standard corpus sets corresponding to each of the corpus segments;
[0036] Determine at least two target standard corpora based on the occurrence frequencies of the respective standard corpora in the standard corpus collection;
[0037] Obtain the representation vectors of each of the target standard corpora based on the standard corpus, and use the representation vectors of the target standard corpora as the standard representation vectors.
[0038] In an exemplary embodiment, before the step of obtaining the standard corpus, a step of constructing an inverted index for the standard corpus is further included, where the inverted index is used to query the standard corpus containing the corpus fragment based on the corpus fragment;
[0039] Correspondingly, the step of querying the standard corpus based on each of the corpus fragments to obtain the standard corpus set corresponding to each corpus fragment includes:
[0040] Query the standard corpus based on the inverted index to obtain the standard corpus set corresponding to each corpus fragment.
[0041] In an exemplary embodiment, the step of constructing an inverted index for the standard corpus includes:
[0042] Perform word segmentation processing on multiple standard corpora respectively to obtain at least one corpus fragment corresponding to each standard corpus;
[0043] Respectively use each of the corpus fragments as an index keyword;
[0044] Determine the standard corpora containing each index keyword;
[0045] Construct a first corpus list based on the standard corpora containing each index keyword;
[0046] Construct an inverted index of each index keyword and the first corpus list corresponding to each index keyword.
[0047] In an exemplary embodiment, the method further includes adjusting the inverted index:
[0048] In response to receiving negative feedback on category determination, detect whether there is abnormal data in the inverted index;
[0049] When there is abnormal data in the inverted index, delete the index keyword corresponding to the abnormal data.
[0050] In an exemplary embodiment, before the step of obtaining the standard corpus, a step of establishing a mapping relationship for the standard corpus is further included, where the mapping relationship is used to query the category pointed to by the standard corpus based on the standard corpus;
[0051] Correspondingly, the step of determining the category of the target object based on the class annotation information corresponding to the matched standard representation vector includes:
[0052] Determine the standard corpus corresponding to the matched standard representation vector;
[0053] Query the standard corpus based on the mapping relationship, determine the category pointed to by the corresponding standard corpus, and use the pointed category as the category of the target object.
[0054] In an exemplary embodiment, the step of establishing the mapping relationship for the standard corpus includes:
[0055] Determine multiple categories based on the class annotation information carried by multiple standard corpora;
[0056] Determine the standard corpus pointing to each category;
[0057] Construct a second corpus list based on the standard corpus pointing to each category;
[0058] Establish a mapping relationship between each category and the second corpus list corresponding to each category.
[0059] According to the third aspect of the embodiments of the present disclosure, there is provided an update device for a corpus processing model, including:
[0060] A sample set acquisition unit configured to acquire the current batch of sample sets;
[0061] A grouping unit configured to group according to the class annotation information carried by the sample corpora in the current batch of sample sets, so that the sample corpora carrying the same class annotation information are located in the same sample corpus group;
[0062] A representation vector obtaining unit configured to obtain the representation vector of the sample corpus based on the current corpus processing model;
[0063] A first correlation calculation unit configured to calculate the correlation between the representation vector of the sample corpus and the representation vectors of the sample corpora in the same group, and obtain the first correlation of the sample corpus, where the sample corpora in the same group are other sample corpora located in the same sample corpus group as the sample corpus;
[0064] A second correlation calculation unit configured to calculate the correlation between the representation vector of the sample corpus and the representation vectors of the sample corpora in different groups, and obtain the second correlation of the sample corpus, where the sample corpora in different groups are other sample corpora located in different sample corpus groups from the sample corpus;
[0065] A parameter adjustment unit, configured to adjust the parameters of the current corpus processing model according to the first relevance and the second relevance to meet the model convergence condition, and use the current corpus processing model that meets the model convergence condition as the target corpus processing model.
[0066] In an exemplary embodiment, the parameter adjustment unit includes:
[0067] A current batch actual global relevance obtaining unit, configured to obtain the current batch actual global relevance of the sample corpus according to the first relevance and the second relevance;
[0068] A current batch expected global relevance obtaining unit, configured to obtain the current batch expected global relevance of the sample corpus;
[0069] A loss function value calculation unit, configured to calculate the loss function value according to the current batch actual global relevance and the current batch expected global relevance;
[0070] A parameter adjustment subunit, configured to adjust the parameters of the current corpus processing model based on the loss function value to meet the model convergence condition.
[0071] In an exemplary embodiment, the current batch actual global relevance obtaining unit includes:
[0072] A normalization processing unit, configured to perform normalization processing on the first relevance and the second relevance to obtain the current batch actual global relevance.
[0073] In an exemplary embodiment, the sample set obtaining unit includes:
[0074] A category expansion instruction receiving unit, configured to receive a category expansion instruction, where the category expansion instruction includes a target category and sample corpus carrying target category annotation information;
[0075] A sample set construction unit, configured to construct the current batch sample set based on the sample corpus carrying the target category annotation information.
[0076] In an exemplary embodiment, the feature vector obtaining unit includes:
[0077] A first corpus processing unit, configured to perform word segmentation on the sample corpus using the first corpus processing structure of the current corpus processing model to obtain at least two sample corpus segments, perform vector transformation on each sample corpus segment, and obtain a matrix representing the sample corpus based on the vectors corresponding to each sample corpus segment;
[0078] A second corpus processing unit, configured to perform encoding processing on the matrix by using a second corpus processing structure of the current corpus processing model to obtain a representation vector of the sample corpus.
[0079] According to a fourth aspect of the embodiments of the present disclosure, there is provided a category determination device, including:
[0080] A to-be-processed corpus acquisition unit, configured to perform acquisition of a to-be-processed corpus indicating a target object;
[0081] A model application unit, configured to perform taking the to-be-processed corpus as an input and using the target corpus processing model described in the first aspect to obtain a representation vector of the to-be-processed corpus;
[0082] A matching unit, configured to perform determination of a standard representation vector that matches the representation vector of the to-be-processed corpus based on the similarity between the representation vector of the to-be-processed corpus and a plurality of standard representation vectors, and each of the standard representation vectors carries corresponding category annotation information;
[0083] A category determination unit, configured to perform determination of the category of the target object based on the category annotation information corresponding to the matched standard representation vector.
[0084] In an exemplary embodiment, the device further includes a standard representation vector determination unit, and the standard representation vector determination unit includes:
[0085] A standard corpus acquisition unit, configured to perform acquisition of a standard corpus, where the standard corpus records standard corpora and representation vectors of the standard corpora, each of the standard corpora carries corresponding category annotation information, and the representation vectors of the standard corpora are obtained by using the target corpus processing model;
[0086] A first word segmentation processing unit, configured to perform word segmentation processing on the to-be-processed corpus to obtain at least two corpus segments;
[0087] A first query unit, configured to perform querying of the standard corpus based on each of the corpus segments to obtain a standard corpus set corresponding to each of the corpus segments, and the standard corpora in the standard corpus set corresponding to the corpus segment all contain the corpus segment;
[0088] A standard corpus aggregation obtaining unit, configured to perform obtaining of a standard corpus aggregation according to the standard corpus sets corresponding to each of the corpus segments;
[0089] A target standard corpus determination unit, configured to perform determination of at least two target standard corpora based on the occurrence frequencies of the standard corpora in the standard corpus aggregation;
[0090] A standard representation vector determination subunit, configured to obtain the representation vector of each target standard corpus based on the standard corpus, and use the representation vector of the target standard corpus as the standard representation vector.
[0091] In an exemplary embodiment, the apparatus further includes an inverted index construction unit, and the inverted index constructed by the inverted index unit is used to query the standard corpus containing the corpus segment based on the corpus segment;
[0092] Correspondingly, the first query unit includes:
[0093] A first query subunit, configured to query the standard corpus based on the inverted index to obtain a set of standard corpora corresponding to each corpus segment.
[0094] In an exemplary embodiment, the inverted index construction unit includes:
[0095] A second word segmentation processing unit, configured to perform word segmentation processing on a plurality of standard corpora respectively to obtain at least one corpus segment corresponding to each standard corpus;
[0096] An index keyword determination unit, configured to perform using each corpus segment as an index keyword respectively;
[0097] A first standard corpus determination unit, configured to determine the standard corpus containing each index keyword;
[0098] A first corpus list construction unit, configured to construct a first corpus list based on the standard corpus containing each index keyword;
[0099] An inverted index construction subunit, configured to construct an inverted index of each index keyword and the first corpus list corresponding to each index keyword.
[0100] In an exemplary embodiment, the apparatus further includes an inverted index adjustment unit, and the inverted index adjustment unit includes:
[0101] A negative feedback receiving unit, configured to detect whether there is abnormal data in the inverted index in response to the received category determination negative feedback;
[0102] A deletion unit, configured to delete the index keyword corresponding to the abnormal data when there is abnormal data in the inverted index.
[0103] In an exemplary embodiment, the device further includes a mapping relationship establishing unit, and the mapping relationship established by the mapping relationship establishing unit is used to query the category pointed to by the standard corpus based on the standard corpus;
[0104] Correspondingly, the category determining unit includes:
[0105] A second standard corpus determining unit, configured to determine the standard corpus corresponding to the matched standard representation vector;
[0106] A category determining subunit, configured to query the standard corpus based on the mapping relationship, determine the category pointed to by the corresponding standard corpus, and use the pointed category as the category of the target object.
[0107] In an exemplary embodiment, the mapping relationship establishing unit includes:
[0108] A standard corpus processing unit, configured to determine a plurality of categories based on the category annotation information carried by a plurality of standard corpora;
[0109] A third standard corpus determining unit, configured to determine the standard corpus pointing to each category;
[0110] A second corpus list constructing unit, configured to construct a second corpus list based on the standard corpus pointing to each category;
[0111] A mapping relationship establishing subunit, configured to establish a mapping relationship between each category and the second corpus list corresponding to each category.
[0112] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0113] A processor;
[0114] A memory for storing instructions executable by the processor;
[0115] Wherein, the processor is configured to execute the instructions to implement the method for updating the corpus processing model as described in the first aspect or the method for determining a category as described in the second aspect.
[0116] According to a sixth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method for updating the corpus processing model as described in the first aspect or the method for determining a category as described in the second aspect.
[0117] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product, which when running on a computer, causes the computer to execute the method for updating the corpus processing model described in the first aspect or the method for determining categories described in the second aspect.
[0118] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0119] Group the sample corpora in the current batch of sample sets according to the category annotation information, use the current corpus processing model to output the representation vectors of the sample corpora, obtain the first correlation degree and the second correlation degree based on the representation vectors of the sample corpora and the grouping information, and then adjust the parameters of the current corpus processing model according to the first correlation degree and the second correlation degree to obtain the target corpus processing model. The first correlation degree reflects the correlation degree between the sample corpora in the same group, and the second correlation degree reflects the correlation degree between the sample corpora in different groups. When using the current batch of sample sets to update the corpus processing model, paying attention to the first correlation degree and the second correlation degree, compared with paying attention to the sample corpora themselves and the category annotation information they carry, the above technical solution does not need to combine the sample sets of the previous batches to update the model, and can improve the efficiency of model update based on changes in relevant business categories. In addition, in an online application, use the target corpus processing model to output the representation vector of the corpus to be processed, determine the standard representation vector that matches the representation vector based on similarity calculation, and use the category annotation information carried by the standard representation vector to determine the category, which improves the efficiency of category determination.
[0120] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0121] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0122] Figure 1 is a flowchart of a method for updating a corpus processing model shown according to an exemplary embodiment.
[0123] Figure 2 is a flowchart of a method for determining categories shown according to an exemplary embodiment.
[0124] Figure 3 is a flowchart of determining multiple standard representation vectors shown according to an exemplary embodiment.
[0125] Figure 4 is a flowchart of constructing an inverted index for a standard corpus shown according to an exemplary embodiment.
[0126] Figure 5 It is a flowchart showing the adjustment of an inverted index according to an exemplary embodiment.
[0127] Figure 6 It is a flowchart showing the establishment of a mapping relationship for a standard corpus according to an exemplary embodiment.
[0128] Figure 7 It is a flowchart showing the adjustment of the parameters of the current corpus processing model according to a first relevance and a second relevance according to an exemplary embodiment.
[0129] Figure 8 It is a block diagram of an update device for a corpus processing model according to an exemplary embodiment.
[0130] Figure 9 It is a block diagram of a category determination device according to an exemplary embodiment.
[0131] Figure 10 It is an architecture design diagram for executing a category determination method according to an exemplary embodiment.
[0132] Figure 11 It is a block diagram of an electronic device according to an exemplary embodiment. Detailed implementation manners
[0133] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0134] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0135] The method for updating the corpus processing model provided by the present disclosure can be applied to a terminal or a server equipped with an update system or a category determination system of the corpus processing model. The category determination method provided by the present disclosure can be applied to a terminal or a server equipped with a category determination system. The terminal and the server can be connected through a wired network or a wireless network. The terminal can specifically be a smart phone, a desktop computer, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a digital assistant, a smart speaker, a smart wearable device, etc. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0136] Figure 1 is a flowchart of a method for updating a corpus processing model shown according to an exemplary embodiment, as Figure 1 shown, the method includes the following steps S101 to S106.
[0137] In step S101, obtain the current batch of sample sets.
[0138] The sample corpus in the current batch of sample sets carries category annotation information, and the category annotation information can indicate the category to which the sample corpus belongs. For example, the sample corpus 1 is "hard disk", and the category annotation information it carries can indicate the "digital accessory category"; the sample corpus 2 is "capsule coffee machine", and the category annotation information it carries can indicate the "kitchen small household appliance category". A sample corpus can carry at least one category annotation information, and the number of categories to which a sample corpus belongs can be greater than or equal to 1. When the number of categories to which it belongs is greater than or equal to 2, the relationship between the two categories to which it belongs can be a hierarchical relationship or a parallel relationship. For example, the sample corpus 3 is "skirt", and the category annotation information it carries can indicate the "women's clothing category" and the "bottom wear category"; the sample corpus 4 is "apple", and the category annotation information it carries can indicate the "fresh food category", the "fresh fruit category", the "mobile phone category", and the "audio-visual entertainment category". Of course, the sample corpus is not limited to natural language corpora in text form, but can also include natural language corpora in image form and voice form.
[0139] In one embodiment, step S101 may include step S1011 of receiving a category expansion instruction, where the category expansion instruction includes a target category and a sample corpus carrying target category annotation information; and step S1012 of constructing the current batch of sample sets based on the sample corpus carrying the target category annotation information.
[0140] The category expansion instruction can indicate the existence of new categories, and it is necessary to use the sample corpus corresponding to the new categories to update the model. The category expansion instruction can also indicate the existence of new sample corpora corresponding to the original categories, and it is necessary to use the new sample corpora to update the model, which can be regarded as the need to expand the sample corpus corresponding to the original categories. For example, the historical batch sample set includes: sample corpus a (belonging to category A), sample corpus b (belonging to category B), and sample corpus c (belonging to category C). The category expansion instruction 1 can indicate the existence of a new category: category D, and it is necessary to use the sample corpus d corresponding to category D to update the model. Here, the sample corpus corresponding to the new category can also be the original sample corpus, such as sample corpus a. The category expansion instruction 2 can indicate the existence of a new sample corpus corresponding to category B: sample corpus e, and it is necessary to use sample corpus e to update the model. Here, the new sample corpus corresponding to the original category can also be the original sample corpus, such as sample corpus a. In the above embodiments of the present disclosure, in response to the category expansion instruction, the current batch sample set is constructed based on the sample corpus carrying the annotation information indicating the target category, which can improve the flexibility and adaptability of constructing the current batch sample set. In the above embodiments of the present disclosure, the current batch sample set constructed based on the category expansion instruction can make the update of the current corpus processing model more flexible and adaptable, and can improve the timeliness of subsequent category recognition using the target corpus processing model.
[0141] In another embodiment, the sample corpus can be randomly obtained from the sample corpus library, and the current batch sample set is constructed based on the randomly obtained sample corpus. The sample corpus library can be divided into multiple sample corpus sets according to different categories, and the sample corpora in each sample corpus set point to the same category. When constructing the current batch sample corpus set, a preset number of sample corpora can be randomly selected from these sample corpus sets. For example, there are 200 categories in the sample corpus library. Correspondingly, there are 200 sample corpus sets. When the preset number is 2 and the target number of sample corpora in the current batch sample set is 128, 64 sample corpus sets are randomly selected from the 200 sample corpus sets, and 2 sample corpora are respectively selected from these 64 sample corpus sets, and then the current sample set composed of 128 sample corpora is constructed. The above method can also be used for the construction of each batch of sample sets. In practical applications, according to business requirements, the current batch sample set can be constructed by obtaining the sample corpus carrying the annotation information of the relevant business category. The business requirements can be to expand the product categories for an e-commerce platform, or to perform fine classification of plant and animal images for a plant and animal identification website, etc. The number of sample corpora in the current batch sample set can be 64 or 128, and of course, it can also be flexibly set according to actual needs.
[0142] In step S102, the sample corpora in the current batch of sample sets are grouped according to the category annotation information carried by the sample corpora, so that the sample corpora carrying the same category annotation information are in the same sample corpus group.
[0143] According to the category annotation information carried by the sample corpus, the sample corpora carrying the same category annotation information will be divided into the same sample corpus group, and correspondingly, the sample corpora carrying different category annotation information will be divided into different sample corpus groups.
[0144] In one embodiment, the sample corpus groups in the current batch of sample sets can be further processed. The number of sample corpus groups in the current batch of sample sets can be limited, and the number of sample corpora in each sample corpus group can be limited. For example, the number of sample corpus groups in the current batch of sample sets is limited to 32 or 64, and the number of sample corpora in each sample corpus group is limited to 2. Based on the above restrictions, the sample corpus used to input the current corpus processing model can be determined. When the sample corpus in the current batch of sample sets does not meet the above restrictions (for example, it is not enough to make up the 32 sample corpus groups that meet the requirements), the sample corpus used to input the current corpus processing model can be determined in combination with the sample corpus obtained later. When the sample corpus in the current batch of sample sets meets the above restrictions, but there are sample corpora that exceed the limit number (for example, it is enough to make up the 32 sample corpus groups that meet the requirements, and there are 6 sample corpora remaining), the sample corpora that meet the restrictions can be input into the current corpus processing model, and the sample corpora that exceed the limit number are combined with the sample corpora obtained later to construct the next batch of sample sets. It should be noted that the limit number of sample corpus groups can be flexibly set according to actual needs, and the limit number of sample corpora in each sample corpus group can be greater than or equal to 2.
[0145] In step S103, a representation vector of the sample corpus is obtained based on the current corpus processing model.
[0146] In one embodiment, step S103 may include step S1031, using the first corpus processing structure of the current corpus processing model to perform word segmentation processing on the sample corpus to obtain at least two sample corpus segments, performing vector conversion on the at least two sample corpus segments respectively, and obtaining a matrix representing the sample corpus based on the vector corresponding to each of the sample corpus segments; and step S1032, using the second corpus processing structure of the current corpus processing model to encode the matrix to obtain a representation vector of the sample corpus.
[0147] The sample corpus can be expressed in the form of sentences, phrases, words, etc. First, the sample corpus is segmented using the first corpus processing structure to obtain at least two sample corpus fragments, such as the two sample corpus fragments obtained by the sample corpus "wireless mouse" after the segmentation process are "wireless" and "mouse"; then, the at least two sample corpus fragments are respectively vectorized using the first corpus processing structure, such as converting the sample corpus fragment "wireless" into vector 1 and converting the sample corpus fragment "mouse" into vector 2; further, the first corpus processing structure is used to obtain a matrix representing the sample corpus based on the vector corresponding to each sample corpus fragment, such as obtaining matrix 1 representing the sample corpus "wireless mouse" based on vector 1 and vector 2; finally, the matrix is encoded using the second corpus processing structure to obtain a representation vector of the sample corpus, such as encoding matrix 1 to obtain a representation vector 1 of the sample corpus "wireless mouse". This representation vector is also called a hidden layer representation. The above-mentioned embodiment of the present disclosure combines the steps of word segmentation, vector conversion, and matrix encoding to achieve better and deeper mining of the sample corpus. The representation vector obtained in this way can more accurately reflect the semantics of the sample corpus and express the semantics of the sample corpus more comprehensively, especially for reflecting and expressing the more hidden semantics in the sample corpus.
[0148] In practical applications, "segmenting the sample corpus to obtain at least two sample corpus fragments" can be expressed as segmenting the sample corpus to obtain at least two sample corpus fragments, such as the four sample corpus fragments obtained by segmenting the sample corpus "wireless mouse" are "none", "line", "mouse" and "label". Correspondingly, the sample corpus fragments obtained by the segmentation are vectorized, such as "none" is converted into vector 11, "line" is converted into vector 12, "mouse" is converted into vector 21, and "label" is converted into vector 22; the vectors corresponding to the sample corpus fragments obtained by the segmentation are used to obtain a matrix representing the sample corpus, such as matrix 1' representing the sample corpus "wireless mouse" based on vector 11, vector 12, vector 21 and vector 22. Compared with the aforementioned word segmentation processing method, segmenting the sample corpus based on the word segmentation processing method can break through the limitation of the word segmentation processing method applied to sample corpora in the form of fixed phrases, words, etc., and refine the segmentation granularity to mine the representation vector of the sample corpus with a higher degree of restoration.
[0149] When the sample corpus is a natural corpus in the form of images or speech, it can be converted into a natural corpus in the form of text, and then the above-mentioned word segmentation (character) processing, vector conversion, and matrix encoding process can be performed to obtain a representation vector of the sample corpus.
[0150] The first corpus processing structure can correspond to a word vector tool (such as a word embedding table), and the second corpus processing structure can correspond to an encoder. When a sample corpus group includes two sample corpora, the two sample corpora respectively obtain their corresponding matrices through the word embedding table, and the two obtained matrices respectively pass through the same encoder to obtain their corresponding vectors, that is, the hidden layer representations of each sample corpus.
[0151] The construction basis of the current corpus processing model can include one of the following models: TextCNN (a convolutional neural network for text classification) model, fastText (a word vector and text classification tool open-sourced based on the word2vec model; the word2vec model, a related model for generating word vectors) model, and Transformer (a classic NLP model). The current corpus processing model can modify the model structure used. The second corpus processing structure in the current corpus processing model can directly utilize the encoder of the Transformer.
[0152] In step S104, calculate the correlation between the representation vector of the sample corpus and the representation vectors of the sample corpora in the same group, and obtain the first correlation of the sample corpus. The sample corpora in the same group are other sample corpora that are in the same sample corpus group as the sample corpus.
[0153] The first correlation reflects the actual correlation degree between a sample corpus and other sample corpora in the same group based on the calculation of the representation vector. That is to say, through the concept of the first correlation, the actual correlation degree between sample corpora carrying the same type of annotation information can be quantified. When a sample corpus group includes two sample corpora, calculate the similarity between the representation vectors of the two sample corpora, and then obtain the first correlation of each sample corpus. When a sample corpus group includes multiple sample corpora, calculate the similarity between the representation vector of the target sample corpus and the representation vectors of other sample corpora in the same group respectively, and then determine the first correlation of the target sample corpus based on the obtained multiple similarity results.
[0154] In step S105, calculate the correlation between the representation vector of the sample corpus and the representation vectors of the sample corpora in different groups, and obtain the second correlation of the sample corpus. The sample corpora in different groups are other sample corpora that are in different sample corpus groups from the sample corpus.
[0155] The second relevance reflects the actual relevance between a sample corpus and other sample corpora in different groups (heterogeneous groups) based on the calculation of the representation vectors. That is to say, through the concept of the second relevance, the actual relevance between sample corpora carrying different types of target annotation information can be quantified. When the number of other sample corpora in different groups is 1, calculate the similarity between the representation vector of the target sample corpus and the representation vector of this one sample corpus in a different group, and then obtain the second relevance of the target sample corpus. When the number of other sample corpora in different groups is greater than or equal to 2, calculate the similarities between the representation vector of the target sample corpus and the representation vectors of other sample corpora in different groups respectively, and then determine the second relevance of the target sample corpus based on at least 2 obtained similarity results.
[0156] The similarity metrics involved in step S104 and step S105 can be cosine similarity, Euclidean distance, relative entropy, etc. Of course, considering the generality of the data, the update method of the corpus processing model provided in this disclosure selects one metric, such as cosine similarity.
[0157] In step S106, according to the first relevance and the second relevance, adjust the parameters of the current corpus processing model until the model convergence condition is met, and use the current corpus processing model that meets the model convergence condition as the target corpus processing model.
[0158] Since the first relevance reflects the actual relevance between sample corpora in the same group, for sample corpora in the same group carrying the same type of target annotation information, the numerical value representing the highest similarity can be used to measure the expected relevance between sample corpora in the same group. Since the second relevance reflects the actual relevance between sample corpora in different groups, for sample corpora in different groups carrying different types of target annotation information, the numerical value representing the lowest similarity can be used to measure the expected relevance between sample corpora in the same group.
[0159] For example, the representation vectors of the sample corpora in the first sample corpus group are a1 (corresponding to sample corpus A1) and a2 (corresponding to sample corpus A2), the representation vectors of the sample corpora in the second sample corpus group are b1 (corresponding to sample corpus B1) and b2 (corresponding to sample corpus B2), the representation vectors of the sample corpora in the third sample corpus group are c1 (corresponding to sample corpus C1) and c2 (corresponding to sample corpus C2), and the representation vectors of the sample corpora in the fourth sample corpus group are d1 (corresponding to sample corpus D1) and d2 (corresponding to sample corpus D2). When the similarity metric is cosine similarity, the first correlation degree of sample corpus A1 is a1.a2, and the second correlation degree of sample corpus A1 is determined by a1.b1, a1.b2, a1.c1, a1.c2, a1.d1, and a1.d2. The expected correlation degree corresponding to the first relative degree can be taken as 1, and the expected correlation degree corresponding to the second relative degree can be taken as 0 (the expected similarity of each of a1.b1, a1.b2, a1.c1, a1.c2, a1.d1, and a1.d2 is 0), and the same applies to other sample corpora. Based on the difference between the first correlation degree of the sample corpus and the expected correlation degree corresponding to the first relative degree, and the difference between the second correlation degree of the sample corpus and the expected correlation degree corresponding to the second relative degree, the parameters of the current corpus processing model are adjusted to meet the model convergence condition, and then the current corpus processing model that meets the model convergence condition is used as the target corpus processing model.
[0160] In one embodiment, as Figure 7 shown, step S106 may include step S1061 of obtaining the current batch actual global correlation degree of the sample corpus according to the first correlation degree and the second correlation degree; step S1062 of obtaining the current batch expected global correlation degree of the sample corpus; step S1063 of calculating a loss function value according to the current batch actual global correlation degree and the current batch expected global correlation degree; and step S1064 of adjusting the parameters of the current corpus processing model based on the loss function value to meet the model convergence condition.
[0161] The current batch actual global correlation degree reflects the actual correlation degree of a sample corpus with the current batch sample set (or the sample corpus that meets the restricted conditions and is input into the current corpus processing model, refer to the relevant description in step S102) based on the calculation of the representation vectors. The current batch expected global correlation degree reflects the expected correlation degree of a sample corpus with the current batch sample set (or the sample corpus that meets the restricted conditions and is input into the current corpus processing model, refer to the relevant description in step S102) based on the calculation of the representation vectors.
[0162] The first relevance and the second relevance can be normalized respectively based on the same normalization rule, and then the actual global relevance of the current batch can be obtained by merging. For example, the normalization result of the similarity between the representation vectors of a sample corpus relative to the representation vectors of other sample corpora in the current batch of sample sets (or the sample corpus that satisfies the restricted conditions and is input into the current corpus processing model, refer to the relevant records in step S102) is integrated to obtain the actual global relevance of the current batch of this sample corpus. Then, the actual global relevance of the current batch of corresponding multiple sample corpora is merged to obtain the actual global relevance of the current batch. As a way to simplify calculations, normalization respectively converts the first relevance and the second relevance into scalar-form similarity normalization results, and then obtains the actual global relevance of the current batch in scalar form. Based on the scalar-form similarity normalization results, the operation efficiency between the first relevance and its corresponding expected similarity, and the operation efficiency between the second relevance and its corresponding expected similarity can be improved. Furthermore, the efficiency of constructing the loss function is improved, ensuring the timeliness of model update.
[0163] The expected global relevance of the current batch of a sample corpus can include the expected relevance between the representation vectors of the sample corpora in the same group (relative to the first relevance), and the expected relevance between the representation vectors of the sample corpora in different groups (relative to the second relevance). Correspondingly, based on the actual global relevance of the current batch of the sample corpus and the expected global relevance of the current batch, a loss function corresponding to the sample corpus is constructed, and then the parameters of the current corpus processing model are adjusted according to the loss function corresponding to each sample corpus and the method of gradient descent. In the above embodiments of the present disclosure, when adjusting the parameters of the current corpus processing model, based on the actual global relevance of the current batch and the expected global relevance of the current batch pointed to by the loss function, the model can fully learn the similarities between the sample corpora in the same group and the differences between the sample corpora in different groups, facilitating the effective extraction of representation vectors that can better reflect semantics from the corpus, and further improving the efficient category determination brought by applying the subsequent model to the category determination scenario.
[0164] Continuing with the example of the first sample corpus group - the fourth sample corpus group above, a similarity square matrix can be constructed based on the similarities between their representation vectors:
[0165] a1.a2 a1.b2 a1.c2 a1.d2 b1.a2 b1.b2 b1.c2 b1.d2 c1.a2 c1.b2 c1.c2 c1.d2 d1.a2 d1.b2 d1.c2 d1.d2
[0166] In practical applications, the first representation vector group and the second representation vector group can be constructed respectively. The first representation vector group includes a1, b1, c1, and d1, and the second representation vector group includes a2, b2, c2, and d2. The representation vectors in each representation vector group correspond to different categories from each other. The dot product operation is performed between each representation vector in the two representation vector groups, and then the above similarity matrix is obtained. Referring to the above similarity matrix, the similarity of pairwise representation vectors (a1.a2, b1.b2, c1.c2, and d1.d2) on the diagonal from the upper left to the lower right corresponds to the first correlation degree reflecting the actual correlation degree between the sample corpora in the same group, and the similarity of pairwise representation vectors in other positions corresponds to the second correlation degree reflecting the actual correlation degree between the sample corpora in different groups.
[0167] The Softmax function (a normalization exponential function) can be used to perform normalization processing on each row (column), and then the actual global correlation degree of the current batch of sample corpora is obtained. The loss function constructed based on the actual global correlation degree of the current batch and the expected global correlation degree of the current batch can be regarded as a cross-entropy loss function. When adjusting the parameters of the current corpus processing model according to the loss function, the adam optimizer (an optimizer) can be used. The current corpus processing model learns during the update that the correlation degree between the sample corpora in the same group should be higher and the correlation degree between the sample corpora in different groups should be lower, and completes the model update based on the current batch of sample sets based on this idea.
[0168] In the update method of the corpus processing model provided in the above embodiments, the sample corpora in the current batch of sample sets are grouped according to the class annotation information, the representation vectors of the sample corpora are output by using the current corpus processing model, the first correlation degree and the second correlation degree are obtained based on the representation vectors of the sample corpora and the grouping information, and then the parameters of the current corpus processing model are adjusted according to the first correlation degree and the second correlation degree to obtain the target corpus processing model. The first correlation degree reflects the correlation degree between the sample corpora in the same group, and the second correlation degree reflects the correlation degree between the sample corpora in different groups. When updating the corpus processing model by using the current batch of sample sets, paying attention to the first correlation degree and the second correlation degree, compared with paying attention to the sample corpora themselves and the class annotation information they carry, the above technical solution does not need to combine the sample sets of the previous batches to perform model update, and can improve the efficiency of model update based on changes in relevant business categories. When there is a need for category expansion, the corresponding sample corpora can be added under the new category to construct the current batch of sample sets, and then the current corpus processing model is updated based on the current batch of sample sets to achieve dynamic update of the model.
[0169] Figure 2 is a flowchart of a category determination method shown according to an exemplary embodiment, as Figure 2As shown, the method includes the following steps S201 to S204.
[0170] In step S201, the to-be-processed corpus indicating the target object is obtained.
[0171] The to-be-processed corpus indicating the target object can be regarded as a corpus that describes the features of the target object. The to-be-processed corpus can be a natural corpus in text form, such as the to-be-processed corpus presented in the form of sentences, phrases, or words. The to-be-processed corpus can also be a natural corpus in image form or voice form.
[0172] In practical applications, according to different business scenarios, the target object can indicate different contents. When the business scenario involves an e-commerce platform, the target object can indicate a product with an undetermined category. When the business scenario involves an animal and plant identification website, the target object can indicate an animal or plant with an undetermined category.
[0173] In step S202, using the to-be-processed corpus as the input, the representation vector of the to-be-processed corpus is obtained by using the target corpus processing model provided by the present disclosure.
[0174] Input the to-be-processed corpus into the target corpus processing model, and use the target corpus processing model to output the representation vector of the to-be-processed corpus. The process of the target corpus processing model outputting the representation vector of the to-be-processed corpus can refer to the description in the foregoing step S103 and will not be elaborated here.
[0175] In step S203, based on the similarity between the representation vector of the to-be-processed corpus and multiple standard representation vectors, the standard representation vector that matches the representation vector of the to-be-processed corpus is determined, and each standard representation vector carries its corresponding class annotation information.
[0176] The standard representation vector is the representation vector of the standard corpus. The standard corpus can be a representative corpus reflecting a preset category, such as the preset category "watch" and the standard corpus "Rolex". There can be certain differences between the standard corpora pointing to the same preset category. For example, for the preset category "vinegar", the standard corpora are "aged vinegar" and "aromatic vinegar".
[0177] The standard representation vector is obtained by inputting the standard corpus into the target corpus processing model. After determining the target corpus processing model, the standard corpus can be input into the model, and a standard corpus library is constructed based on the obtained representation vector and the standard corpus. The standard corpus library can be stored in redis (a key-value pair database). In practical applications, sample corpora can also be used as standard corpora and put into the standard corpus library. Correspondingly, the representation vectors of the sample corpora are used as standard representation vectors and put into the standard corpus library. The representation vectors of the sample corpora are obtained by using the target corpus processing model.
[0178] In one embodiment, as Figure 3 shown, before step S203, it further includes determining the multiple standard representation vectors: Step S301, obtaining a standard corpus, which records standard corpus and the representation vectors of the standard corpus, each standard corpus carries its corresponding class annotation information, and the representation vector of the standard corpus is obtained by using the target corpus processing model; Step S302, performing word segmentation processing on the corpus to be processed to obtain at least two corpus segments; Step S303, querying the standard corpus based on each corpus segment to obtain the standard corpus set corresponding to each corpus segment, and the standard corpus in the standard corpus set corresponding to the corpus segment all contain the corpus segment; Step S304, obtaining a standard corpus collection according to the standard corpus set corresponding to each corpus segment; Step S305, determining at least two target standard corpus based on the occurrence frequencies of the various standard corpus in the standard corpus collection; Step S306, obtaining the representation vector of each target standard corpus based on the standard corpus, and taking the representation vector of the target standard corpus as the standard representation vector.
[0179] The form of the corpus to be processed can be a sentence, a phrase, etc. Performing word segmentation processing on the corpus to be processed to obtain corpus segments can be performed based on the jieba word segmentation component (a word segmentation component). The "word segmentation processing" involved here can also be manifested as "character segmentation processing", especially for the corpus to be processed in the form of fixed phrases or words. When the corpus to be processed is a natural corpus in the form of an image or speech, it can be converted into a natural corpus in text form, and then the above word (character) segmentation processing can be performed to obtain corpus segments. For example, after the corpus to be processed is subjected to word (character) segmentation processing, corpus segments 1, 2, and 3 are obtained.
[0180] When querying the standard corpus based on the corpus segment and determining the standard corpus containing the corpus segment in the standard corpus, the standard corpus set corresponding to each corpus segment can also be obtained. If the standard corpus containing corpus segment 1 in the standard corpus are standard corpus 1, 2, and 3, then standard corpus 1, 2, and 3 constitute the standard corpus set 1 corresponding to corpus segment 1; the standard corpus containing corpus segment 2 in the standard corpus are standard corpus 2 and 4, then standard corpus 2 and 4 constitute the standard corpus set 2 corresponding to corpus segment 2; the standard corpus containing corpus segment 3 in the standard corpus are standard corpus 2 and 3, then standard corpus 2 and 3 constitute the standard corpus set 3 corresponding to corpus segment 3.
[0181] According to the corresponding standard corpus sets for each corpus segment, a standard corpus collection is obtained, that is, the union of these standard corpus sets is obtained. For example, the standard corpus collection obtained from the above standard corpus sets 1-3 includes standard corpus 1, 2, 3, and 4. In the standard corpus collection, the occurrence frequency of standard corpus 1 is 1, the occurrence frequency of standard corpus 2 is 3, the occurrence frequency of standard corpus 3 is 2, and the occurrence frequency of standard corpus 4 is 1. Based on the occurrence frequencies of the various standard corpora in the standard corpus collection, at least two target standard corpora can be determined. Then, sorting by the occurrence frequency from high to low, we have standard corpus 2 > standard corpus 3 > standard corpus 1 = standard corpus 4. Standard corpus 2 and standard corpus 3 can be used as the target standard corpora. Of course, standard corpora 1-4 can also be used as the target standard corpora. Here, the number of target standard corpora can be determined according to a preset number. Correspondingly, the representation vectors of the target standard corpora can be obtained, and the representation vectors of the target standard corpora are used as the standard representation vectors.
[0182] Correspondingly, step S203 may include the following steps: calculating the similarity between the representation vector of each of the target standard corpora and the representation vector of the to-be-processed corpus respectively, and determining the standard representation vector that matches the representation vector of the to-be-processed corpus based on the similarity calculation result. For example, taking standard corpus 2 and standard corpus 3 as the target standard corpora, the representation vector of standard corpus 2 is standard representation vector 2, and the representation vector of standard corpus 3 is standard representation vector 3. Calculate the similarity between the representation vector of the to-be-processed corpus and standard representation vector 2 to obtain similarity 2, calculate the similarity between the representation vector of the to-be-processed corpus and standard representation vector 3 to obtain similarity 3, and determine the standard representation vector that is more matched with the representation vector of the to-be-processed corpus from standard representation vector 2 and standard representation vector 3 according to similarity 2 and similarity 3. Here, the similarity metrics involved in the similarity calculation can use cosine similarity, Euclidean distance, relative entropy, etc.
[0183] In the above embodiments of the present disclosure, when selecting multiple standard representation vectors, the corpus segments of the to-be-processed corpus are used for the initial screening of the corpus in the standard corpus library, and then the target standard corpora are determined based on the occurrence frequencies of the standard corpora to integrate the initial screening results. Furthermore, the representation vectors of the target standard corpora are used as the standard representation vectors. Compared with all the standard representation vectors in the standard corpus library, this can effectively narrow the range of the standard representation vectors to be matched, improve the matching efficiency, and at the same time reduce the consumption of computing resources. Compared with other standard corpora in the standard corpus library, the standard corpora corresponding to the standard representation vectors are more relevant to the to-be-processed corpus, which can ensure the accuracy of subsequent category determination.
[0184] In practical applications, cosine similarity calculations can also be performed separately on the representation vectors of the corpus to be processed and the retrieved target standard corpus representation vectors. The TOPN categories with the highest matching degree can be selected and returned from the category set corresponding to the retrieved target standard corpus according to the sorting result of the obtained similarity levels and the occurrence frequency of the hit categories, where N is a positive integer greater than 0.
[0185] In another embodiment, for the standard corpus, the present disclosure also provides an embodiment for constructing an inverted index for the standard corpus. The construction and application of the inverted index will be introduced separately below.
[0186] 1) As Figure 4 shown, the steps for constructing an inverted index for the standard corpus include: Step S401, performing word segmentation processing on multiple standard corpora respectively to obtain at least one corpus segment corresponding to each of the standard corpora; Step S402, taking each of the corpus segments as an index keyword respectively; Step S403, determining the standard corpora containing each index keyword; Step S404, constructing a first corpus list based on the standard corpora containing each index keyword; Step S405, constructing an inverted index of each index keyword and the first corpus list corresponding to each index keyword.
[0187] First, the representation form of the standard corpus can be a sentence, a phrase, etc. Performing word segmentation processing on the standard corpus to obtain corresponding corpus segments can be performed based on the jieba word segmentation component (a word segmentation component). The "word segmentation processing" involved here can also be manifested as "character segmentation processing", especially for standard corpora in the form of fixed phrases or words. When the standard corpus is a natural corpus in image form or speech form, it can be converted into a natural corpus in text form, and then the above-mentioned word (character) segmentation processing can be performed to obtain standard corpus segments. For example, after word (character) segmentation processing of standard corpus 1, standard corpus segments 1, 2, and 3 are obtained; after word (character) segmentation processing of standard corpus 2, standard corpus segments 1 and 4 are obtained; after word (character) segmentation processing of standard corpus 3, standard corpus segment 2 is obtained.
[0188] Then, the corpus segments can be used as index keywords. Standard corpus segment 1 can be used as index keyword 1, standard corpus segment 2 can be used as index keyword 2, standard corpus segment 3 can be used as index keyword 3, and standard corpus segment 4 can be used as index keyword 4.
[0189] Furthermore, the standard corpus containing each index keyword can be determined, and then a first corpus list can be constructed based on the standard corpus containing the index keyword. The first corpus list 1 is constructed based on the standard corpus 1 and the standard corpus 2 containing the index keyword 1, the first corpus list 2 is constructed based on the standard corpus 1 and the standard corpus 3 containing the index keyword 2, the first corpus list 1 is constructed based on the standard corpus 1 containing the index keyword 3, and the first corpus list 4 is constructed based on the standard corpus 2 containing the index keyword 4.
[0190] Next, an inverted index of each index keyword and the first corpus list corresponding to each index keyword is constructed. The index keyword can be used as the key, and the first corpus list corresponding to the index keyword can be used as the value, and then an inverted index is established with key-value. For example, the index keyword 1 is used as the key, and the first corpus list 1 is used as the value to construct the inverted index 1. Correspondingly, the inverted index 2 of the index keyword 2 and the first corpus list 2 is constructed, the inverted index 3 of the index keyword 3 and the first corpus list 3 is constructed, and the inverted index 4 of the index keyword 4 and the first corpus list 4 is constructed.
[0191] In the above embodiment of constructing the inverted index of the present disclosure, the corpus segment of the standard corpus in the standard corpus library is used as the index keyword, and the first corpus list is obtained based on the standard corpus containing the index keyword, so as to construct an inverted index of each index keyword and the first corpus list corresponding to each index keyword. Further, the corpus segment can be used as the query object to perform a corpus segment matching query in the inverted index. With the update of the standard corpus in the standard corpus library, the update of the original inverted index also has good flexibility and real-time performance, ensuring the accuracy and efficiency of the corpus segment matching query.
[0192] Furthermore, as Figure 5 shown, the constructed inverted index can also be adjusted according to business requirements: Step S501, in response to receiving the category determination negative feedback, detect whether there is abnormal data in the inverted index; Step S502, when there is abnormal data in the inverted index, delete the index keyword corresponding to the abnormal data.
[0193] The category determination negative feedback can be generated in the test link or in the online application. The category determination negative feedback indicates that there is a difference between the category determined by the system and the expected category. The manifestation of this difference can be that the category determined by the system ("food") is too broad relative to the expected category ("drip coffee"), or there is no intersection between the category determined by the system ("food") and the expected category ("jewelry"), etc. When there is a difference between the category of the target object determined in step S204 and the expected category, the category determination negative feedback can be generated by reporting an error.
[0194] Based on the category to determine negative feedback, abnormal data detection can be performed on the inverted index. The abnormal data is the data in the inverted index that causes the above differences. The abnormal data can be dirty data, and dirty data is often inevitable. When abnormal data is detected, the index keywords corresponding to the abnormal data can be deleted. The abnormal data can exist in the index keywords in the inverted index or in the first corpus list. Deleting the index keywords corresponding to the abnormal data can dynamically adjust the errors between categories and effectively avoid the influence of abnormal data on category determination when retrieving based on the inverted index, especially for some dirty data that significantly affects the online effect. According to business requirements, the index keywords of the inverted index can also be dynamically fine-tuned without updating the model, which can further reduce the coupling between offline model updates and online category determination applications and improve the flexibility of online category determination applications.
[0195] 2) The inverted index can be used to query the standard corpus containing the corpus fragment based on the corpus fragment. When implementing step S303 to obtain the standard corpus set corresponding to each corpus fragment, it can be obtained by querying the standard corpus library based on the inverted index. For the corpus fragments obtained by tokenizing (characterizing) the corpus to be processed, these corpus fragments can be retrieved based on the inverted index to obtain the standard corpus containing these corpus fragments, which can effectively avoid the consumption of computing resources by invalid calculations, improve computing efficiency, and improve the online feedback time.
[0196] In step S204, the category of the target object is determined based on the class annotation information corresponding to the matched standard representation vector.
[0197] The class annotation information carried by the standard representation vector indicates a preset category. First, the standard representation vector and its corresponding preset category can be stored; after determining the matched standard representation vector, the preset category corresponding to the matched standard representation vector is obtained based on the corresponding relationship in the stored data; and this preset category is used as the category of the target object.
[0198] In one embodiment, for the standard corpus library, the present disclosure also provides an embodiment for establishing a mapping relationship for the standard corpus library. The establishment and application of the mapping relationship will be introduced separately below.
[0199] 1) As Figure 6 shown, the steps for establishing a mapping relationship for the standard corpus library include: step S601, determining multiple categories based on the class annotation information carried by multiple standard corpora; step S602, determining the standard corpora pointing to each category; step S603, constructing a second corpus list based on the standard corpora pointing to each category; step S604, establishing a mapping relationship between each category and the second corpus list corresponding to each category.
[0200] First, based on the class annotation information carried by multiple standard corpora, multiple categories are determined. The class annotation information carried by multiple standard corpora can be different from each other, and there may also be two standard corpora carrying the same class annotation information among multiple standard corpora. Determining multiple categories here can be regarded as constructing a category library based on the class annotation information carried by multiple standard corpora.
[0201] Then, the standard corpora pointing to each category are determined, and a second corpus list is constructed based on the standard corpora pointing to each category. For example, the class annotation information carried by standard corpus 1 can indicate category 1, the class annotation information carried by standard corpus 2 can indicate category 2, the class annotation information carried by standard corpus 3 can indicate category 3, the class annotation information carried by standard corpus 4 can indicate category 1, and the class annotation information carried by standard corpus 5 can indicate category 2. Then, a second corpus list 1 is constructed based on standard corpus 1 and standard corpus 4 pointing to category 1, a second corpus list 2 is constructed based on standard corpus 2 and standard corpus 5 pointing to category 2, and a second corpus list 3 is constructed based on standard corpus 3 pointing to category 3.
[0202] Furthermore, a mapping relationship between each category and the second corpus list corresponding to each category is established. The category can be used as the key, and the second corpus list corresponding to the category can be used as the value, and then a mapping relationship is established with key-value. For example, category 1 is used as the key and the second corpus list 1 is used as the value to establish mapping relationship 1. Correspondingly, mapping relationship 2 between category 2 and the second corpus list 2 is established, and mapping relationship 3 between category 3 and the second corpus list 3 is established.
[0203] In the above embodiment of establishing the mapping relationship of the present disclosure, categories are determined based on the class annotation information carried by the standard corpora in the standard corpus library, and the second corpus list is obtained based on the standard corpora including the categories, so as to establish the mapping relationship between each category and the second corpus list corresponding to each category. Further, the standard corpora can be used as the query object to perform query in the mapping relationship in the standard corpus matching manner. With the update of the standard corpora in the standard corpus library, the update of the original mapping relationship also has good flexibility and real-time performance, ensuring the accuracy and efficiency of the standard corpus matching manner query.
[0204] 2) The mapping relationship is used to query the category pointed to by the standard corpus based on the standard corpus. Step S204 may include the following steps: determining the standard corpus corresponding to the matched standard representation vector; querying the standard corpus library based on the mapping relationship to determine the category pointed to by the corresponding standard corpus, and using the pointed category as the category of the target object. On the basis of determining the standard representation vector that matches the representation vector of the corpus to be processed, the standard corpus corresponding to the matched standard representation vector is determined, and then the category corresponding to the standard corpus is determined based on the mapping relationship, so as to use this category as the category of the target object. For the standard corpus corresponding to the matched standard representation vector, the corresponding category can be determined by retrieving the standard corpus based on the mapping relationship, which can effectively avoid the consumption of computing resources caused by invalid calculations, improve the computing efficiency, and improve the online feedback time.
[0205] In an exemplary embodiment, as Figure 10 shown, the category determination method provided by the present disclosure can be applied to the scenario of e-commerce category retrieval and recommendation. For example, when a merchant uploads a product to an e-commerce platform, the merchant enters product feature keywords, and the e-commerce platform recommends the category to which the product belongs based on the product feature keywords. When there is a need for category expansion, corresponding sample corpora can be added under the newly added categories to construct the current batch of sample sets, and then the current corpus processing model is updated based on the current batch of sample sets. The update of the current corpus processing model is performed in an offline environment. The online service part includes the representation vectors of the training entries and the inverted index constructed for the training entries. The online response part responds to the received online request entries and performs category determination together with the online service part.
[0206] E-commerce platforms often involve hundreds of categories, and each category contains hundreds of keywords or phrases. By using the method for updating the corpus processing model provided by the present disclosure, dynamic and efficient model updates can be achieved when there is a need for category expansion, and the impact on online applications caused by category expansion can be reduced. Further, for e-commerce platforms with imperfect categories, it is often necessary to frequently perform category expansion, category modification, etc. By using the method for updating the corpus processing model provided by the present disclosure, the model update cycle can be reduced, and the computing resource overhead brought by model updates can be reduced.
[0207] In the category determination method provided in the above embodiment, the representation vector of the corpus to be processed is output by using the target corpus processing model, the standard representation vector that matches the representation vector is determined based on similarity calculation, and the category determination is performed by using the category annotation information carried by the standard representation vector, which improves the efficiency of category determination.
[0208] Figure 8 is a block diagram of an apparatus for updating a corpus processing model shown according to an exemplary embodiment. Refer to Figure 8, the apparatus includes a sample set acquisition unit 810, a grouping unit 820, a feature vector obtaining unit 830, a first correlation calculation unit 840, a second correlation calculation unit 850, and a parameter adjustment unit 860.
[0209] The sample set acquisition unit 810 is configured to acquire the current batch of sample sets.
[0210] The grouping unit 820 is configured to group according to the class annotation information carried by the sample corpus in the current batch of sample sets, so that the sample corpora carrying the same class annotation information are located in the same sample corpus group;
[0211] The feature vector obtaining unit 830 is configured to obtain the feature vector of the sample corpus based on the current corpus processing model;
[0212] The first correlation calculation unit 840 is configured to calculate the correlation between the feature vector of the sample corpus and the feature vectors of the sample corpora in the same group, and obtain the first correlation of the sample corpus, where the sample corpora in the same group are other sample corpora located in the same sample corpus group as the sample corpus;
[0213] The second correlation calculation unit 850 is configured to calculate the correlation between the feature vector of the sample corpus and the feature vectors of the sample corpora in different groups, and obtain the second correlation of the sample corpus, where the sample corpora in different groups are other sample corpora located in different sample corpus groups from the sample corpus;
[0214] The parameter adjustment unit 860 is configured to adjust the parameters of the current corpus processing model according to the first correlation and the second correlation until the model convergence condition is satisfied, and use the current corpus processing model that satisfies the model convergence condition as the target corpus processing model.
[0215] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0216] Figure 9 is a block diagram of a category determination apparatus shown according to an exemplary embodiment. Refer to Figure 9 , the apparatus includes a to-be-processed corpus acquisition unit 910, a model application unit 920, a matching unit 930, and a category determination unit 940.
[0217] The to-be-processed corpus acquisition unit 910 is configured to acquire the to-be-processed corpus indicating the target object;
[0218] The model application unit 920 is configured to execute to obtain a representation vector of the to-be-processed corpus by using the current corpus processing model with the to-be-processed corpus as the input as described in the foregoing steps S101 to S106;
[0219] The matching unit 930 is configured to execute to determine a standard representation vector that matches the representation vector of the to-be-processed corpus based on the similarity between the representation vector of the to-be-processed corpus and multiple standard representation vectors, and each of the standard representation vectors carries its corresponding class annotation information;
[0220] The category determination unit 940 is configured to execute to determine the category of the target object based on the class annotation information corresponding to the matched standard representation vector.
[0221] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0222] In an exemplary embodiment, an electronic device is further provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored on the memory, the steps of any of the corpus processing model update methods or category determination methods in the above embodiments are implemented.
[0223] The electronic device may be a terminal, a server, or a similar computing device. Taking the electronic device being a server as an example, Figure 11FIG. 0 is a block diagram of an electronic device for updating a corpus processing model or determining a category according to an exemplary embodiment. The electronic device 1100 may vary significantly depending on configuration or performance, and may include one or more central processing units (CPUs) 1110 (the processor 1110 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA), a memory 1130 for storing data, and one or more storage media 1120 (such as one or more mass storage devices) for storing application programs 1123 or data 1122. Among them, the memory 1130 and the storage media 1120 may be transient storage or persistent storage. The program stored in the storage media 1120 may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the central processor 1110 may be configured to communicate with the storage media 1120 and execute a series of instruction operations in the storage media 1120 on the electronic device 1100. The electronic device 1100 may further include one or more power supplies 1160, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1140, and / or one or more operating systems 1121, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0224] The input / output interface 1140 may be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the electronic device 1100. In one example, the input / output interface 1140 includes a network interface controller (NIC), which may be connected to other network devices through a base station and thus communicate with the Internet. In an exemplary embodiment, the input / output interface 110 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0225] Those of ordinary skill in the art can understand that Figure 11 The structure shown is only illustrative and does not limit the structure of the above electronic device. For example, the electronic device 1100 may further include more or fewer components than Figure 11 shown, or have a different configuration from Figure 11 shown.
[0226] In an exemplary embodiment, a storage medium is further provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the steps of the update method or the category determination method of any corpus processing model in the above embodiments.
[0227] In an exemplary embodiment, a computer program product is further provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the update method or the category determination method provided in any of the above embodiments.
[0228] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0229] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0230] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for updating a corpus processing model, characterized in that, the method includes: Obtain the current batch of sample sets; Group according to the class annotation information carried by the sample corpus in the current batch of sample sets, so that the sample corpora carrying the same class annotation information are located in the same sample corpus group; Obtain the representation vector of the sample corpus based on the current corpus processing model; Calculate the correlation degree between the representation vector of the sample corpus and the representation vectors of the sample corpora in the same group, and obtain the first correlation degree of the sample corpus. The sample corpora in the same group are other sample corpora that are in the same sample corpus group as the sample corpus; Calculate the correlation degree between the representation vector of the sample corpus and the representation vectors of the sample corpora in different groups, and obtain the second correlation degree of the sample corpus. The sample corpora in different groups are other sample corpora that are in different sample corpus groups from the sample corpus; Perform normalization processing on the first correlation degree and the second correlation degree to obtain the actual global correlation degree of the current batch; obtain the expected global correlation degree of the current batch of the sample corpus; calculate the loss function value according to the actual global correlation degree of the current batch and the expected global correlation degree of the current batch; and adjust the parameters of the current corpus processing model based on the loss function value until the model convergence condition is satisfied; Use the current corpus processing model that satisfies the model convergence condition as the target corpus processing model.
2. The method according to claim 1, characterized in that, the obtaining the current batch of sample sets includes: Receive a category expansion instruction, where the category expansion instruction includes a target category and a sample corpus carrying target class annotation information; Construct the current batch of sample sets based on the sample corpus carrying the target class annotation information.
3. The method according to claim 1, characterized in that, the obtaining the representation vector of the sample corpus based on the current corpus processing model includes: Use the first corpus processing structure of the current corpus processing model to perform word segmentation on the sample corpus to obtain at least two sample corpus segments, perform vector transformation on each sample corpus segment, and obtain a matrix representing the sample corpus based on the vectors corresponding to each sample corpus segment; Use the second corpus processing structure of the current corpus processing model to perform encoding processing on the matrix to obtain the representation vector of the sample corpus.
4. A method for determining a category, characterized in that, the method includes: Obtain the corpus to be processed indicating the target object; Use the target corpus processing model according to any one of claims 1 to 3 as the input with the corpus to be processed to obtain the representation vector of the corpus to be processed; Based on the similarity between the representation vector of the corpus to be processed and multiple standard representation vectors, determine the standard representation vector that matches the representation vector of the corpus to be processed. Each standard representation vector carries its corresponding class annotation information; Determine the category of the target object based on the class annotation information corresponding to the matched standard representation vector.
5. The method according to claim 4, characterized in that, Before determining the standard representation vector that matches the representation vector of the to-be-processed corpus based on the similarity between the representation vector of the to-be-processed corpus and multiple standard representation vectors, the method further includes a step of determining the multiple standard representation vectors; The step of determining the multiple standard representation vectors includes: Obtain a standard corpus, where the standard corpus records standard corpora and the representation vectors of the standard corpora, each standard corpus carries its corresponding class annotation information, and the representation vector of the standard corpus is obtained by using the target corpus processing model; Perform word segmentation on the to-be-processed corpus to obtain at least two corpus segments; Query the standard corpus based on each corpus segment to obtain the standard corpus set corresponding to each corpus segment, and the standard corpora in the standard corpus set corresponding to the corpus segment all contain the corpus segment; Obtain a standard corpus collection according to the standard corpus set corresponding to each corpus segment; Based on the occurrence frequencies of the respective standard corpora in the standard corpus collection, determine at least two target standard corpora; Obtain the representation vector of each target standard corpus based on the standard corpus, and use the representation vector of the target standard corpus as the standard representation vector.
6. The method according to claim 5, wherein, Before obtaining the standard corpus, it further includes a step of constructing an inverted index for the standard corpus, and the inverted index is used to query the standard corpora containing the corpus segment based on the corpus segment; Correspondingly, the querying the standard corpus based on each corpus segment to obtain the standard corpus set corresponding to each corpus segment includes: Query the standard corpus based on the inverted index to obtain the standard corpus set corresponding to each corpus segment.
7. The method according to claim 6, wherein, The step of constructing an inverted index for the standard corpus includes: Perform word segmentation on multiple standard corpora respectively to obtain at least one corpus segment corresponding to each standard corpus; Respectively use each corpus segment as an index keyword; Determine the standard corpora containing each index keyword; Construct a first corpus list based on the standard corpora containing each index keyword; Construct an inverted index of each index keyword and the first corpus list corresponding to each index keyword.
8. The method according to claim 7, wherein, The method further includes adjusting the inverted index: In response to the received negative feedback for category determination, detect whether there is abnormal data in the inverted index; When there is abnormal data in the inverted index, delete the index keyword corresponding to the abnormal data.
9. The method according to claim 5, wherein, Before obtaining the standard corpus, it further includes a step of establishing a mapping relationship for the standard corpus, and the mapping relationship is used to query the category pointed to by the standard corpus based on the standard corpus; Correspondingly, the determining the category of the target object based on the class annotation information corresponding to the matched standard representation vector includes: Determine the standard corpus corresponding to the matched standard characterization vector; Query the standard corpus based on the mapping relationship, determine the category pointed to by the corresponding standard corpus, and use the pointed category as the category of the target object.
10. The method according to claim 9, characterized in that the step of establishing the mapping relationship for the standard corpus includes: Determine a plurality of categories based on the category annotation information carried by a plurality of standard corpora; Determine the standard corpus pointing to each category; Construct a second corpus list based on the standard corpus pointing to each category; Establish a mapping relationship between each category and the second corpus list corresponding to each category.
11. An update device for a corpus processing model, characterized in that the device includes: A sample set acquisition unit configured to acquire the current batch of sample sets; A grouping unit configured to group according to the category annotation information carried by the sample corpora in the current batch of sample sets, so that the sample corpora carrying the same category annotation information are located in the same sample corpus group; A characterization vector obtaining unit configured to obtain the characterization vector of the sample corpus based on the current corpus processing model; A first correlation calculation unit configured to calculate the correlation between the characterization vector of the sample corpus and the characterization vectors of the sample corpora in the same group, to obtain the first correlation of the sample corpus, and the sample corpora in the same group are other sample corpora located in the same sample corpus group as the sample corpus; A second correlation calculation unit configured to calculate the correlation between the characterization vector of the sample corpus and the characterization vectors of the sample corpora in different groups, to obtain the second correlation of the sample corpus, and the sample corpora in different groups are other sample corpora located in different sample corpus groups from the sample corpus; A parameter adjustment unit configured to adjust the parameters of the current corpus processing model to meet the model convergence condition according to the first correlation and the second correlation, and use the current corpus processing model that meets the model convergence condition as the target corpus processing model; wherein, the parameter adjustment unit includes: a current batch actual global correlation obtaining unit configured to obtain the current batch actual global correlation of the sample corpus according to the first correlation and the second correlation; a current batch expected global correlation obtaining unit configured to obtain the current batch expected global correlation of the sample corpus; a loss function value calculation unit configured to calculate the loss function value according to the current batch actual global correlation and the current batch expected global correlation; a parameter adjustment subunit configured to adjust the parameters of the current corpus processing model to meet the model convergence condition based on the loss function value; The current batch actual global correlation obtaining unit includes: a normalization processing unit configured to perform normalization processing on the first correlation and the second correlation to obtain the current batch actual global correlation.
12. The device according to claim 11, characterized in that the sample set acquisition unit includes: A category expansion instruction receiving unit, configured to receive a category expansion instruction, where the category expansion instruction includes a target category and a sample corpus carrying target category annotation information; A sample set construction unit, configured to construct the current batch of sample sets based on the sample corpus carrying target category annotation information.
13. The apparatus according to claim 11, wherein, the characterization vector obtaining unit includes: A first corpus processing unit, configured to perform word segmentation on the sample corpus using the first corpus processing structure of the current corpus processing model to obtain at least two sample corpus segments, perform vector transformation on each sample corpus segment, and obtain a matrix representing the sample corpus based on the vectors corresponding to each sample corpus segment; A second corpus processing unit, configured to perform encoding processing on the matrix using the second corpus processing structure of the current corpus processing model to obtain the characterization vector of the sample corpus.
14. A category determination apparatus, wherein, the apparatus includes: A to-be-processed corpus acquisition unit, configured to acquire a to-be-processed corpus indicating a target object; A model application unit, configured to use the to-be-processed corpus as an input and obtain the characterization vector of the to-be-processed corpus using the target corpus processing model according to any one of claims 1 to 3; A matching unit, configured to determine a standard characterization vector matching the characterization vector of the to-be-processed corpus based on the similarity between the characterization vector of the to-be-processed corpus and multiple standard characterization vectors, and each standard characterization vector carries its corresponding category annotation information; A category determination unit, configured to determine the category of the target object based on the category annotation information corresponding to the matched standard characterization vector.
15. The apparatus according to claim 14, wherein, the apparatus further includes a standard characterization vector determination unit, and the standard characterization vector determination unit includes: A standard corpus acquisition unit, configured to acquire a standard corpus, where the standard corpus records standard corpora and the characterization vectors of the standard corpora, each standard corpus carries its corresponding category annotation information, and the characterization vector of the standard corpus is obtained using the target corpus processing model; A first word segmentation processing unit, configured to perform word segmentation on the to-be-processed corpus to obtain at least two corpus segments; A first query unit, configured to query the standard corpus based on each corpus segment to obtain a standard corpus set corresponding to each corpus segment, and the standard corpora in the standard corpus set corresponding to the corpus segment all contain the corpus segment; A standard corpus collection obtaining unit, configured to obtain a standard corpus collection according to the standard corpus sets corresponding to each corpus segment; A target standard corpus determination unit, configured to determine at least two target standard corpora based on the occurrence frequencies of the standard corpora in the standard corpus collection; The standard representation vector determination subunit is configured to obtain the representation vector of each target standard corpus based on the standard corpus, and use the representation vector of the target standard corpus as the standard representation vector.
16. The apparatus according to claim 15, wherein, the apparatus further includes an inverted index construction unit, and the inverted index constructed by the inverted index unit is used to query the standard corpus containing the corpus segment based on the corpus segment; Correspondingly, the first query unit includes: The first query subunit is configured to query the standard corpus based on the inverted index to obtain the standard corpus set corresponding to each corpus segment.
17. The apparatus according to claim 16, wherein, the inverted index construction unit includes: The second word segmentation processing unit is configured to perform word segmentation processing on multiple standard corpora respectively to obtain at least one corpus segment corresponding to each standard corpus; The index keyword determination unit is configured to perform using each corpus segment as an index keyword respectively; The first standard corpus determination unit is configured to determine the standard corpus containing each index keyword; The first corpus list construction unit is configured to perform constructing a first corpus list based on the standard corpus containing each index keyword; The inverted index construction subunit is configured to perform constructing an inverted index of each index keyword and the first corpus list corresponding to each index keyword.
18. The apparatus according to claim 17, wherein, the apparatus further includes an inverted index adjustment unit, and the inverted index adjustment unit includes: The negative feedback receiving unit is configured to perform detecting whether there is abnormal data in the inverted index in response to the received category determination negative feedback; The deletion unit is configured to perform deleting the index keyword corresponding to the abnormal data when there is abnormal data in the inverted index.
19. The apparatus according to claim 15, wherein, the apparatus further includes a mapping relationship establishment unit, and the mapping relationship established by the mapping relationship establishment unit is used to query the category pointed to by the standard corpus based on the standard corpus; Correspondingly, the category determination unit includes: The second standard corpus determination unit is configured to determine the standard corpus corresponding to the matching standard representation vector; The category determination subunit is configured to perform querying the standard corpus based on the mapping relationship to determine the category pointed to by the corresponding standard corpus, and using the pointed category as the category of the target object.
20. The apparatus according to claim 19, wherein, the mapping relationship establishment unit includes: The standard corpus processing unit is configured to perform determining multiple categories based on the category annotation information carried by multiple standard corpora; The third standard corpus determination unit is configured to determine the standard corpus pointing to each category; The second corpus list construction unit is configured to perform constructing a second corpus list based on the standard corpus pointing to each category; A mapping relationship establishing subunit, configured to establish a mapping relationship between each of the categories and a second corpus list corresponding to each of the categories.
21. An electronic device, characterized in that it includes: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the corpus processing model updating method according to any one of claims 1 to 3 or the category determination method according to any one of claims 4 to 10.
22. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the corpus processing model updating method according to any one of claims 1 to 3 or the category determination method according to any one of claims 4 to 10.
Citation Information
Patent Citations
A method of generating a word vector from a multi-task model
CN109325231A
Corpus classification method and device, computer equipment and storage medium
CN109902285A