Text processing method, device, equipment and storage medium
Through the optimized hierarchical classification model and the matching processing of the preset word vector set, the category information of the text is accurately determined, which solves the problem of inaccurate category information in traditional methods and improves the accuracy and efficiency of search results.
Patent Information
- Application Number
- CN202110978569.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-08-24
AI Technical Summary
Traditional text processing methods cannot accurately determine the category information of text content, resulting in inaccuracy and inefficiency of search results.
The optimized hierarchical classification model is adopted to obtain the semantic vectors of the text to be predicted, and match the preset word vector set for matching, determine the target secondary category of the text to be predicted, and generate category information based on the pre-established correspondence relationship.
It improves the accuracy of category information and the accuracy of search results, and improves the efficiency of data retrieval.
Smart Images

Figure CN114328807B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a text processing method, apparatus, device, and storage medium. Background Art
[0002] Category information within text content (such as primary and secondary categories) plays a crucial role in information retrieval. For applications that need to provide data search services, accurately extracting category information that matches the search text can improve the accuracy of the application's search results, thereby increasing the application's search efficiency. However, traditional text processing methods are unable to accurately determine the category information within text content. Therefore, accurately extracting this information has become a research hotspot. Summary of the Invention
[0003] The embodiments of the present application provide a text processing method, apparatus, device, and storage medium, which can improve the accuracy of predicted category information.
[0004] In one aspect, an embodiment of the present application provides a text processing method, comprising:
[0005] Obtaining a first semantic vector of the text to be predicted, and using the optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector, to obtain one or more first-level predicted categories;
[0006] Matching the first semantic vector with a word vector of each secondary category in a preset word vector set to obtain one or more first matching features, where the preset word vector set includes: one or more secondary categories and a word vector of each secondary category in the one or more secondary categories;
[0007] Using the optimized hierarchical classification model, based on the one or more first matching features and the one or more first-level prediction categories, determine a target second-level category of the text to be predicted from the one or more second-level categories;
[0008] Obtaining a target first-level category corresponding to the target second-level category based on a pre-established correspondence between the second-level category and the first-level category;
[0009] Category information of the text to be predicted is generated, where the category information includes the target primary category and the target secondary category.
[0010] In one aspect, an embodiment of the present application provides a text processing device, comprising:
[0011] a first prediction unit, configured to obtain a first semantic vector of a text to be predicted, and use an optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector, to obtain one or more first-level predicted categories;
[0012] a matching processing unit, configured to match the first semantic vector with a word vector of each secondary category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes: one or more secondary categories and a word vector of each secondary category in the one or more secondary categories;
[0013] a second prediction unit, configured to determine a target secondary category of the text to be predicted from the one or more secondary categories based on the one or more first matching features and the one or more first-level prediction categories, using the optimized hierarchical classification model;
[0014] a processing unit, configured to obtain a target primary category corresponding to the target secondary category based on a pre-established correspondence between the secondary categories and the primary categories;
[0015] A generating unit is configured to generate category information of the text to be predicted, wherein the category information includes the target primary category and the target secondary category.
[0016] In one aspect, the present application provides a text processing device, comprising:
[0017] a processor adapted to execute one or more computer programs;
[0018] A computer storage medium storing one or more computer programs, wherein the one or more computer programs are suitable for being loaded and executed by the processor:
[0019] Obtain a first semantic vector of a text to be predicted, and use an optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector to obtain one or more first-level predicted categories; match the first semantic vector with a word vector of each second-level category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes: one or more second-level categories, and a word vector of each second-level category in the one or more second-level categories; use the optimized hierarchical classification model to determine a target second-level category of the text to be predicted from the one or more second-level categories based on the one or more first matching features and the one or more first-level predicted categories; obtain a target first-level category corresponding to the target second-level category based on a pre-established correspondence between the second-level category and the first-level category; generate category information of the text to be predicted, wherein the category information includes the target first-level category and the target second-level category.
[0020] In one aspect, the present application provides a storage medium storing one or more computer programs, wherein the one or more computer programs are suitable for being loaded and executed by a processor:
[0021] Obtain a first semantic vector of a text to be predicted, and use an optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector to obtain one or more first-level predicted categories; match the first semantic vector with a word vector of each second-level category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes: one or more second-level categories, and a word vector of each second-level category in the one or more second-level categories; use the optimized hierarchical classification model to determine a target second-level category of the text to be predicted from the one or more second-level categories based on the one or more first matching features and the one or more first-level predicted categories; obtain a target first-level category corresponding to the target second-level category based on a pre-established correspondence between the second-level category and the first-level category; generate category information of the text to be predicted, wherein the category information includes the target first-level category and the target second-level category.
[0022] In one aspect, the present application provides a computer program product or computer program, the computer program product comprising a computer program, the computer program being stored in a computer storage medium; a processor of a text processing device reading the computer program from the computer storage medium and executing the computer program, causing the text processing device to perform:
[0023] Obtain a first semantic vector of a text to be predicted, and use an optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector to obtain one or more first-level predicted categories; match the first semantic vector with a word vector of each second-level category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes: one or more second-level categories, and a word vector of each second-level category in the one or more second-level categories; use the optimized hierarchical classification model to determine a target second-level category of the text to be predicted from the one or more second-level categories based on the one or more first matching features and the one or more first-level predicted categories; obtain a target first-level category corresponding to the target second-level category based on a pre-established correspondence between the second-level category and the first-level category; generate category information of the text to be predicted, wherein the category information includes the target first-level category and the target second-level category.
[0024] In an embodiment of the present application, the text processing device constructs a first matching feature between the word vector of the secondary category and the semantic vector of the text to be predicted based on a variety of different matching features, so that the accuracy of the target secondary category predicted by the hierarchical classification model is higher; in addition, since the text processing device is based on the word vectors of each secondary category, the first matching feature determined can enable the text processing device to better utilize the semantic information of the secondary category, thereby further improving the accuracy of the category information. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0026] Figure 1 is a schematic diagram of a training sample provided in an embodiment of the present application;
[0027] Figure 2 This is a schematic diagram of an optimization method for a hierarchical classification model provided in an embodiment of the present application;
[0028] Figure 3 This is a schematic diagram of the structure of a hierarchical classification model provided in an embodiment of the present application;
[0029] Figure 4 This is a flowchart of a text processing method provided by an embodiment of the present application;
[0030] Figure 5 This is a schematic diagram of the architecture of a text processing device provided in an embodiment of the present application;
[0031] Figure 6 This is a schematic diagram of the architecture of a text processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] This application proposes a text processing method based on natural language processing technology. The text processing method can generate category information of the text to be predicted, so that each business party applying the text processing method can classify and process the data content corresponding to the text to be predicted based on the category information. Among them, the text to be predicted can include but is not limited to: title text (such as: video title, music title, document title, etc.), search text (such as: keywords in the search bar), then the data content corresponding to the text to be predicted can include but is not limited to: video data, music data, document data, etc.; In addition, the category information mentioned above includes but is not limited to: first-level category, second-level category, third-level category, etc. It should be noted that there is a hierarchical relationship between the category information at each level, and the upper-level category is the parent of the lower-level category. The lower the level, the finer the category granularity; for example, the first-level category is the parent of the second-level category, and the granularity of the second-level category is finer than the first-level category; the second-level category is the parent of the third-level category, and the granularity of the third-level category is finer than the second-level category.
[0033] Among them, the categories mentioned above can be understood as: categories, then the first-level categories can be understood as: categories with broad meanings, which cover more information. For example, the first-level categories can include: games, digital products, home appliances, daily necessities, clothing, etc.; there is a hierarchical relationship between the second-level categories and the first-level categories. Specifically, the second-level categories are subcategories of the first-level categories. Based on the above description, it is not difficult to understand that the division of the second-level categories is more refined than that of the first-level categories. Then, the information covered by each second-level category can be the information covered by the corresponding first-level category of the second-level category. For example, for the first-level category "Games", its subcategories (i.e., second-level categories) can include: "Mini-Program Games", "Stand-alone Games", "Mobile Games", "PC Games", etc.; further, assuming that the information included in "Mobile Games" is: Mobile Game A and Mobile Game B, and the information included in "PC Games" is: PC Game A and PC Game B, then the information included in the "Games" category can be: Mobile Game A, Mobile Game B, as well as PC Game A and PC Game B. It is not difficult to see that the information included in the "Mobile Games" category is part of the information included in the "Games" category. For an understanding of the third-level categories, please refer to the "Second-level Categories", which will not be repeated in this application.
[0034] It should be noted that the above-mentioned natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. Natural language processing mainly studies various theories and methods that can achieve effective communication between people and computers using natural language (that is, the language people use in daily life). In general, natural language processing is a science that integrates linguistics, computer science, and mathematics. It is mainly used in machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, Chinese OCR (Optical Character Recognition), etc. This application mainly uses the relevant technologies of text classification and text semantic comparison in natural language processing.
[0035] This application makes full use of the natural language processing technology mentioned above and proposes a text processing solution, which can be executed by a text processing device. The general principle of the solution is explained below: the text processing device adopts an optimized hierarchical classification model to determine the first-level prediction category of the text to be predicted based on the semantic vector of the text to be predicted, and adopts the optimized hierarchical classification model to construct the multi-matching features of the word vectors of each second-level category in the preset word vector set and the semantic vector, so that the text processing device can determine the target second-level category of the text to be predicted based on the constructed multiple matching features. Then, it further enables the text processing device to backtrack and obtain the target first-level category of the text to be predicted based on the target second-level category, thereby generating category information of the text to be predicted. It can be understood that the category information includes the determined target second-level category and target first-level category.
[0036] Among them, the preset word vector set mentioned above is pre-stored in the text processing device. Optionally, the preset word vector set can also include: one or more secondary categories, each of the one or more secondary categories has a corresponding relationship with its parent category (i.e., the primary category) to support the text processing device to trace back to the target primary category based on the target secondary category, thereby supporting the text processing device to generate complete category information. For example, the secondary category "mobile games" can establish a corresponding relationship with the primary category "games", and the secondary category "client games" can also establish a corresponding relationship with the primary category "games". Then, when the target secondary category determined by the text processing device is "client games" or "mobile games", it can be traced back to obtain the target primary category "games" based on the corresponding relationship between "client games" and "games", or based on the corresponding relationship between "mobile games" and "games". Then, based on the above description, that is to say, the preset word vector set can include: one or more secondary categories, and the word vectors of each secondary category in one or more secondary categories. It can be understood that these one or more primary categories (or one or more secondary categories) refer to: categories that can be output by the text processing device, or understood as: categories that can be predicted by the text processing device. For example: assuming that the natural language processing technology includes a total of 3 primary categories (primary category A, primary category B, primary category C), and the text processing device only stores primary category A and primary category B, then for any text to be predicted, after the text processing device performs text processing on the text to be predicted, the target primary category determined can only be one of the primary category A and the primary category B, and it is impossible to determine the primary category C. Generally speaking, the categories stored in the text processing device are all categories included in the natural language processing technology. Of course, when this application is applied to category prediction in a specific field, the text processing device can also only store categories related to the specific field. This processing method can improve the processing efficiency of the text processing device to a certain extent.
[0037] In actual applications, the text processing device performs the above-mentioned text processing on the title text, which can be considered as the completion of the parsing of the title text. Then, based on the parsing results (i.e., the generated category information), the text processing device can strengthen its understanding of the semantic information of the multimedia data (such as pictures, videos, text, voice, etc.) corresponding to the title text, which can effectively improve the efficiency and accuracy of semantic understanding to a certain extent. For example, in a video playback application, determining the basic subject category of the video corresponding to the video title based on the video title, and modeling and completing the title word weight task are both basic capabilities that the video playback application needs to have in order to understand the video content. Then, the video playback application can apply the above-mentioned text processing method to perform text processing on the video title, and use the generated category information as the video semantic information of the video corresponding to the video title. This method can improve the video understanding efficiency of the video playback application to a certain extent. In addition, since the category information includes the category information of the video title of the video, the video playback application can accurately and quickly classify the video based on the target first-level category and target second-level category in the category information, so that when the video playback application detects a retrieval for the target first-level category and / or target second-level category, it can quickly find the video from the database corresponding to the video playback application, which can effectively improve the efficiency and accuracy of video retrieval.
[0038] In one embodiment, the above-described text processing method can be implemented within a single text processing device, which can be either a server or a terminal. For example, in the case of a server, the server can first obtain a semantic vector for the text to be predicted, and then, using an optimized hierarchical classification model, perform text processing on the text to be predicted based on the semantic vector to obtain category information corresponding to the predicted text.
[0039] In another embodiment, the above-mentioned text processing method can also be applied in different text processing devices. For example, the semantic vector of the text to be predicted can be acquired in a first text processing device, and in a second text processing device, an optimized hierarchical classification model is used to perform hierarchical classification processing on the text to be predicted based on the semantic vector to obtain category information corresponding to the predicted text, wherein the first text processing device can be a terminal device, and the second text processing device can be a server that establishes a communication connection with the first processing device.
[0040] It should be noted that the servers mentioned above may include, but are not limited to, independent physical servers, server clusters or distributed systems consisting of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal devices mentioned above may include, but are not limited to, smartphones, computers (such as laptops, tablets, and desktops), smart watches, smart home appliances, and in-vehicle terminals.
[0041] In actual application, before the text processing device applies the above-mentioned text processing method, the present application can also use a deep learning algorithm to optimize the hierarchical classification model based on the training samples through the model optimization device to obtain the above-mentioned optimized hierarchical classification model. Among them, the model optimization device and the text processing device can be the same device or different devices, and the present application does not impose specific restrictions on this; the training samples used for model optimization can include N (N is a positive integer) training texts, and the standard first-level categories and standard second-level categories corresponding to each training text in the N training texts. That is to say, the training samples can include: N training texts, N standard first-level categories, and N standard second-level categories. For example, the training text can be a title text; wherein, the standard first-level category can be understood as: the correct first-level category corresponding to the title text, and similarly, the standard second-level category can be understood as: the correct second-level category corresponding to the title text. For example, if Figure 1 As shown, assuming that the first-level category "Game" corresponds to the second-level categories "Mobile Games", "Mini Games" and "PC Games", then if Figure 1 The title text of the mobile game video 100 in the example is used as training text. The standard first-level category of the training text is "Games"; the standard second-level category of the training text is "Mobile Games." Therefore, it can be understood that when a text processing device processes the title text, it is expected that the text processing device will generate category information including the target first-level category of "Games" and the target second-level category of "Mobile Games."
[0042] In one embodiment, the aforementioned model optimization device (or devices) and text processing device (or devices) may form a blockchain, with any of the model optimization device and text processing device acting as a node on the blockchain. A blockchain is a distributed, shared ledger and database characterized by decentralization, immutability, full traceability, traceability, collective maintenance, and transparency. Therefore, any node on the blockchain comprised of the model optimization device and text processing device can access data stored on the blockchain (e.g., training samples, category information, etc.).
[0043] For ease of explanation, this application uses the text processing device and the model optimization device as the same device as an example. In addition, unless otherwise specified, the following uses the category information including the first-level category and the second-level category as an example to describe in detail the text processing method provided by this application. In this case, for example, the first-level category can be 44 thematic coarse-grained categories, which can specifically include: sports, games, entertainment, etc.; in addition, each of the 44 first-level categories is further subdivided into fine-grained categories, namely: one or more second-level categories. Specifically, the number of one or more second-level categories can be 305, such as: gymnastics, table tennis, basketball, mobile games, PC games, singing, reading, etc.
[0044] See Figure 2 , Figure 2 This is a schematic diagram of an optimization method for a hierarchical classification model provided in an embodiment of the present application. The method is executed by the above-mentioned text processing device, such as Figure 2 As shown, the method includes steps S201-S206.
[0045] S201, obtaining training samples.
[0046] Among them, the training samples can be composed of first-level categories, second-level categories, and training texts. In the optimization process of the hierarchical classification model, the first-level categories in the training samples can be called standard first-level categories, and the standard first-level categories indicate: the first-level categories that are expected to be used as target first-level categories by the hierarchical classification model. Then, correspondingly, the second-level categories in the training samples can be called standard second-level categories, and the standard second-level categories indicate: the second-level categories that are expected to be used as target second-level categories by the hierarchical classification model. It should be noted that the standard second-level categories are subordinate categories of the standard first-level categories, that is, the granularity of the standard second-level categories is finer than that of the standard first-level categories, wherein the so-called granularity refers to: the level of refinement or comprehensiveness of the data stored in the data unit of the database; specifically, the higher the degree of refinement of the data, the smaller the granularity level of the data, and correspondingly, the lower the degree of refinement of the data, the larger the granularity level of the data.
[0047] For example, the training samples used in this application can be found in a training sample table, which can be shown in Table 1.
[0048] Table 1
[0049] Training text Standard first-level category Standard secondary category Text A game Mini Games Text B dance Square Dance Text C science and technology cell phone ... ... ... Text N First-level category M Secondary category m
[0050] Based on Table 1, it is not difficult to see that there can be N training texts used to optimize the hierarchical classification model, where N is a positive integer. However, in practical applications, in order to obtain a hierarchical classification model with more stable performance and higher accuracy, N can usually be a larger positive integer (such as 5000, 10000, etc.).
[0051] It should be noted that in order to better understand the optimization method of the hierarchical classification model proposed in this application, the following steps S202-S205 are all taken as an example of N=1, that is, the training sample includes: a training text, and the standard first-level category and standard second-level category corresponding to the training text as an example, to elaborate on the model optimization method of the hierarchical classification model proposed in this application.
[0052] S202: Decode the training text to obtain a second semantic vector of the training text.
[0053] In practical applications, the text processing device may use any one of the following decoding methods to decode the training text, thereby obtaining the second semantic vector of the training text.
[0054] Optionally, the first decoding method can be: the text processing device adopts LSTM (Long Short-Term Memory, long short-term memory network) + Attention (attention mechanism) technology to realize the decoding processing of the training text, thereby obtaining the second semantic vector of the training text. Then, in this case, the text processing device can first vectorize the training text, and then use the LSTM network to extract sequence information of words (i.e., words in the training text) and sentences (i.e., sentences corresponding to the training text). Furthermore, the text processing device can add an attention mechanism layer after the LSTM network at the word level and the sentence level to obtain the feature vector of the training text, i.e., the second semantic vector. It is not difficult to understand that in this way, the text processing device can effectively extract the semantic structure feature information between words and words, and between sentences in the training text. That is to say, the text processing device adopts the decoding method proposed in this application to obtain a more accurate second semantic vector.
[0055] Optionally, a second decoding method may be: the text processing device uses a CNN (Convolutional Neural Network) to decode the training text to obtain a second semantic vector of the training text. The so-called convolutional neural network is a type of feedforward neural network (Feedforward Neural Network) that includes convolution calculations and has a deep structure. It is one of the representative algorithms of deep learning. Convolutional neural networks have representation learning capabilities and can perform translation-invariant classification of input information according to their hierarchical structure.
[0056] Optionally, a third decoding method may be: the text processing device uses an LSTM network to decode the training text to obtain a second semantic vector of the training text. The so-called LSTM network can be understood as: a time-recursive neural network; this LSTM network is suitable for processing and predicting important events in time series with relatively long intervals and delays. Based on this, it can be understood that if the present application constructs a feature extraction module based on LSTM, the second semantic vector of the training text can be accurately extracted.
[0057] It should be noted that, in other embodiments, the text processing device may also adopt other decoding methods to decode the training text and then extract the second semantic vector of the training text. This application does not list other methods one by one here.
[0058] S203 : Using a hierarchical classification model, determine a second predicted probability of the standard first-level category based on the second semantic vector.
[0059] The second predicted probability of the standard first-level category can be used to indicate the probability that the text processing device uses the standard first-level category as the target first-level category. In practical applications, the text processing device can use a hierarchical classification model to predict the first-level category of the training text. For example, the structure of the hierarchical classification model can be seen in Figure 3 .like Figure 3As shown, the hierarchical classification model includes a first-level category prediction structure 31 and a second-level category prediction structure 32. It can be seen that the first-level category prediction structure 31 used in this application can be a conventional multi-classification structure. Then, when the text processing device performs first-level category prediction on the training text through the first-level category prediction structure 31, it can specifically include the following steps: the text processing device first inputs the training text into the encoder in the first-level category prediction structure 31, and then obtains the semantic vector corresponding to the training text (such as: the second semantic vector). Furthermore, the text processing device can input the semantic vector into the classifier for multi-classification processing, thereby obtaining the prediction probability of each first-level category, and further obtaining the second prediction probability of the standard first-level category.
[0060] Specifically, if 44 first-level categories are stored in the hierarchical classification model, then when the text processing device uses the hierarchical classification model to predict the first-level categories of the training text, it can be achieved in the following way: the text processing device performs first-level category prediction processing on the training text based on the second semantic vector of the training text by calling the first-level category prediction structure. During the prediction process, the classifier in the first-level category prediction structure 31 will determine the prediction probability of each of the 44 first-level categories mentioned above based on the second semantic vector. It is not difficult to understand that these 44 first-level categories include standard first-level categories, so that is to say, the prediction probability of each first-level category in the 44 first-level categories includes: the second prediction probability of the standard first-level category. Then, combined with the above description of the second prediction probability of the standard first-level category, it is obvious that the prediction probability of each first-level category can be used to indicate: the probability that the text processing device uses the first-level category as the target first-level category.
[0061] Optionally, the text processing device may further determine one or more first-level predicted categories based on the predicted probabilities of each first-level category, and then perform vectorization processing on the one or more first-level predicted categories to obtain one or more first-level predicted category vectors. These one or more first-level predicted category vectors may be used to assist the second-level category prediction module 32 in determining a target second-level category.
[0062] In one embodiment, the one or more first-level predicted categories obtained may be some of the 44 first-level categories mentioned above, for example, after sorting the predicted probabilities from high to low, the first-level categories corresponding to the predicted probabilities of the top three predicted probabilities are arranged. Of course, in other embodiments, the one or more first-level predicted categories obtained may also be all of the 44 first-level categories mentioned above, and this application does not impose specific limitations on this.
[0063] S204 : Match the second semantic vector with the word vector of each secondary category in one or more secondary categories to obtain one or more second matching features.
[0064] Among them, the word vector of each secondary category can be obtained by the text processing device based on general corpus training. Corpus refers to: a collection of text resources of a certain number and scale; then, general corpus can be understood as: text resources that can be understood and used in various fields. The second matching feature refers to: the matching feature between the second semantic vector and the word vector of the secondary category. Then, based on this, it is not difficult to understand that: the number of one or more second matching features is the same as the number of one or more secondary categories. That is to say, for any secondary category in one or more secondary categories, the text processing device can match the word vector of any secondary category with the second semantic vector, and then obtain a second matching feature.
[0065] In a specific embodiment, the general corpus can be obtained in advance by technical personnel from various businesses, such as: obtaining secondary categories related to games from the game business, obtaining secondary categories related to music from the music playback business, and so on. It can be understood that the general corpus in this application mainly refers to secondary categories that are applicable to various fields. Then, in order to facilitate the subsequent secondary category prediction processing of the training text, this application can convert each secondary category into a word vector by means of Word2vector (converting words into vectors), and generate a preset word vector set based on the converted word vector. Then, after storing the preset word vector set in the hierarchical classification model, the text processing device can obtain one or more second matching features by matching the word vectors of each secondary category in the preset word vector set with the second semantic vector.
[0066] Among them, the text processing device can construct the second matching feature based on a multi-matching mechanism. The so-called multi-matching mechanism can be understood as: the text processing device determines the second matching feature between the word vector of the secondary category and the second semantic vector based on a variety of matching processing methods. Exemplarily, the multiple matching processing methods mentioned above may include but are not limited to: vector matching processing, interactive matching processing, semantic matching processing, etc., among which, the vector matching processing can be implemented by the matching module 1 in the secondary category prediction module, the interactive matching processing can be implemented by the matching module 2 in the secondary category prediction module, and the semantic matching processing can be implemented by the matching module 3 in the secondary category prediction module.
[0067] It should be noted that the relevant principles of the text processing device constructing the second matching feature based on the multi-matching mechanism can be found in the relevant description of the subsequent step S403, and this application will not repeat them here.
[0068] S205 , using a hierarchical classification model, determining a third predicted probability of the standard secondary category based on the one or more second matching features.
[0069] In actual applications, the text processing device can determine the predicted probability of the secondary category corresponding to the second matching feature based on the second matching feature, wherein the predicted probability of the secondary category is used to indicate: the probability that the text processing device uses the secondary category as the target secondary category; for example, assuming that one or more second matching features include a second matching feature A, and the second matching feature A is determined based on the word vector b of the secondary category B and the second semantic vector of the training text; then, the text processing device can determine the predicted probability of the secondary category B based on the second matching feature A. Based on this, it can be understood that the text processing device can determine the predicted probability of each secondary category in one or more secondary categories based on one or more second matching features; and since the standard secondary category is also included in one or more first-level categories stored in the hierarchical classification model, the text processing device can determine the third predicted probability of the standard secondary category based on the predicted probability of each secondary category in one or more secondary categories.
[0070] S206 , determining a target loss value based on the second predicted probability and the third predicted probability, and optimizing the model parameters of the hierarchical classification model in a direction of reducing the target loss value to obtain an optimized hierarchical classification model.
[0071] In one embodiment, when N=1, the number of second predicted probabilities is one, and the number of third predicted probabilities is also one. In this case, the text processing device can determine the target loss value in the following manner: the text processing device determines the first loss value based on the second predicted probability and the third predicted probability, determines the second loss value based on the second predicted probability, and determines the third loss value based on the third predicted probability; further, the text processing device can reconcile the first loss value, the second loss value, and the third loss value to obtain the target loss function. Exemplarily, the reconciliation process can specifically refer to: weighted summation process. That is to say, the text processing device can obtain the target loss value by performing a weighted summation process on the first loss value, the second loss value, and the third loss value.
[0072] In a specific application, the first loss value can be calculated based on the first loss function. If the predicted probability of the secondary category is greater than the probability of the secondary category corresponding to the primary category, the loss of the hierarchical classification model will increase. Then, in order to reduce the loss value of the model, it is necessary to avoid the situation where the predicted probability of the secondary category is greater than the predicted probability of the primary category corresponding to the secondary category. Then, it can be understood that in this application, in order to ensure the consistency of the processing results (i.e., the prediction results of the primary category and the prediction results of the secondary category), it can be assumed that the prediction of the coarse-grained primary category is always easier than the prediction of the fine-grained secondary category, that is, to ensure that the prediction probability of the secondary category is less than the prediction probability of the secondary category corresponding to the primary category. Then, exemplarily, the first loss function can be a hinge loss function, which can be used to represent the difference between the primary classification result and the secondary classification result. Then, it can be understood that the purpose of citing the hinge loss function is to hope that the predicted probability corresponding to the secondary category predicted by the hierarchical classification model is greater than the predicted probability of the secondary category corresponding to the primary category. Wherein, exemplarily, the expression of the hinge loss function can refer to Formula 1:
[0073] loss h =max(0,λ+l2_score-l1 score ) Formula 1
[0074] Among them, loss h represents the first loss value, λ is a hyperparameter, l2_score represents the third prediction probability (i.e., the prediction probability of the standard secondary category), l1 score Represents the second predicted probability (i.e., the predicted probability of the standard first-level category).
[0075] In a specific application, the second loss value can be calculated based on the second loss function, and the second loss value can be used to represent the loss value of the primary category prediction module. For example, the second loss function can be a negative logarithmic loss function, and the expression of the negative logarithmic loss function can be seen in Formula 2:
[0076]
[0077] Among them, Loss cls1 Represents the second loss value, n represents the number of first-level categories stored in the hierarchical classification model, n represents the number of first-level categories stored in the hierarchical classification model, n is a positive integer, for example, n can be 44; i represents the i-th first-level category, i is a positive integer, y i Represents the true value of the i-th level category, a iRepresents the predicted probability of the i-th first-level category. For example, the true value corresponding to the standard first-level category can be 1, and the true values of other first-level categories except the standard first-level category can all be 0. Then, it can be understood that when the i-th first-level category is the standard first-level category, y i =1,a i is the second predicted probability mentioned above.
[0078] In a specific application, the third loss value can be determined based on a third loss function, and the third loss value can be used to indicate the loss value of the secondary category prediction module. For example, the third loss function can be a negative logarithmic loss function as shown in Formula 3:
[0079]
[0080] Among them, Loss cls2 represents the third loss value, m represents the number of secondary categories stored in the hierarchical classification model, m is a positive integer, for example, m can be 305; j represents the jth secondary category, j is a positive integer, y j represents the true value of the j-th secondary category, a j Represents the predicted probability of the jth secondary category. For example, the true value corresponding to the standard secondary category can be 1, and among the m secondary categories, the true values of the other secondary categories except the standard secondary category can all be 0. Then, it can be understood that when the jth secondary category is the standard secondary category, y j =1,a j is the third predicted probability mentioned above.
[0081] Then, based on the above description, the text processing device can reconcile the first loss value, the second loss value and the third loss value in a manner as shown in Formula 4 to obtain a target loss value.
[0082] Loss=λ1Loss cls1 +λ2Loss cls2 +λ3loss h Formula 4
[0083] Among them, Loss represents the target loss value, Loss cls1 Represents the second loss value, λ1 is the weight of the second loss value, Loss cls2 represents the third loss value, λ2 is the weight of the third loss value, loss hrepresents the first loss value, λ3 is the weight of the first loss value, and λ1+λ2+λ3=1. It should be noted that λ1, λ2, and λ3 can be set according to actual needs. For example, the weight of the loss value with a higher reference value is set to a larger value, and the weight of the loss value with a smaller reference value is set to a smaller value, as long as λ1+λ2+λ3=1 is satisfied.
[0084] In another embodiment, when N is greater than 1, it is not difficult to understand that the second prediction probability determined by the text processing device may include: N second prediction probabilities; the N second prediction probabilities include the second prediction probability of the standard first-level category of each training text in the N training texts; correspondingly, it is understood that the third prediction probability determined by the text processing device may include: N third prediction probabilities; the N third prediction probabilities include the third prediction probability of the standard second-level category of each training text in the N training texts. The second prediction probability of the standard first-level category of each training text is determined by the text processing device based on the second semantic vector of each training text; and the third prediction probability of the standard second-level category of each training text is determined by the text processing device based on one or more second matching features corresponding to each training text.
[0085] It is not difficult to understand from the above embodiment that the N second prediction probabilities correspond one-to-one to the N third prediction probabilities, that is, one second prediction probability corresponds to one third prediction probability, and these two corresponding prediction probabilities are prediction probabilities obtained for the same training text; for example, if the standard first-level category of training text A is a and the standard second-level category is b, then it can be understood that the second prediction probability corresponding to a and the third prediction probability corresponding to b have a corresponding relationship. Then, in this case, the text processing device can determine the target loss value based on the second prediction probability and the third prediction probability in the following manner: the text processing device obtains a first loss value based on the N second prediction probabilities and the N third prediction probabilities, and determines the second loss value based on the second prediction probability of the standard first-level category of any training text in the N training texts; and determines the third loss value based on the third prediction probability of the standard second-level category of any training text; then, further, the text processing device can reconcile the first loss value, the second loss value and the third loss value to finally obtain the target loss value.
[0086] Specifically, the text processing device can calculate the first loss values corresponding to the N second prediction probabilities and the N third prediction probabilities using the loss function shown in Formula 5.
[0087]
[0088] Where N is the number of training texts; k represents the kth training text, k is a positive integer; l2_scorek represents the third predicted probability of the standard secondary category of the k-th training text; represents the second predicted probability of the standard first-level category of the k-th training text, and λ is a hyperparameter.
[0089] In addition, when the text processing device determines the second loss value based on the second predicted probability of the standard first-level category of any training text among the N training texts, and determines the third loss value based on the third predicted probability of the standard second-level category of any training text, the text processing device can specifically determine the following method: the text processing device determines the first true value of the standard first-level category of any training text based on the standard first-level category of the training text; for example, the true value of the standard first-level category can be 1. Further, the text processing device can determine the second loss value based on the first true value and the second predicted probability of the standard first-level category of the training text. The loss function used by the text processing device when determining the second loss value can be specifically described in the relevant description of Formula 2 above, which will not be repeated in this application. Based on the relevant description of Formula 2 above, it can be seen that the second loss value is negatively correlated with the second predicted probability of the standard first-level category of the training text. Similarly, the text processing device can determine the second true value of the standard second-level category of any training text based on the standard second-level category of the training text; and further determine the third loss value based on the second true value and the third predicted probability of the standard second-level category of the training text. When the text processing device determines the third loss value, the loss function used can be found in the relevant description of Formula 2 above. This application will not go into details here. Based on the relevant description of Formula 2 above, it can be seen that the third loss value is negatively correlated with the third predicted probability of the standard secondary category of any training text.
[0090] Based on the above description, it is not difficult to understand that the text processing device can determine N target loss values based on N training texts. In other words, the text processing device can perform N model optimization processes on the hierarchical classification model, or in other words, the text processing device can perform N parameter adjustments on the hierarchical classification model, and the target loss value on which each parameter adjustment depends is determined based on different training texts. It should be noted that, since the second prediction probability of the standard first-level category and the third prediction probability of the standard second-level category corresponding to different training texts may be different, the second loss value and the third loss value used by the text processing device when determining each target loss value may be different; at the same time, since the first loss value is determined based on N second prediction probabilities and N third prediction probabilities, the first loss value used by the text processing device when determining each target loss value is the same.
[0091] In the present application, when the text processing device optimizes the hierarchical classification model, it does so by referring to the weighted sum of the first loss value, the second loss value, and the third loss value. Then, since the first loss value is determined based on the third probability value of the standard secondary category of the training text and the second probability value of the standard primary category of the training text, and there is a hierarchical relationship between the standard secondary category and the standard primary category, it can be seen that the text processing device optimizes the hierarchical classification model using the model optimization method provided by the present application, so that the hierarchical classification model can make full use of the hierarchical relationship between the upper category (such as the first category) and the lower category (such as the second category) when performing text processing, so that the consistency of the first category and the second category predicted by the optimized hierarchical classification model can be effectively improved, thereby improving the accuracy of the optimized hierarchical classification model when performing text processing.
[0092] Based on the above description, after the text processing device completes the model optimization process and obtains the optimized hierarchical classification model, the optimized hierarchical classification model can be used to perform text processing (such as hierarchical classification processing) on the text to be predicted. Figure 4 , Figure 4 This is a text processing method provided by an embodiment of the present application. The text processing method can be executed by the above-mentioned text processing device, such as Figure 4 As shown, the method includes:
[0093] S401: Obtain a first semantic vector of the text to be predicted.
[0094] In one embodiment, the specific implementation of step S401 can refer to the relevant description of step S201, and this application will not elaborate on it here.
[0095] S402: Match the first semantic vector with the word vector of each secondary category in the preset word vector set to obtain one or more first matching features.
[0096] In a specific application, each of the one or more first matching features is obtained by fusing multiple different matching features. In other words, the text processing device can construct the first matching features through a multi-matching mechanism. For example, the multiple different matching features mentioned above may include semantic difference features, vector similarity features, and interaction features. Semantic difference features can be used to indicate the degree of difference between the semantics of the text to be predicted and the semantics of the secondary category; vector similarity features can be used to indicate the degree of similarity between the text to be predicted and the secondary category; and interaction features can be used to indicate the correlation between the text to be predicted and the secondary category. Specifically, interaction features can refer to features obtained by interactively matching the first semantic vector with the word vector of the secondary category. Interactive matching can be understood as follows: for two text contents (e.g., sentences, words, etc.), when the text processing device extracts the features of the first character of text content A, it can compare the word vector corresponding to the first character with the word vector of each word in text content B, thereby extracting the correlation between the first character and text content B, and then allowing the text processing device to determine the interaction features based on the correlation.
[0097] It can be seen from the above that the preset word vector set includes: one or more secondary categories, and the word vector of each secondary category in the one or more secondary categories. Based on this, the text processing device can obtain one or more first matching features in the following manner: the text processing device first traverses one or more secondary categories in the preset word vector set, and performs vector matching processing on the word vector of the currently traversed secondary category and the first semantic vector to obtain vector similarity features; then, the text processing device can interactively match the word vector of the currently traversed secondary category with the first semantic vector to obtain interactive features; and, the text processing device can semantically match the word vector of the currently traversed secondary category with the first semantic vector to obtain semantic difference features; then, after the text processing device obtains vector similarity features, interactive features and semantic difference features, the text processing device can further perform feature fusion processing on the vector similarity features, interactive features and semantic difference features to obtain the first matching features corresponding to the currently traversed secondary category; illustratively, the text processing device can realize feature fusion processing on the vector similarity features, interactive features and semantic difference features by performing dot multiplication or addition on the vector similarity features, interactive features and semantic difference features. Then, it can be understood that in this way, after the text processing device finishes traversing the preset word vector set, it can obtain one or more first matching features mentioned above.
[0098] In order to facilitate a better understanding of the text processing method provided by the present application, the following describes in detail the method in which the text processing device in the present application determines any first matching feature among one or more first matching features, in combination with the specific calculation method of each matching feature.
[0099] In a specific implementation, the vector similarity feature may be cosine similarity. Then, the text processing device may calculate the vector similarity feature using the calculation method shown in Formula 6:
[0100]
[0101] Among them, Logits1 represents cosine similarity, L1 emb Represents the first semantic vector, which belongs to the vector space R m*1 (m and 1 are used to represent the dimension of the vector space), f l2_cls_emb Represents a word vector of a secondary category, which belongs to the vector space R l2 _num*d (l2_num represents the total number of secondary categories, d represents the dimension of the word vector for each secondary category, usually ranging from 200 to 500); It can be understood as a bilinear decomposition method; U and V are both parameter mapping matrices, U∈R m*l2_num , V∈R d*1 So, based on this, (U*L1 emb ) and (V*f l2_cls_emb ) can be understood as two vectors with dimensions l2_num*1; (U*L1 emb )*(V*f l2_cls_emb ) T It can be understood as the Hadamard product.
[0102] Optionally, the aforementioned interaction features may be calculated by a text processing device using a calculation method as shown in Formula 7:
[0103] Logits2=L1 emb *W*f l2_cls_emb Formula 7
[0104] Among them, Logits2 represents the interaction feature, L1 emb Represents the first semantic vector, which belongs to the vector space R m*1 (m and 1 are used to represent the dimension of the vector space), f l2_cls_emb Represents a word vector of a secondary category, which belongs to the vector space R l2_num*d (l2_num represents the total number of secondary categories, d represents the dimension of the word vector for each secondary category, usually ranging from 200 to 500); W is the parameter matrix.
[0105] Optionally, the semantic difference feature mentioned above can be obtained by performing vector difference processing on the first semantic vector and the word vector of the secondary category through a text processing device. Specifically, the text processing device can obtain the semantic difference feature by using the calculation method shown in Formula 8:
[0106]
[0107] S403 , using the optimized hierarchical classification model, based on the one or more first matching features and the one or more first-level prediction categories, determining a target second-level category of the text to be predicted from the one or more second-level categories.
[0108] In a specific implementation, a text processing device can determine a first prediction probability for each of one or more secondary categories based on one or more first matching features and one or more first-level prediction categories. The one or more first-level prediction categories are primarily used to assist the text processing device in determining the prediction probabilities of each secondary category. Specifically, based on the one or more first-level prediction categories, the secondary categories with higher prediction probabilities determined by the text processing device can be concentrated in the secondary categories corresponding to the one or more first-level prediction categories. In addition, the first prediction probability of any secondary category is used to indicate the probability that the optimized hierarchical classification model will use any secondary category as the target secondary category for the text to be predicted. Furthermore, the text processing device can determine the target secondary category for the text to be predicted from the one or more secondary categories based on the first prediction probability of each of the one or more second-level categories. The first prediction probability of the target secondary category is greater than the first prediction probability of any secondary category other than the target secondary category in the one or more second-level categories. In other words, the first prediction probability of the target secondary category is the largest first prediction probability among the first prediction probabilities corresponding to the one or more second-level categories.
[0109] S404: Obtain a target first-level category corresponding to the target second-level category according to a pre-established correspondence between the second-level category and the first-level category.
[0110] It can be seen from the above that each second-level category has a corresponding first-level category, that is, there is a corresponding relationship between any second-level category and its corresponding first-level category. Then, it is not difficult to understand that the text processing device can determine the first-level category corresponding to the target second-level category based on this corresponding relationship, and further use the determined first-level category as the target first-level category.
[0111] S405: Generate category information of the text to be predicted.
[0112] The category information includes target first-level categories and target second-level categories.
[0113] Based on the relevant description of the above steps S401-S405, in one embodiment, the text to be predicted can be a text obtained by the text processing device based on the analysis of the data query request; then, specifically, when the present application is applied to a data query scenario, the text processing device can first obtain the data query request, which carries query information. Exemplarily, the query information can be any multimedia information, such as voice information, video information, etc.; then, the text processing device can parse the query information to obtain the text to be predicted. It can be understood that the query information is the same as the query keyword indicated by the text to be predicted, that is, when the query information is voice information, the text processing device can perform voice recognition on the voice information to obtain text information corresponding to the voice information, and the text information can be used as the text to be predicted; when the query information is video information, the text processing device can perform video understanding on the video information to obtain video semantic information corresponding to the video information. The video semantic information can be text content, and the text content can be used as the text to be predicted. Based on this, the text processing device can further generate the category information of the text to be predicted according to the above steps S401-S405; after the text processing device obtains the category information, it can also search the database based on the category information for data content that matches the target first-level category included in the category information and also matches the target second-level category included in the category information, and can output the matched data content after the search is successful. Then, because the generated category information is more accurate, the user (or device) who initiates the data query request can receive the data content to be queried more quickly. Therefore, it can be seen that the text processing method provided by this application can improve data query efficiency and user experience to a certain extent.
[0114] Based on the above description, it is not difficult to understand that the present application can be applied to all business scenarios that require the extraction of text category information, such as: search content classification processing in search scenarios, product title classification of various e-commerce systems in online shopping scenarios, classification processing of various video contents in video playback scenarios, etc. Its application principle is similar to the above-mentioned data query scenario, and this application will no longer give examples one by one here.
[0115] In the present application, the text processing device adopts a multi-matching mechanism to construct multi-angle matching features, thereby determining the first matching feature based on the multi-angle matching features, strengthening the interaction between the first semantic vector and the word vector of the secondary category, so that the text processing device can capture the deeper semantic information between the text to be predicted and each secondary category, and further enable the text processing device to use the optimized hierarchical classification model to predict more accurate category information. In addition, since the secondary category itself also has semantic information, the present application can also enhance the ability of the hierarchical classification model to recognize semantic information by matching the first semantic vector and the word vector of the secondary category.
[0116] Based on the description of the above text processing method, the present application also discloses a text processing device, which can be a computer program (including program code) running in the above-mentioned text processing device. The text processing device can execute the following steps: Figure 2 and Figure 4 For the text processing method shown, see Figure 5 The text processing device may include at least: a first prediction unit 501, a matching processing unit 502, a second prediction unit 503, a processing unit 504, and a generation unit 505.
[0117] A first prediction unit 501 is configured to obtain a first semantic vector of a text to be predicted, and use an optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector, thereby obtaining one or more first-level predicted categories.
[0118] a matching processing unit 502 configured to match the first semantic vector with a word vector of each secondary category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes one or more secondary categories and a word vector of each secondary category in the one or more secondary categories;
[0119] A second prediction unit 503 is configured to use the optimized hierarchical classification model to determine a target secondary category of the text to be predicted from the one or more secondary categories based on the one or more first matching features and the one or more first-level prediction categories;
[0120] The processing unit 504 is configured to obtain a target primary category corresponding to the target secondary category based on a pre-established correspondence between the secondary category and the primary category;
[0121] The generating unit 505 is configured to generate category information of the text to be predicted, where the category information includes the target primary category and the target secondary category.
[0122] In one implementation, the matching processing unit 502 may be specifically configured to:
[0123] Traversing one or more secondary categories in the preset word vector set, performing vector matching processing on the word vector of the currently traversed secondary category and the first semantic vector to obtain a vector similarity feature, wherein the vector similarity feature is used to indicate: the similarity between the to-be-predicted text and the currently traversed secondary category;
[0124] Interactively matching the word vector of the currently traversed secondary category with the first semantic vector to obtain an interactive feature, where the interactive feature is used to indicate the correlation between the to-be-predicted text and the currently traversed secondary category;
[0125] Performing semantic matching processing on the word vector of the currently traversed secondary category and the first semantic vector to obtain a semantic difference feature, where the semantic difference feature is used to indicate the degree of difference between the semantics of the text to be predicted and the semantics of the currently traversed secondary category;
[0126] Feature fusion processing is performed on the vector similarity feature, the interaction feature, and the semantic difference feature to obtain a first matching feature corresponding to the currently traversed secondary category, so as to obtain the one or more first matching features after the traversal is completed.
[0127] In another embodiment, the text processing apparatus further includes a model optimization unit 506, which can be configured to perform:
[0128] Obtaining a training sample, the training sample comprising: a training text, and standard first-level categories and standard second-level categories of the training text;
[0129] Decoding the training text to obtain a second semantic vector of the training text;
[0130] Using a hierarchical classification model, determining a second predicted probability of the standard first-level category based on the second semantic vector, the second predicted probability being used to indicate: the probability of the hierarchical classification model outputting the standard first-level category;
[0131] Matching the second semantic vector with the word vector of each of the one or more secondary categories to obtain one or more second matching features;
[0132] Determining, using the hierarchical classification model, a third predicted probability of the standard secondary category based on the one or more second matching features, the third predicted probability being used to indicate: a probability of the hierarchical classification model outputting the standard secondary category;
[0133] Based on the second predicted probability and the third predicted probability, a target loss value is determined, and the model parameters of the hierarchical classification model are optimized in a direction of reducing the target loss value to obtain the optimized hierarchical classification model.
[0134] In another embodiment, the training sample includes: N training texts, and a standard first-level category and a standard second-level category of each of the N training texts, where N is a positive integer and N>1;
[0135] The determined second prediction probability includes N second prediction probabilities, the N second prediction probabilities including a second prediction probability of a standard first-level category of each training text in the N training texts, the second prediction probability of the standard first-level category of each training text being determined based on a second semantic vector of each training text;
[0136] The determined third prediction probability includes N third prediction probabilities, where the N third prediction probabilities include a third prediction probability of a standard secondary category of each training text in the N training texts, where the third prediction probability of the standard secondary category of each training text is determined based on one or more second matching features corresponding to each training text;
[0137] The model optimization unit 506 is further configured to execute:
[0138] Obtaining a first loss value based on the N second prediction probabilities and the N third prediction probabilities;
[0139] Determining a second loss value based on a second predicted probability of a standard first-level category of any training text among the N training texts, and determining a third loss value based on a third predicted probability of a standard second-level category of any training text;
[0140] The target loss value is determined based on the first loss value, the second loss value, and the third loss value.
[0141] In yet another embodiment, the model optimization unit 506 is further configured to perform:
[0142] Determining a first true value of the standard first-level category of any one of the training texts based on the standard first-level category of the any one of the training texts;
[0143] determining the second loss value based on the first true value and the second predicted probability of the standard first-level category of the any training text, wherein the second loss value is negatively correlated with the second predicted probability of the standard first-level category of the any training text;
[0144] Determining a second true value of the standard secondary category of any one of the training texts based on the standard secondary category of the any one of the training texts;
[0145] The third loss value is determined based on the second true value and the third predicted probability of the standard secondary category of any one of the training texts, and the third loss value is negatively correlated with the third predicted probability of the standard secondary category of any one of the training texts.
[0146] In yet another embodiment, the second prediction unit 503 is further configured to perform:
[0147] Determining a first prediction probability for each of the one or more second-level categories based on the one or more first matching features and the one or more first-level prediction categories, wherein the first prediction probability of any second-level category indicates a probability that the optimized hierarchical classification model uses the any second-level category as a target second-level category for the text to be predicted;
[0148] Based on the first prediction probability of each secondary category in the one or more secondary categories, a target secondary category of the text to be predicted is determined from the one or more secondary categories, the first prediction probability of the target secondary category being greater than the first prediction probability of any secondary category in the one or more secondary categories other than the target secondary category.
[0149] In yet another embodiment, the processing unit 504 may also be configured to execute:
[0150] Obtaining a data query request, where the data query request carries query information;
[0151] parsing the query information to obtain the text to be predicted, wherein the query information and the query keyword indicated by the text to be predicted are the same;
[0152] After the generating unit 505 generates the category information of the text to be predicted, the processing unit 504 may further be configured to:
[0153] Searching for data content that matches the target primary category and the target secondary category;
[0154] Output the matched data content.
[0155] According to one embodiment of the present application, Figure 2 and Figure 4 The steps involved in the method shown can be performed by Figure 5 The text processing device shown in FIG. Figure 2 Steps S201 to S206 shown in FIG. Figure 5The model optimization unit 506 in the text processing device shown is executed. Figure 4 Step S401 shown can be performed by Figure 5 The first prediction unit 501 in the text processing device shown in FIG. 4 is executed; step S402 can be performed by Figure 5 The matching processing unit 502 in the text processing device shown in FIG. 4 is used to perform step S403. Figure 5 The second prediction unit 503 in the text processing apparatus shown in FIG. 4 is executed, and step S404 can be performed by Figure 5 The processing unit 504 in the text processing device shown in FIG. 4 is executed, and step S405 can be performed by Figure 5 The text processing apparatus shown is executed by the generating unit 505.
[0156] According to another embodiment of the present application, Figure 5 The various units in the text processing device shown are divided based on logical functions. The above-mentioned units can be individually or all combined into one or more other units to form a structure, or one (or some) of the units can be further divided into multiple functionally smaller units to form a structure, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. In other embodiments of the present application, the above-mentioned text processing device can also include other units. In actual applications, these functions can also be assisted by other units and can be achieved by the collaboration of multiple units.
[0157] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. Figure 2 or Figure 4 A computer program (including program code) for each step of the method shown is used to construct Figure 5 The text processing device shown in and the text processing method of the embodiment of the present application are implemented. The computer program can be recorded on a computer storage medium, for example, and loaded into the above-mentioned computing device through the computer storage medium and run therein.
[0158] In an embodiment of the present application, the text processing device constructs a first matching feature between the word vector of the secondary category and the semantic vector of the text to be predicted based on a variety of different matching features, so that the accuracy of the target secondary category predicted by the hierarchical classification model is higher; in addition, since the text processing device is based on the word vectors of each secondary category, the first matching feature determined can enable the text processing device to better utilize the semantic information of the secondary category, thereby further improving the accuracy of the category information.
[0159] Based on the relevant descriptions of the above method embodiment and apparatus embodiment, the present application embodiment further provides a text processing device, see Figure 6 The text processing device at least includes a processor 601 and a computer storage medium 602, and the processor 601 and the computer storage medium 602 of the text processing device can be connected via a bus or other means.
[0160] The aforementioned computer storage medium 602 is a memory device within the text processing device, used to store programs and data. It is understood that the computer storage medium 602 herein may include both the built-in storage medium within the text processing device and, of course, the extended storage medium supported by the text processing device. The computer storage medium 602 provides storage space, which stores the operating system of the text processing device. Furthermore, this storage space also stores one or more computer programs suitable for loading and executing by the processor 601. These computer programs may be one or more program codes. It should be noted that the computer storage medium herein may be high-speed RAM memory or non-volatile memory, such as at least one disk drive; optionally, it may be at least one computer storage medium located remotely from the aforementioned processor. The processor 601 (also known as the CPU (Central Processing Unit)) is the computing and control core of the text processing device, adapted to execute one or more computer programs, specifically, to load and execute one or more computer programs to implement the corresponding method flow or corresponding function.
[0161] In one embodiment, the processor 601 may load and execute one or more computer programs stored in the computer storage medium 602 to implement the above-mentioned Figure 2 and Figure 4 The corresponding method steps in the method embodiment shown; in a specific implementation, one or more computer programs in the computer storage medium 602 are loaded by the processor 601 and execute the following steps:
[0162] Obtain a first semantic vector of a text to be predicted, and use an optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector to obtain one or more first-level predicted categories; match the first semantic vector with a word vector of each second-level category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes: one or more second-level categories, and a word vector of each second-level category in the one or more second-level categories; use the optimized hierarchical classification model to determine a target second-level category of the text to be predicted from the one or more second-level categories based on the one or more first matching features and the one or more first-level predicted categories; obtain a target first-level category corresponding to the target second-level category based on a pre-established correspondence between the second-level category and the first-level category; generate category information of the text to be predicted, wherein the category information includes the target first-level category and the target second-level category.
[0163] In one embodiment, the processor 601 may be specifically configured to load and execute:
[0164] Traversing one or more secondary categories in the preset word vector set, performing vector matching processing on the word vector of the currently traversed secondary category and the first semantic vector to obtain a vector similarity feature, wherein the vector similarity feature is used to indicate: the similarity between the to-be-predicted text and the currently traversed secondary category;
[0165] Interactively matching the word vector of the currently traversed secondary category with the first semantic vector to obtain an interactive feature, where the interactive feature is used to indicate the correlation between the to-be-predicted text and the currently traversed secondary category;
[0166] Performing semantic matching processing on the word vector of the currently traversed secondary category and the first semantic vector to obtain a semantic difference feature, where the semantic difference feature is used to indicate the degree of difference between the semantics of the text to be predicted and the semantics of the currently traversed secondary category;
[0167] Feature fusion processing is performed on the vector similarity feature, the interaction feature, and the semantic difference feature to obtain a first matching feature corresponding to the currently traversed secondary category, so as to obtain the one or more first matching features after the traversal is completed.
[0168] In yet another embodiment, the processor 601 may be specifically configured to load and execute:
[0169] Obtaining a training sample, the training sample comprising: a training text, and standard first-level categories and standard second-level categories of the training text;
[0170] Decoding the training text to obtain a second semantic vector of the training text;
[0171] Using a hierarchical classification model, determining a second predicted probability of the standard first-level category based on the second semantic vector, the second predicted probability being used to indicate: the probability of the hierarchical classification model outputting the standard first-level category;
[0172] Matching the second semantic vector with the word vector of each of the one or more secondary categories to obtain one or more second matching features;
[0173] Determining, using the hierarchical classification model, a third predicted probability of the standard secondary category based on the one or more second matching features, the third predicted probability being used to indicate: a probability of the hierarchical classification model outputting the standard secondary category;
[0174] Based on the second predicted probability and the third predicted probability, a target loss value is determined, and the model parameters of the hierarchical classification model are optimized in a direction of reducing the target loss value to obtain the optimized hierarchical classification model.
[0175] In another embodiment, the training sample includes: N training texts, and a standard first-level category and a standard second-level category of each of the N training texts, where N is a positive integer and N>1;
[0176] The determined second prediction probability includes N second prediction probabilities, the N second prediction probabilities including a second prediction probability of a standard first-level category of each training text in the N training texts, the second prediction probability of the standard first-level category of each training text being determined based on a second semantic vector of each training text;
[0177] The determined third prediction probability includes N third prediction probabilities, where the N third prediction probabilities include a third prediction probability of a standard secondary category of each training text in the N training texts, where the third prediction probability of the standard secondary category of each training text is determined based on one or more second matching features corresponding to each training text;
[0178] The processor 601 may also be specifically configured to load and execute:
[0179] Obtaining a first loss value based on the N second prediction probabilities and the N third prediction probabilities;
[0180] Determining a second loss value based on a second predicted probability of a standard first-level category of any training text among the N training texts, and determining a third loss value based on a third predicted probability of a standard second-level category of any training text;
[0181] The target loss value is determined based on the first loss value, the second loss value, and the third loss value.
[0182] In yet another embodiment, the processor 601 may further be specifically configured to load and execute:
[0183] Determining a first true value of the standard first-level category of any one of the training texts based on the standard first-level category of the any one of the training texts;
[0184] determining the second loss value based on the first true value and the second predicted probability of the standard first-level category of the any training text, wherein the second loss value is negatively correlated with the second predicted probability of the standard first-level category of the any training text;
[0185] Determining a second true value of the standard secondary category of any one of the training texts based on the standard secondary category of the any one of the training texts;
[0186] The third loss value is determined based on the second true value and the third predicted probability of the standard secondary category of any one of the training texts, and the third loss value is negatively correlated with the third predicted probability of the standard secondary category of any one of the training texts.
[0187] In yet another embodiment, the processor 601 may further be specifically configured to load and execute:
[0188] Determining a first prediction probability for each of the one or more second-level categories based on the one or more first matching features and the one or more first-level prediction categories, wherein the first prediction probability of any second-level category indicates a probability that the optimized hierarchical classification model uses the any second-level category as a target second-level category for the text to be predicted;
[0189] Based on the first prediction probability of each secondary category in the one or more secondary categories, a target secondary category of the text to be predicted is determined from the one or more secondary categories, the first prediction probability of the target secondary category being greater than the first prediction probability of any secondary category in the one or more secondary categories other than the target secondary category.
[0190] In yet another embodiment, the processor 601 may further be specifically configured to load and execute:
[0191] Obtaining a data query request, where the data query request carries query information;
[0192] parsing the query information to obtain the text to be predicted, wherein the query information and the query keyword indicated by the text to be predicted are the same;
[0193] The processor 601 is specifically configured to load and execute: after generating the category information of the text to be predicted, the processor 601 may also be specifically configured to load and execute:
[0194] Searching for data content that matches the target primary category and the target secondary category;
[0195] Output the matched data content.
[0196] In an embodiment of the present application, the text processing device constructs a first matching feature between the word vector of the secondary category and the semantic vector of the text to be predicted based on a variety of different matching features, so that the accuracy of the target secondary category predicted by the hierarchical classification model is higher; in addition, since the text processing device is based on the word vectors of each secondary category, the first matching feature determined can enable the text processing device to better utilize the semantic information of the secondary category, thereby further improving the accuracy of the category information.
[0197] The present application also provides a storage medium storing a computer program for the above-mentioned text processing method. The computer program includes program instructions. When one or more processors load and execute the program instructions, the text processing method described in the embodiment can be implemented, which is not repeated here. The description of the beneficial effects of adopting the same method is not repeated here. It is understood that the program instructions can be deployed and executed on one or more devices that can communicate with each other.
[0198] It should be noted that, according to one aspect of the present application, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor in a text processing device reads the computer instructions from the computer-readable storage medium and then executes the computer instructions, thereby enabling the text processing device to perform the above-mentioned Figure 2 and Figure 4 The text processing method shown in the embodiment is provided in various optional ways.
[0199] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed, the computer program can include the processes in the above-described text processing method embodiments. The computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0200] The above disclosure is only a partial embodiment of the present application, and it is certainly not intended to limit the scope of the rights of the present application. A person skilled in the art can understand that all or part of the processes of the above embodiments and equivalent changes made in accordance with the claims of the present application are still within the scope of the invention.
Claims
1. A text processing method, characterized in that: include: Obtaining a first semantic vector of the text to be predicted, and using the optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector, to obtain one or more first-level predicted categories; Matching the first semantic vector with the word vector of each secondary category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes: one or more secondary categories, and the word vector of each secondary category in the one or more secondary categories; the first matching feature is obtained by fusing multiple matching features, wherein the multiple matching features include: a vector similarity feature obtained by vector matching, an interaction feature obtained by interactive matching, and a semantic difference feature obtained by semantic matching; the vector similarity feature is used to indicate: the similarity between the text to be predicted and the secondary category, the interaction feature is used to indicate: the correlation between the text to be predicted and the secondary category, and the semantic difference feature is used to indicate: the difference between the semantics of the text to be predicted and the semantics of the secondary category; Using the optimized hierarchical classification model, based on the one or more first matching features and the one or more first-level prediction categories, determining a target second-level category of the text to be predicted from the one or more second-level categories; the optimized hierarchical classification model constrains the prediction probability of the second-level category to be less than the prediction probability of the first-level category corresponding to the second-level category by using a first loss function; Obtaining a target first-level category corresponding to the target second-level category based on a pre-established correspondence between the second-level category and the first-level category; Category information of the text to be predicted is generated, where the category information includes the target primary category and the target secondary category.
2. The method according to claim 1, characterized in that The matching process of the first semantic vector with the word vector of each secondary category in the preset word vector set to obtain one or more first matching features includes: Traversing one or more secondary categories in the preset word vector set, performing vector matching processing on the word vector of the currently traversed secondary category and the first semantic vector to obtain vector similarity features; Interactively matching the word vector of the currently traversed secondary category with the first semantic vector to obtain an interactive feature; Performing semantic matching processing on the word vector of the currently traversed secondary category and the first semantic vector to obtain a semantic difference feature; Feature fusion processing is performed on the vector similarity feature, the interaction feature, and the semantic difference feature to obtain a first matching feature corresponding to the currently traversed secondary category, so as to obtain the one or more first matching features after the traversal is completed.
3. The method according to claim 1, characterized in that The method further comprises: Acquire a training sample, the training sample comprising: a training text, and a standard first-level category and a standard second-level category of the training text; Decoding the training text to obtain a second semantic vector of the training text; Using a hierarchical classification model, determining a second predicted probability of the standard first-level category based on the second semantic vector, the second predicted probability being used to indicate: the probability of the hierarchical classification model outputting the standard first-level category; Matching the second semantic vector with the word vector of each of the one or more secondary categories to obtain one or more second matching features; Determining, using the hierarchical classification model, a third predicted probability of the standard secondary category based on the one or more second matching features, the third predicted probability being used to indicate: a probability of the hierarchical classification model outputting the standard secondary category; Based on the second predicted probability and the third predicted probability, a target loss value is determined, and the model parameters of the hierarchical classification model are optimized in a direction of reducing the target loss value to obtain the optimized hierarchical classification model.
4. The method according to claim 3, characterized in that The training sample includes: N training texts, and a standard first-level category and a standard second-level category of each training text in the N training texts, where N is a positive integer and N>1; The determined second prediction probability includes N second prediction probabilities, the N second prediction probabilities including a second prediction probability of a standard first-level category of each training text in the N training texts, the second prediction probability of the standard first-level category of each training text being determined based on a second semantic vector of each training text; The determined third prediction probability includes N third prediction probabilities, where the N third prediction probabilities include a third prediction probability of a standard secondary category of each training text in the N training texts, where the third prediction probability of the standard secondary category of each training text is determined based on one or more second matching features corresponding to each training text; The determining a target loss value based on the second predicted probability and the third predicted probability includes: Obtaining a first loss value based on the N second prediction probabilities and the N third prediction probabilities; Determining a second loss value based on a second predicted probability of a standard first-level category of any training text among the N training texts, and determining a third loss value based on a third predicted probability of a standard second-level category of any training text; The target loss value is determined based on the first loss value, the second loss value, and the third loss value.
5. The method according to claim 4, characterized in that Determining a second loss value based on a second predicted probability of a standard first-level category of any training text among the N training texts, and determining a third loss value based on a third predicted probability of a standard second-level category of any training text; include: Determining a first true value of the standard first-level category of any one of the training texts based on the standard first-level category of the any one of the training texts; determining the second loss value based on the first true value and the second predicted probability of the standard first-level category of the any training text, wherein the second loss value is negatively correlated with the second predicted probability of the standard first-level category of the any training text; Determining a second true value of the standard secondary category of any one of the training texts based on the standard secondary category of the any one of the training texts; The third loss value is determined based on the second true value and the third predicted probability of the standard secondary category of any one of the training texts, and the third loss value is negatively correlated with the third predicted probability of the standard secondary category of any one of the training texts.
6. The method according to claim 1, wherein The determining, based on the one or more first matching features and the one or more first-level prediction categories, a target second-level category of the text to be predicted from the one or more second-level categories includes: Determining a first prediction probability for each of the one or more second-level categories based on the one or more first matching features and the one or more first-level prediction categories, wherein the first prediction probability of any second-level category indicates a probability that the optimized hierarchical classification model uses the any second-level category as a target second-level category for the text to be predicted; Based on the first prediction probability of each secondary category in the one or more secondary categories, a target secondary category of the text to be predicted is determined from the one or more secondary categories, the first prediction probability of the target secondary category being greater than the first prediction probability of any secondary category in the one or more secondary categories other than the target secondary category.
7. The method according to claim 1, characterized in that Before obtaining the first semantic vector of the text to be predicted, the method further includes: Obtaining a data query request, where the data query request carries query information; parsing the query information to obtain the text to be predicted, wherein the query information and the query keyword indicated by the text to be predicted are the same; After generating the category information of the text to be predicted, the method further includes: Searching for data content that matches the target primary category and the target secondary category; Output the matched data content.
8. A text processing device, characterized in that: include: a first prediction unit, configured to obtain a first semantic vector of a text to be predicted, and use an optimized hierarchical classification model to predict a first-level category of the text to be predicted based on the first semantic vector, to obtain one or more first-level predicted categories; a matching processing unit, configured to match the first semantic vector with the word vector of each secondary category in a preset word vector set to obtain one or more first matching features, wherein the preset word vector set includes: one or more secondary categories, and the word vector of each secondary category in the one or more secondary categories; the first matching feature is obtained based on the fusion of multiple matching features, wherein the multiple matching features include: a vector similarity feature obtained through vector matching processing, an interaction feature obtained through interactive matching processing, and a semantic difference feature obtained through semantic matching processing; the vector similarity feature is used to indicate: the similarity between the text to be predicted and the secondary category, the interaction feature is used to indicate: the correlation between the text to be predicted and the secondary category, and the semantic difference feature is used to indicate: the difference between the semantics of the text to be predicted and the semantics of the secondary category; a second prediction unit, configured to use the optimized hierarchical classification model to determine a target secondary category of the text to be predicted from the one or more secondary categories based on the one or more first matching features and the one or more first-level prediction categories; the optimized hierarchical classification model constraining the prediction probability of the secondary category to be less than the prediction probability of the first-level category corresponding to the secondary category by using a first loss function; a processing unit, configured to obtain a target primary category corresponding to the target secondary category based on a pre-established correspondence between the secondary categories and the primary categories; A generating unit is configured to generate category information of the text to be predicted, wherein the category information includes the target primary category and the target secondary category.
9. A text processing device, characterized in that: include: a processor adapted to execute one or more computer programs; A computer storage medium storing one or more computer programs, wherein the one or more computer programs are suitable for being loaded by the processor and executing the text processing method according to any one of claims 1 to 7.
10. A storage medium storing one or more computer programs, wherein the one or more computer programs are suitable for being loaded by a processor and executing the text processing method according to any one of claims 1 to 7.
11. A computer program product comprising computer instructions, characterized in that The computer instructions are stored in a computer-readable storage medium and are suitable for being read and executed by a processor of a text processing device, so that the text processing device executes the text processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Text hierarchy classification method, electronic equipment and storage medium
CN112069321A
Commodity retrieval method and device, computer equipment and storage medium
CN112241493A