Text processing method, device, medium and electronic device

By obtaining text samples and their corresponding sample tag sequences, using the label prediction model to predict the Nth tag, and adjusting the model parameters, the problem of difficulty in utilizing the tag hierarchy relationship in the prior art is solved, and more accurate text hierarchy classification is achieved.

CN113553401BActive Publication Date: 2025-05-06NETEASE MEDIA TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110857343.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-28
Publication Date
2025-05-06
Estimated Expiration
2041-07-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the hierarchical relationship between labels in text-level classification tasks, resulting in poor prediction results.

Method used

By obtaining text samples and their corresponding sample tag sequences, the label prediction model is used to predict the Nth tag based on the first N-1 sample tags, and the model parameters are adjusted to combine hierarchical relationships.

Benefits of technology

It realizes the more accurate combination of the hierarchical relationship between multiple labels in the text-level classification task, and improves the effect of label prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113553401B_ABST
    Figure CN113553401B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure provides a text processing method. The method may include: obtaining a text sample and a corresponding sample label sequence, wherein the arrangement order of the multiple sample labels included in the sample label sequence indicates the hierarchical relationship. The text sample and the sample label sequence are input into a label prediction model, and based on the text sample and at least part of the sample labels in the first N-1 sample labels in the sample label sequence, the Nth predicted label in the predicted label sequence corresponding to the text sample is obtained to predict each predicted label in the predicted label sequence step by step; based on the sample label sequence and the predicted label sequence, the parameters of the label prediction model are adjusted. In this way, the input text can be accurately labeled in combination with the hierarchical relationship between multiple labels. In addition, an embodiment of the present disclosure provides a text processing device, a medium and an electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of natural language processing, and more specifically, the embodiments of the present disclosure relate to a text processing method, device, medium and electronic device. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the present disclosure as recited in the claims. No description herein is admitted to be related art by inclusion in this section.

[0003] The text level classification task is a multi-label classification task in natural language processing. In this task, it is necessary to use a neural model to process the input text data and predict the multiple labels corresponding to the text data.

[0004] The label can be used to indicate the specific type of the text data in the preset classification system. For example, in the news classification system, the label can indicate the type of news corresponding to the text data. It should be noted that the preset classification system recorded in the present disclosure is not limited to the news classification system, but can also be other classification systems.

[0005] There is a hierarchical relationship between the multiple tags. The hierarchical relationship may include a hierarchical relationship and a same-level relationship in the classification system. For example, in a news classification system, three levels may be included. The three levels are represented by a progressive type division, that is, the lower the level, the finer the type division. Tags at the same level have a same-level relationship. Two tags across levels have a hierarchical relationship.

[0006] Currently, in text-level classification tasks, label prediction mainly relies on flattened models, local models or global models. Summary of the invention

[0007] However, these methods cannot make good use of the hierarchical relationship between labels when predicting labels, and it is difficult to achieve good prediction results.

[0008] Therefore, in the related art, the text level classification task cannot be performed well. This is very annoying.

[0009] To this end, an improved text processing method is urgently needed to combine the hierarchical relationship between multiple labels and accurately predict the labels of the input text.

[0010] In this context, embodiments of the present disclosure are intended to provide a text processing method, apparatus, medium, and electronic device.

[0011] In a first aspect of the embodiments of the present disclosure, a text processing method is provided, comprising: obtaining a text sample and a sample label sequence corresponding to the text sample, the sample label sequence comprising a plurality of sample labels having a hierarchical relationship, and the arrangement order of the plurality of sample labels indicating the hierarchical relationship; taking the text sample and the sample label sequence as input, and using a label prediction model, based on the text sample and at least part of the sample labels among the first N-1 sample labels in the sample label sequence, obtaining the Nth predicted label in the predicted label sequence corresponding to the text sample, so as to predict each predicted label in the predicted label sequence step by step; the N represents the sample label sequence or the label order in the predicted label sequence; and adjusting the parameters of the label prediction model based on the sample label sequence and the predicted label sequence.

[0012] In some embodiments, the method of obtaining the Nth predicted label in the predicted label sequence corresponding to the text sample based on the text sample and at least part of the sample labels among the first N-1 sample labels in the sample label sequence comprises: obtaining the weights corresponding to at least part of the sample labels among the first N-1 sample labels in the sample label sequence based on the attention mechanism; obtaining the Nth predicted label in the predicted label sequence corresponding to the text sample based on the text sample, the part of the sample labels, and the weights corresponding to the part of the sample labels.

[0013] In some embodiments, before adjusting the parameters of the label prediction model based on the sample label sequence and the predicted label sequence, it also includes: maintaining the correspondence between each sample label and its ancestral label according to the ancestor-descendant relationship between each sample label included in the sample label sequence; based on the correspondence, screening out the weights corresponding to the ancestral labels of each predicted label in the predicted label sequence; inputting the weights corresponding to the ancestral labels of each predicted label into a second loss function to obtain second loss information; the second loss function is used to increase the weights corresponding to the ancestral labels of the predicted labels during the training stage of the label prediction model.

[0014] In some embodiments, the parameters of the label prediction model are adjusted based on the sample label sequence and the predicted label sequence, including: inputting the sample label sequence and the predicted label sequence into a first loss function to obtain first loss information; the first loss function indicates the error between the predicted label sequence and the sample label sequence; obtaining the second loss information; and adjusting the parameters of the label prediction model based on the first loss information and the second loss information.

[0015] In some embodiments, the method for obtaining the sample label sequence includes: obtaining a label tree; the label tree includes multiple label nodes corresponding to multiple labels respectively; the hierarchical relationship between the multiple label nodes indicates the hierarchical relationship between the multiple labels; traversing the multiple label nodes level by level to obtain the target label node corresponding to the text sample to obtain the sample label sequence.

[0016] In some embodiments, the level-by-level traversal includes a breadth-first traversal.

[0017] In some embodiments, the label prediction model includes a conversion model.

[0018] In some embodiments, the natural language processing model comprises a text-to-text transfer conversion model.

[0019] In some embodiments, the method further includes: acquiring a target text; inputting the target text into the trained label prediction model to obtain a predicted label sequence corresponding to the target text.

[0020] In a second aspect of the embodiments of the present disclosure, a text processing device is provided, comprising: an acquisition module, for acquiring a text sample, and a sample label sequence corresponding to the text sample; the sample label sequence comprises a plurality of sample labels having a hierarchical relationship; the arrangement order of the plurality of sample labels indicates the hierarchical relationship; a first prediction module, for taking the text sample and the sample label sequence as input, and using a label prediction model, based on the text sample and at least part of the sample labels among the first N-1 sample labels in the sample label sequence, obtaining the Nth predicted label in the predicted label sequence corresponding to the text sample, so as to predict each predicted label in the predicted label sequence step by step; the N represents the sample label sequence or the label order in the predicted label sequence; an adjustment module, for adjusting the parameters of the label prediction model based on the sample label sequence and the predicted label sequence.

[0021] In a third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the medium stores a computer program, and the computer program is used to enable a processor to execute the text processing method as shown in any of the aforementioned embodiments.

[0022] In a fourth aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor executable instructions; wherein the processor implements the text processing method as shown in any of the foregoing embodiments by running the executable instructions.

[0023] In the technical solution described above, the acquired text sample and sample label sequence can be used as input, and the label prediction model can be used to obtain the Nth predicted label in the predicted label sequence corresponding to the text sample based on the text sample and at least part of the sample labels in the first N-1 sample labels in the sample label sequence, so as to predict each predicted label in the predicted label sequence step by step; wherein the sample label sequence includes multiple sample labels with a hierarchical relationship, and the arrangement order of the multiple sample labels indicates the hierarchical relationship. Then, based on the sample label sequence and the predicted label sequence, the parameters of the label prediction model are adjusted.

[0024] Therefore, during the label prediction model training process, the model can learn the correspondence between text samples and each sample label, and when obtaining the Nth label, it can make full use of the label information before the label. Therefore, when using the model for label prediction, the hierarchical relationship between multiple labels can be combined to achieve better label prediction effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The foregoing and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:

[0026] Figure 1 A schematic diagram of a tag tree structure shown in an embodiment of the present disclosure;

[0027] Figure 2 A schematic diagram of an application scenario of text processing shown in an embodiment of the present disclosure;

[0028] Figure 3 A method flow chart of a text processing method shown in an embodiment of the present disclosure;

[0029] Figure 4 A flowchart of a tag prediction method according to an embodiment of the present disclosure is shown;

[0030] Figure 5 A schematic diagram of a label prediction model structure shown in an embodiment of the present disclosure;

[0031] Figure 6 A schematic diagram of a mask matrix shown in an embodiment of the present disclosure;

[0032] Figure 7 A flowchart of a method for determining loss information according to an embodiment of the present disclosure is shown;

[0033] Figure 8 A schematic diagram of a method for obtaining a sample tag sequence according to an embodiment of the present disclosure;

[0034] Fig. 9 A flowchart of a label prediction method shown in an embodiment of the present disclosure;

[0035] Fig.10 A schematic diagram of a model training method flow chart shown in an embodiment of the present disclosure;

[0036] Fig.11 A structural schematic diagram of a text processing device shown in an embodiment of the present disclosure;

[0037] Fig.12 A program product applied to a text processing method shown in an embodiment of the present disclosure;

[0038] Fig.13 The present invention is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.

[0039] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0040] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0041] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0042] According to an embodiment of the present disclosure, a text processing method, medium, device and electronic device are proposed.

[0043] In this document, it is to be understood that the terms involved are as follows.

[0044] The neural model is an algorithmic mathematical model that imitates the behavioral characteristics of animal neural models and performs distributed parallel information processing. This model can achieve the purpose of processing information by adjusting the interconnected relationships between a large number of internal nodes.

[0045] Natural language processing is the process of mathematically modeling human language, analyzing and processing it using computers, and exploring the laws and patterns in language to mine its value based on actual needs.

[0046] The label tree refers to a label classification system stored in a tree structure. The label tree includes a plurality of label nodes corresponding to a plurality of sample labels respectively; the hierarchical relationship between the plurality of label nodes indicates the hierarchical relationship between the plurality of sample labels.

[0047] See also Figure 1 , Figure 1 A schematic diagram of the structure of a tag tree shown in an embodiment of the present disclosure.

[0048] Figure 1 The label tree shown is a label tree of a news classification system. The label tree is divided into three levels. Among them, the first level can include label nodes corresponding to the three labels of special topics, comments, and news. These three labels are in the same level relationship. The second level can include label nodes corresponding to the three labels of art, business, and sports. Among them, special topics and art have a hierarchical relationship. That is, art is a subcategory of special topics. Business and sports have a hierarchical relationship with news. The third level can include label nodes corresponding to the five labels of dance, music, basketball, football, and hockey. Among them, dance and music have a hierarchical relationship with art. Basketball, football, and hockey have a hierarchical relationship with sports.

[0049] Furthermore, any number of elements in the drawings is for illustrative purposes only and not limiting, and any naming is for distinction only and does not have any limiting meaning.

[0050] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION

[0052] On the one hand, the present inventors have discovered that if the labels in the multi-label are arranged in descending order according to the hierarchical relationship to obtain a label sequence, the text level classification task can be transformed into a sequence-to-sequence text processing process, so that the text processing model can be used to learn the mapping relationship between the input text sequence and the label sequence, and perform the text level classification task.

[0053] On the other hand, the present inventors have discovered that the arrangement order of each label in the label sequence can indicate the hierarchical relationship between the labels. If the label before the Nth label can be combined when predicting the Nth label, the hierarchical relationship between multiple labels can be combined in the label prediction process to accurately predict the label of the input text.

[0054] Therefore, the present inventors consider enabling the model to learn the correspondence between text samples and sample labels during the model training phase, and to fully utilize the label information before the Nth label when the model is obtained. This allows the model to be used for label prediction by combining the hierarchical relationship between multiple labels to achieve better label prediction results.

[0055] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.

[0056] Application Scenario Overview

[0057] Please refer to Figure 2 , Figure 2 A schematic diagram of an application scenario of text processing shown in an embodiment of the present disclosure.

[0058] like Figure 2 As shown, schematically, the aforementioned application scenario may include client terminals such as a mobile phone 2012, a tablet computer 2013, a computer 2011, and a server 202 equipped with text level classification logic.

[0059] Schematically, the aforementioned terminal can collect text data in the form of text, image, voice, etc., and transmit the collected text data to the server 202 for processing.

[0060] The server 202 may be equipped with model training logic and label prediction logic. The model training logic may use the input text sample set to train the label prediction model carried by the server, so that the model can learn the correspondence between text samples and each sample label, and when obtaining the Nth label, it can fully utilize the label information before the label.

[0061] The label prediction logic can perform label prediction on the input target text to obtain a predicted label sequence 203 of the input text. The predicted label sequence 203 may include multiple labels with a hierarchical relationship. When predicting the Nth label of the sequence 203, the label information before the label can be fully utilized, so that when the model is used for label prediction, the hierarchical relationship between multiple labels can be combined to achieve a better label prediction effect.

[0062] Exemplary Methods

[0063] See also Figure 3 , Figure 3 A method flow chart of a text processing method shown in an embodiment of the present disclosure.

[0064] Figure 3The text processing method shown can be applied to an electronic device. The electronic device can execute the method by carrying software logic corresponding to the text processing method. The type of the electronic device can be a laptop, a computer, a mobile phone, a PAD (Packet Assembler And Disassembler, providing a terminal to host link service) terminal, etc. The type of the electronic device is not particularly limited in the present disclosure. The electronic device can also be a user-end device or a server-end device, which is not particularly limited here.

[0065] like Figure 3 As shown, the text processing method may include S302-S306. The text processing method is described in an embodiment from the perspective of model training.

[0066] In S302, a text sample and a sample label sequence corresponding to the text sample are obtained, wherein the sample label sequence includes a plurality of sample labels having a hierarchical relationship, and an arrangement order of the plurality of sample labels indicates the hierarchical relationship.

[0067] The text sample may include any combination of words. For example, the text sample may be a news report or a manuscript. The text sample is used for supervised model training. In general, multiple labels with hierarchical relationships corresponding to the text sample may be obtained in advance.

[0068] The sample label sequence may be a sequence obtained by sorting the multiple labels corresponding to the text samples in descending order of hierarchy.

[0069] by Figure 1 Take the label tree shown in the figure as an example. Assume that the multiple labels corresponding to the text sample A include topics, art, news, sports, and football. According to the hierarchical relationship between these labels, we can get the sample label sequence {topics, news, art, sports, football} by sorting them from high to low.

[0070] In some embodiments, special symbols indicating hierarchical relationships and special symbols indicating the end of a sequence may be added between tags. For example, "_ (underscore)" may be used to indicate that the two tags are in the same hierarchical relationship. " / (slash)" may be used to indicate that the two tags are in a hierarchical relationship. "EOS" may be used to indicate the end of a sequence.

[0071] by Figure 1 Take the label tree shown as an example. Assume that the multi-labels corresponding to the text sample A include topics, art, news, sports, football ( Figure 1Dark label nodes in the middle). According to the hierarchical relationship between these labels, we can sort them from high to low to get the sample label sequence {topic, _, news, / , art, _, sports, / , football, EOS}. Therefore, the hierarchical relationship of each label can be accurately reflected in the sample label sequence, which helps to improve the model learning effect.

[0072] S304, taking the text sample and the sample label sequence as input, using a label prediction model, based on the text sample and at least part of the sample labels among the first N-1 sample labels in the sample label sequence, obtain the Nth predicted label in the predicted label sequence corresponding to the text sample, so as to predict each predicted label in the predicted label sequence step by step; N represents the label order in the sample label sequence or the predicted label sequence.

[0073] The label prediction model may include a text processing model built based on a neural model. The label prediction model may be used to obtain predicted labels with a hierarchical relationship corresponding to the input text, and obtain a predicted label sequence.

[0074] In some embodiments, the label prediction model may be an NLP (Natural Language Processing) model. The NLP model may be used to implement sequence-to-sequence mapping, thereby obtaining a predicted label sequence corresponding to the input text.

[0075] In some embodiments, the label prediction model may be a Transformer model in NLP. The model has the ability to learn sequence-to-sequence mapping relationships, thereby enabling the text-level classification task described in the present disclosure to be implemented.

[0076] In some embodiments, the tag prediction model (hereinafter referred to as prediction model) may include T5 (Tansfer Text-to-Text Transformer, text-to-text transmission conversion model).

[0077] T5 transforms most natural language processing tasks into a model that uses sequence generation by using a sequence-to-sequence loss function design. For label prediction tasks, by fine-tuning the pre-trained T5 using some corpus with label sequence annotation information, T5 can be made suitable for label prediction tasks. This can simplify the label prediction task training process.

[0078] In the model training stage, the input of the prediction model may be the acquired text samples and the sample label sequence corresponding to the text samples, and then the prediction model may obtain the prediction label sequence corresponding to the input text.

[0079] It should be noted here that in order to clearly and briefly explain the embodiments, in the subsequent embodiments, the Nth predicted label in the predicted label sequence corresponding to the text sample is obtained based on the text sample and the first N-1 sample labels in the sample label sequence as an example.

[0080] For example, suppose the text sample sequence obtained after sentence segmentation and vectorization is {w 1 , w 2 , …w t}, and its corresponding sample label sequence is {s 1 ,s 2 ,…s m The present disclosure can input the two sequences into the prediction model for processing to obtain the prediction label sequence {p 1 , p 2 ,…p m}. Among them, when predicting the Nth label p n When the corresponding operation logic can be expressed as: L(p n )=M({w 1 , w 2 , …w t},{s 1 ,s 2 ,…s n-1}). L means p n The probability of predicting various possible labels. M represents the mapping function inside the model, that is, using the input text sample and the first n-1 labels in the sample label sequence, the probability of predicting the nth label as various possible labels is obtained. Then, the label corresponding to the maximum probability can be selected as p n The corresponding predicted label.

[0081] Therefore, when the Nth label of the prediction sequence is obtained, the label information before the label can be fully utilized to obtain accurate prediction results.

[0082] S306: Adjust the parameters of the label prediction model based on the sample label sequence and the predicted label sequence.

[0083] After obtaining the predicted label sequence corresponding to the sample text, the sample label sequence and the predicted label sequence can be input into the first loss function to obtain the first loss information. Then, the first loss information can be used to determine the model descent gradient and adjust the parameters of the prediction model by back propagation.

[0084] The first loss function may indicate the error between the predicted label sequence and the sample label sequence. For example, the first loss function may be a cross entropy loss function. The present disclosure does not specifically limit the type of the first loss function.

[0085] In some embodiments, the first loss function may summarize the errors between the labels at the same position in the predicted label sequence and the sample label sequence as the first loss information to update the model parameters.

[0086] In some embodiments, the prediction model can be trained using a text sample set. In this case, the first loss function can summarize the errors between the predicted label sequence and the sample label sequence corresponding to each text sample in the text sample set as the first loss information to update the model parameters.

[0087] In the technical solution disclosed in the present disclosure, the acquired text sample and sample label sequence can be used as input, and a label prediction model can be used to obtain the Nth predicted label in the predicted label sequence corresponding to the text sample based on the text sample and at least part of the sample labels in the first N-1 sample labels in the sample label sequence, so as to predict each predicted label in the predicted label sequence step by step; wherein the sample label sequence includes multiple sample labels with a hierarchical relationship, and the arrangement order of the multiple sample labels indicates the hierarchical relationship. Then, based on the sample label sequence and the predicted label sequence, the parameters of the label prediction model are adjusted.

[0088] Therefore, during the label prediction model training process, the model can learn the correspondence between text samples and each sample label, and when obtaining the Nth label, it can make full use of the label information before the label. Therefore, when using the model for label prediction, the hierarchical relationship between multiple labels can be combined to achieve better label prediction effects.

[0089] The present disclosure also proposes a text processing method. The execution steps of the method can be seen in S302-S306.

[0090] Among them, when executing S304, based on the attention mechanism, the weights corresponding to at least some of the sample labels in the first N-1 sample labels in the sample label sequence can be obtained, so that attention can be reasonably allocated to at least some of the sample labels according to the degree of influence of the at least some of the sample labels on the prediction of the Nth label, and then when predicting the Nth label, the information brought by the previous labels can be reasonably combined to improve the prediction accuracy. It should be noted that the steps of S302 and S306 are not repeated below.

[0091] See also Figure 4 , Figure 4 The figure is a flowchart of a label prediction method shown in an embodiment of the present disclosure.

[0092] like Figure 4 As shown, when S304 is executed, S402 to S404 may be executed.

[0093] In S402, based on the attention mechanism, weights corresponding to at least some of the first N-1 sample labels in the sample label sequence are obtained.

[0094] In some embodiments, when executing S402, the similarity between the Nth label in the sample label sequence and the first N-1 labels can be determined first (the similarity can be cosine similarity, Mahalanobis similarity, etc.), and then the similarity is normalized (for example, using a softmax (normalization) function to normalize) to obtain the weights corresponding to the first N-1 labels. It should be noted that the attention mechanism can be a single-head or multi-head attention mechanism.

[0095] It is understandable that in the model training process described in the present disclosure, the labels in the prediction label sequence correspond to the sample label sequence one by one. The similarity between the Nth sample label and the N-1 sample labels before it can be understood as the similarity between the Nth prediction label and the N-1 sample labels. Therefore, the weights corresponding to the N-1 sample labels can represent the influence of the N-1 sample labels on the prediction of the Nth prediction label.

[0096] S404: Based on the text sample, the partial sample labels, and the weights corresponding to the partial sample labels, obtain an Nth predicted label in a predicted label sequence corresponding to the text sample.

[0097] In some embodiments, when executing S404, the weighted sum of the partial sample labels and the weights corresponding to the partial sample labels can be performed to obtain a first sequence. The sequence can carry the information brought by the partial sample labels. Then it can be fused with the text sample, and based on the fusion result, the Nth predicted label is determined. In some embodiments, the fusion step can also be an attention mechanism to fuse the information brought by the partial sample labels and the information carried by the text sample to facilitate the correct prediction of the Nth predicted label.

[0098] It should be noted that when predicting the first predicted label, it may be impossible to combine the sample label information because it is the first label in the sequence. In this case, a starting character can be set before the first sample label to facilitate calculation.

[0099] See also Figure 5 , Figure 5 A schematic diagram of a label prediction model structure shown in an embodiment of the present disclosure.

[0100] The following combination Figure 5 The process of predicting the Nth prediction label using the prediction model is described.

[0101] like Figure 5 As shown, the prediction model may include an encoder 510 and a decoder 520. The decoder 520 includes an attention mechanism unit 521 and a fusion unit 522. In some embodiments, the encoder 510 may also include an attention mechanism unit. Since the present disclosure mainly relates to the improvement of the attention mechanism unit of the decoder, the encoder will not be described in detail.

[0102] The encoder 510 is used to encode a text sample sequence (ie, model input text) to obtain an encoding vector of the text sample.

[0103] The attention mechanism unit 521 is used to execute S402, determine the influence of the first N-1 labels of the sample label sequence on the Nth label, and determine the weights corresponding to the first N-1 labels respectively.

[0104] The unit can first determine the similarity between the Nth label in the sample label sequence and the first N-1 labels (the similarity can be cosine similarity, Mahalanobis similarity, etc.), and then normalize the similarity (for example, using a softmax (normalization) function to normalize) to obtain the weights corresponding to the first N-1 labels.

[0105] The fusion unit 522 is used to execute S404, first weighted summing the encoding vectors of the first N-1 labels according to the weight, and then fusing the encoding vectors of the input text samples using the attention mechanism to obtain a fused vector to achieve the aggregation of sample label information and text sample information. Then, the fused vector is used for classification processing to obtain the prediction result of the Nth predicted label pn.

[0106] Therefore, attention can be reasonably allocated to at least part of the sample labels according to the degree of influence of at least part of the sample labels on the prediction of the Nth label, and then when predicting the Nth label, the information brought by its previous labels can be reasonably combined to improve the prediction accuracy.

[0107] The present disclosure also proposes a text processing method. The execution steps of the method can be referred to S302-S306. Among them, the aforementioned S402-S404 can be executed when S304 is executed. It should be noted that the steps of S302-S306 and S402-S404 are not repeated below.

[0108] The following are some concepts.

[0109] Ancestor-descendant relationships are used to indicate the relationship between two labels that have a direct or indirect hierarchical relationship.

[0110] The ancestor tag of a tag refers to the direct or indirect parent tag of the tag.

[0111] See also Figure 1 , sports and football have a direct hierarchical relationship, that is, sports and football have an ancestor and descendant relationship. Sports is the ancestor label of football. News and football have an indirect hierarchical relationship through sports, that is, news and football have an ancestor and descendant relationship. News is the ancestor label of football. It should be noted that, first, the special characters representing the hierarchical relationship are determined according to the adjacent labels before it, so it can be considered that the status of the special characters is the same as the adjacent labels before it. That is, if label A is the ancestor label of label B, the special character after label A is also the ancestor label of label B. Among them, A and B represent any labels in the label sequence. Second, the labels before the special character indicating the end of the sequence are all ancestor labels of the special character indicating the end of the sequence.

[0112] In the above example, it is not difficult to find that when predicting the Nth label, the previous labels include both ancestral labels and non-ancestral labels of the label, and only the ancestral labels are beneficial to the predicted labels, while the non-ancestral labels will bring noise and affect the prediction of the Nth label.

[0113] See also Figure 1 , assuming that the multi-labels corresponding to the text sample A include topics, art, news, sports, football ( Figure 1 The dark label nodes in the middle). According to the hierarchical relationship between these labels, we can get the sample label sequence {topic,_, news, / , art,_, sports, / , football, EOS}. When predicting the label "football", all the previous labels will be combined. Obviously, the previous labels include both its ancestor labels and its non-ancestor labels. Therefore, the noise brought by the non-ancestor labels may affect the prediction of the label "football".

[0114] In order to solve this problem, a second loss function can be introduced during the training process; the second loss function is used to increase the weight corresponding to the ancestral label of the predicted label during the training phase of the label prediction model. Therefore, when the value of the second loss function operation approaches 0 during the training process, the prediction model can learn to increase the weight corresponding to the ancestral label of the predicted label when predicting the label, so that when the prediction model is used for label prediction, the weight corresponding to the ancestral label of the predicted label can be increased, and the weight of its non-ancestral label can be reduced, thereby increasing the influence of its ancestral label and reducing the noise caused by the non-ancestral label, thereby improving the accuracy of label prediction.

[0115] See also Figure 7 , Figure 7 A flowchart of a method for determining loss information according to an embodiment of the present disclosure is shown.

[0116] like Figure 7 As shown, before executing S306, S3501-S3503 can be executed.

[0117] In S3501, the corresponding relationship between each sample label and its ancestor label can be maintained according to the ancestor-descendant relationship between each sample label included in the sample label sequence.

[0118] It should be noted that in the model training process recorded in the present disclosure, the labels in the prediction label sequence correspond to the sample label sequence one by one. The ancestral label of the Nth sample label can be understood as the ancestral label of the Nth prediction label. In the present disclosure, according to the direct correspondence between each sample label and its ancestral label, it can be applied to determine the ancestral label of the prediction label. That is, the ancestral label of the prediction label can be understood as the ancestral label of the sample label in the sample label sequence that is at the same position as the prediction label.

[0119] In some embodiments, the above correspondence relationship can be maintained by maintaining a mask matrix M. The mask matrix can be a lower triangular matrix. Figure 6 , Figure 6 A schematic diagram of a mask matrix shown in an embodiment of the present disclosure.

[0120] like Figure 6As shown, the horizontal direction of the mask matrix M indicates the sample label sequence input by the encoder, and the vertical direction refers to the predicted label sequence output by the encoder. The diagonal elements correspond to each label in the predicted label sequence. The upper right corner elements of the matrix are all 0. In each row of elements in the matrix, the elements before the diagonal elements of the row indicate whether each label corresponding to the element is the ancestor label of the label corresponding to the diagonal element. If the element value is 1, it is an ancestor label, and if the element value is 0, it is a non-ancestor label. In this way, the correspondence between each sample label in the sample label sequence and its ancestor label can be maintained through the matrix M. Figure 6 If no character is marked at the position corresponding to an element, it means that the value of the element is 0.

[0121] For example, suppose the sample label sequence input to the model decoder is Figure 1 Shown are {features,_,news, / ,arts,_,sports, / ,football,EOS}. Figure 1 The ancestor-descendant relationship between the labels shown can be obtained as Figure 6 The corresponding relationship shown.

[0122] Among them, maintenance Figure 6 Take the O8 tag in the example, which is the ancestor tag corresponding to the "football" tag. Figure 1 It can be seen that the ancestor tags of "football" include "sports" and "news", which correspond to Figure 6 I3 and I7 in . Since the status of special characters is the same as that of their previous adjacent labels, it can be determined that I4 and I8 are also ancestor labels of O8. Figure 6 In the row of elements corresponding to O8, the positions corresponding to I3, I4, I7, and I8 are marked as 1, which completes the maintenance of the correspondence between O8 and its ancestor labels.

[0123] After the corresponding relationship is maintained, in step S3052 , based on the corresponding relationship, the weights corresponding to the ancestor labels of each predicted label in the predicted label sequence are screened out.

[0124] In some embodiments, after completing step S402 for each sample label in the sample sequence, a weight matrix T may be formed. The weight matrix T may indicate the weights of each sample label used when predicting each prediction label.

[0125] The weight matrix T can be a lower triangular matrix similar to the mask matrix M. The horizontal direction of the weight matrix T indicates the sample label sequence input by the encoder, and the vertical direction refers to the predicted label sequence output by the encoder. The diagonal elements correspond to each label in the predicted label sequence. The upper right corner elements of the matrix are all 0. In each row of elements in the matrix, the elements before the diagonal elements of the row indicate the weights of each label corresponding to the elements. When executing S3052, the weight matrix T can be dot-multiplied with the mask matrix M to filter out the weights corresponding to the ancestral labels of each predicted label in the predicted label sequence.

[0126] S3053, inputting the weights corresponding to the ancestral labels of the predicted labels into the second loss function to obtain second loss information.

[0127] In some embodiments, the second loss function may include at least two calculation steps. First, the difference between the sum of the weights of the ancestor labels of each predicted label and 1 may be calculated. Second, the difference between the sum of the weights of the ancestor labels of each predicted label and 1 may be added to obtain the second loss information.

[0128] For example, the second loss function may be L2=b*∑ i (1-∑ j∈C a i,j ). Where b is a preset parameter. It should be noted that b may be related to the structure of the label prediction model and the number of heads of the multi-head attention mechanism. i,j Indicates the weight of the jth ancestor label of the i-th predicted label. 1-∑ j∈C a i,j Indicates the difference between the sum of the weights of the ancestral labels of each predicted label and 1, ∑ i (1-∑ j∈C a ij ) indicates that the difference between the sum of the weights of the ancestral labels of each predicted label and 1 is added.

[0129] Therefore, by using the second loss information to train the prediction model during the training process, the second loss information can be gradually approached to 0, and the sum of the weights of the ancestral labels of each predicted label can approach 1, that is, the ability to increase the weight corresponding to the ancestral label of the predicted label can be increased. Therefore, when using the prediction model to predict labels, the weight corresponding to the ancestral label of the predicted label can be increased, and the weight of its non-ancestral label can be reduced, thereby increasing the influence of its ancestral label and reducing the noise caused by the non-ancestral label, thereby improving the accuracy of label prediction.

[0130] After obtaining the second loss information, when executing S306, S3062-S3066 can be executed.

[0131] In S3062, the sample label sequence and the predicted label sequence are input into a first loss function to obtain first loss information; the first loss function indicates the error between the predicted label sequence and the sample label sequence.

[0132] The first loss function may indicate the error between the predicted label sequence and the sample label sequence. For example, the first loss function may be a cross entropy loss function. The present disclosure does not specifically limit the type of the first loss function.

[0133] The first loss function can summarize the errors between the labels at the same position in the predicted label sequence and the sample label sequence as the first loss information to update the model parameters.

[0134] S3064: Obtain the second loss information.

[0135] S3066: Adjust the parameters of the label prediction model based on the first loss information and the second loss information.

[0136] In some embodiments, the weighted sum of the first loss information and the second loss information can be used as the total loss information of a round of training. The total loss information can then be used to determine the model descent gradient and adjust the parameters of the prediction model by back propagation.

[0137] For example, the total loss calculation formula can be: Ltotal = L1 + λL2. Among them, lambda can be a pre-defined hyperparameter. L1 represents the first loss information. L2 represents the second loss information. Therefore, by using the total loss calculation formula to update the model parameters, on the one hand, the model can learn the correspondence between text samples and each sample label, and when obtaining the Nth label, it can make full use of the label information before the label, so that when using the model for label prediction, the hierarchical relationship between multiple labels can be combined to achieve better label prediction effect. On the other hand, the second loss information can be gradually approached to 0, and the sum of the weights of the ancestral labels of each predicted label can be approached to 1, that is, the ability to increase the weight corresponding to the ancestral label of the predicted label can be increased, so that when using the prediction model for label prediction, the weight corresponding to the ancestral label of the predicted label can be increased, and the weight of its non-ancestral label can be reduced, thereby increasing the influence of its ancestral label and reducing the noise caused by the non-ancestral label, thereby improving the accuracy of label prediction.

[0138] The present disclosure also proposes a text processing method. The execution steps of the method can be referred to S302-S306. It should be noted that the steps S302 to S306 are not described repeatedly below.

[0139] Before executing S302 , a sample label sequence corresponding to the text sample may be obtained first.

[0140] See also Figure 8 , Figure 8 The present invention is a flow chart of a method for obtaining a sample tag sequence according to an embodiment of the present invention.

[0141] like Figure 8 As shown, the method for obtaining the sample tag sequence may include S802-S804.

[0142] Among them, S802, obtain a tag tree; the tag tree includes a plurality of tag nodes corresponding to the plurality of tags respectively; the hierarchical relationship between the plurality of tag nodes indicates the hierarchical relationship between the plurality of tags.

[0143] The description of the label tree and the hierarchical relationship can be referred to the above content and will not be described in detail here. It should be noted that in the present disclosure, labels and label nodes correspond to each other, so the hierarchical relationship between labels can be considered to be equivalent to the hierarchical relationship between label nodes.

[0144] S804, traversing the multiple label nodes level by level to obtain a target label node corresponding to the text sample to obtain the sample label sequence.

[0145] Multiple target tags corresponding to the text sample have been obtained in advance. In this example, it is necessary to determine the target tag nodes that match the multiple target tags in the tag nodes included in the tag tree to obtain a sample tag sequence that can characterize the tag hierarchical relationship.

[0146] In some embodiments, the target tag nodes matching the multiple target tags can be screened layer by layer in the order of the tag tree from top to bottom, and the sample tag sequence can be generated in the order in which the target tag nodes are screened. Thus, the arrangement order of the sample tag sequence can indicate the hierarchical relationship. In some embodiments, the level-by-level traversal can be implemented in a breadth-first traversal manner, thereby quickly and efficiently generating the sample tag sequence.

[0147] The present disclosure also proposes a text processing method. The steps of training the label prediction model in the method can be referred to S302-S306. The specific training process can refer to any of the above embodiments. It will not be described in detail here.

[0148] See also Fig. 9 , Fig. 9 A flowchart of a label prediction method shown in an embodiment of the present disclosure.

[0149] like Fig. 9 As shown, the method may include S902-S904.

[0150] Among them, S902, obtain the target text.

[0151] In different scenarios, different methods may be used to obtain the target text. The present disclosure does not particularly limit the specific method for obtaining the target text.

[0152] For example, in some scenarios, text feedback can be provided in the form of voice such as by phone. In this case, the target text can be obtained by recognizing the voice. For another example, in some scenarios, text feedback can be provided in the form of text such as email, text message, chat software, etc. In this case, the feedback text can be directly used as the target text. For another example, in some scenarios, the target text can be obtained from an image containing text content through text recognition (for example, OCR (Optical Character Recognition)).

[0153] S904: Input the target text into the trained label prediction model to obtain a predicted label sequence corresponding to the target text.

[0154] The label prediction model has been trained using the training method shown in any of the aforementioned embodiments. After the target text is input into the model, multiple labels (predicted label sequences) with hierarchical relationships corresponding to the target text can be predicted. When predicting the Nth label, the label information before the label can be fully utilized, so that when the model is used for label prediction, the hierarchical relationship between multiple labels can be combined to achieve a better label prediction effect.

[0155] Combine the following Figure 2 The implementation scheme is described in detail with reference to the application scenarios. It should be noted that the aforementioned application scenarios are only shown to facilitate understanding of the spirit and principle of the present disclosure, and the implementation scheme of the present disclosure is not limited in this respect. On the contrary, the implementation scheme of the present disclosure can be applied to any applicable scenario.

[0156] exist Figure 2 In the illustrated scenario, the training logic and the use logic of the news tag prediction model may be deployed in the server 202. It is understandable that the training logic and the use logic of the model may also be deployed in different servers according to business requirements.

[0157] The news tag prediction model (hereinafter referred to as the prediction model) can be used to predict news tags. The model can use a pre-trained T5 model. The model includes an encoder and a decoder. The model structure can be found in Figure 5 It should be noted that T5 may include multiple encoders and decoders.

[0158] The training logic may use some text samples annotated with sample label sequences to perform multiple rounds of fine-tuning on the parameters of the prediction model to obtain a converged trained prediction model.

[0159] See also Fig.10 , Fig.10 A flow chart of a model training method shown in an embodiment of the present disclosure.

[0160] like Fig.10 As shown, during a round of training, S1002 - S1010 may be included.

[0161] In S1002, a sample label sequence corresponding to the text sample is generated according to a pre-maintained news label tree. The arrangement order of the multiple sample labels in the sample label sequence indicates the hierarchical relationship.

[0162] S1004: Input the text sample and the corresponding sample label sequence into the prediction model to obtain a pre-label sequence corresponding to the text sample.

[0163] Among them, the text sample can be encoded by the encoder of the model to obtain a first encoding vector. When predicting the Nth predicted label, the decoder can be used to determine the weights of the first N-1 sample labels in the sample label sequence, and the weighted sum of the encoding vectors corresponding to the first N-1 sample labels is fused with the first encoding vector to obtain a fused vector. Then, classification is performed based on the fused vector to obtain the Nth predicted label.

[0164] S1006, the predicted label sequence and the sample label sequence may be input into a first loss function (cross entropy loss function) corresponding to T5 to obtain first loss information.

[0165] S1008, based on the maintained mask matrix M and weight matrix T, the ancestor label of each predicted label is screened out, and then the ancestor label of each predicted label is input into the second loss function to obtain second loss information.

[0166] The mask matrix M represents the correspondence between each label in the input sample label sequence and its ancestor label. The weight matrix T represents the weights corresponding to each label before the predicted label. The second loss function is L2=b*∑ i (1-∑ j∈C a i,j ). Where b is a preset parameter. It should be noted that b may be related to the structure of the label prediction model and the number of heads of the multi-head attention mechanism. i,j Indicates the weight of the jth ancestor label of the i-th predicted label. 1-∑ j∈C a i,j Indicates the difference between the sum of the weights of the ancestral labels of each predicted label and 1, ∑i (1-∑ j∈C a i,j ) indicates that the difference between the sum of the weights of the ancestral labels of each predicted label and 1 is added.

[0167] S1010, based on the weighted sum of the first loss information and the second loss information, determine the total loss, and use back propagation to complete the parameter adjustment of the prediction model. Wherein, Ltotal = L1 + λL2. Wherein, L can be a pre-defined hyperparameter. L1 represents the first loss information. L2 represents the second loss information. Thus, using the total loss calculation formula to update the model parameters can enable the model to learn the correspondence between text samples and each sample label. When the Nth label is obtained, it can make full use of the label information before the label, and can increase the weight corresponding to the ancestral label of the predicted label.

[0168] After the model training is completed, the prediction model can be deployed in the news label prediction device to perform label prediction by using logic.

[0169] The usage logic can obtain the target text for which label prediction is required. Then the target text can be input into the prediction model to obtain a predicted news label sequence corresponding to the target text. Thus, by using the prediction model for prediction, the hierarchical relationship between the predicted multiple labels can be combined to achieve a better label prediction effect, and the weight corresponding to the ancestor label of the predicted label can be increased, and the weight of its non-ancestor label can be reduced, thereby increasing the influence of its ancestor label and reducing the noise caused by the non-ancestor label, thereby improving the accuracy of label prediction.

[0170] Exemplary Devices

[0171] After introducing the method of the exemplary embodiment of the present disclosure, next, refer to Fig.11 The text processing device disclosed in the present disclosure is described as an example. The text processing device is used to implement the text processing method shown in any of the above embodiments.

[0172] See also Fig.11 , Fig.11 The present invention is a schematic diagram of the structure of a text processing device according to an embodiment of the present invention.

[0173] like Fig.11 As shown, the device 110 may include: a first acquisition module 111, used to acquire a text sample and a sample label sequence corresponding to the text sample; the sample label sequence includes a plurality of sample labels having a hierarchical relationship; the arrangement order of the plurality of sample labels indicates the hierarchical relationship;

[0174] The first prediction module 112 is used to take the text sample and the sample label sequence as input, and use the label prediction model to obtain the Nth predicted label in the predicted label sequence corresponding to the text sample based on the text sample and at least part of the sample labels in the first N-1 sample labels in the sample label sequence, so as to predict each predicted label in the predicted label sequence step by step; N represents the label sequence in the sample label sequence or the predicted label sequence;

[0175] The adjustment module 113 adjusts the parameters of the label prediction model based on the sample label sequence and the predicted label sequence.

[0176] In some embodiments, the first prediction module 112 is specifically configured to:

[0177] Based on the attention mechanism, obtaining weights corresponding to at least some of the first N-1 sample labels in the sample label sequence;

[0178] Based on the text sample, the partial sample labels, and the weights corresponding to the partial sample labels, an Nth predicted label in the predicted label sequence corresponding to the text sample is obtained.

[0179] In some embodiments, the device 110 further includes:

[0180] A second loss information determination module, configured to maintain a correspondence between each sample label and its ancestor label according to an ancestor-descendant relationship between each sample label included in the sample label sequence;

[0181] Based on the corresponding relationship, the weights corresponding to the ancestral labels of each predicted label in the predicted label sequence are screened out; the weights corresponding to the ancestral labels of each predicted label are input into the second loss function to obtain second loss information; the second loss function is used to increase the weights corresponding to the ancestral labels of the predicted labels during the training stage of the label prediction model.

[0182] In some embodiments, the adjustment module 113 is specifically used to:

[0183] Inputting the sample label sequence and the predicted label sequence into a first loss function to obtain first loss information; the first loss function indicates the error between the predicted label sequence and the sample label sequence;

[0184] Acquire the second loss information; and adjust the parameters of the label prediction model based on the first loss information and the second loss information.

[0185] In some embodiments, the device 110 further includes:

[0186] A sample label sequence acquisition module is used to acquire a label tree; the label tree includes a plurality of label nodes corresponding to the plurality of labels respectively; the hierarchical relationship between the plurality of label nodes indicates the hierarchical relationship between the plurality of labels;

[0187] The multiple label nodes are traversed level by level to obtain a target label node corresponding to the text sample to obtain the sample label sequence.

[0188] In some embodiments, the level-by-level traversal includes a breadth-first traversal.

[0189] In some embodiments, the label prediction model includes a conversion model.

[0190] In some embodiments, the conversion model comprises a text-to-text transfer conversion model.

[0191] In some embodiments, the device 110 further includes:

[0192] A second acquisition module 114, used to acquire a target text;

[0193] The second prediction module 115 is used to input the target text into the trained label prediction model to obtain a predicted label sequence corresponding to the target text.

[0194] Therefore, during the training process, on the one hand, the model can learn the correspondence between text samples and each sample label, and when the Nth label is obtained, it can make full use of the label information before the label, so that when the model is used for label prediction, the hierarchical relationship between multiple labels can be combined to achieve better label prediction results. On the other hand, the second loss information can be gradually approached to 0, and the sum of the weights of the ancestral labels of each predicted label can be approached to 1, that is, the ability to increase the weight corresponding to the ancestral label of the predicted label can be increased, so that when the prediction model is used for label prediction, the weight corresponding to the ancestral label of the predicted label can be increased, and the weight of its non-ancestral label can be reduced, thereby increasing the influence of its ancestral label and reducing the noise caused by non-ancestral labels, thereby improving the accuracy of label prediction.

[0195] Exemplary Media

[0196] After introducing the method and apparatus of the exemplary embodiments of the present disclosure, next, reference is made to Fig.12 A readable storage medium exemplarily disclosed in the present disclosure is described. The storage medium stores a computer program, and the computer program is used to enable a processor to execute a text processing method as shown in any of the above embodiments.

[0197] See also Fig.12 , Fig.12A program product 120 applied to a text processing method according to an embodiment of the present disclosure is shown.

[0198] In some embodiments shown, the aforementioned text processing method can be implemented by a program product 70, such as a portable compact disk read-only memory (CD-ROM) and including program code, and can be run on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto, and in this document, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device.

[0199] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0200] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0201] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RE, etc., or any suitable combination of the foregoing.

[0202] Program code for performing the disclosed operations may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and also conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user electronic device, partially on the user electronic device, partially on the remote electronic device, or entirely on the remote electronic device or server. In the case of a remote electronic device, the remote electronic device may be connected to the user electronic device through any type of model, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external electronic device (e.g., using an Internet service provider to connect through the Internet).

[0203] Exemplary Electronic Devices

[0204] After introducing the method, apparatus and medium of the exemplary embodiments of the present disclosure, next, reference is made to Fig.13 An electronic device disclosed in the present disclosure is described. The device includes: a processor; a memory for storing processor executable instructions; wherein the processor implements the text processing method shown in any of the above embodiments by running the executable instructions.

[0205] See also Fig.13 , Fig.13 The present invention is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.

[0206] Fig.13 The electronic device 1300 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0207] like Fig.13 As shown, the electronic device 1300 is in the form of a general electronic device. The components of the electronic device 1300 may include but are not limited to: the aforementioned at least one processor 1301, the aforementioned at least one storage processor 1302, and a bus 1303 connecting different system components (including the processor 1301 and the storage processor 1302).

[0208] The bus 1303 includes a data bus, a control bus, and an address bus.

[0209] The storage processor 1302 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 13021 and / or a cache memory 13022 , and may further include a readable medium in the form of a non-volatile memory, such as a read-only memory (ROM) 13023 .

[0210] The storage processor 1302 may also include a program / utility 13025 having a set (at least one) of program modules 13024, such program modules 13024 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a model environment.

[0211] The electronic device 1300 may also communicate with one or more external devices 1304 (eg, keyboards, pointing devices, etc.).

[0212] Such communication may be performed via an input / output (I / O) interface 1305. Furthermore, the electronic device 1300 may also communicate with one or more models (e.g., a local area network (LAN), a wide area network (WAN), and / or a public model, such as the Internet) via a model adapter 1306. Fig.13 As shown, the model adapter 1306 communicates with other modules of the electronic device 1300 via the bus 1303. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 1300, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0213] It should be noted that although several units / modules or sub-units / modules of the text processing device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided to be embodied by multiple units / modules.

[0214] In addition, although the operations of the disclosed method are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0215] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.

Claims

1. A text processing method, comprising: Acquire a text sample and a sample label sequence corresponding to the text sample, wherein the sample label sequence includes a plurality of sample labels having a hierarchical relationship, and an arrangement order of the plurality of sample labels indicates the hierarchical relationship; Taking the text sample and the sample label sequence as input, using a label prediction model, based on the text sample and at least part of the sample labels in the first N-1 sample labels in the sample label sequence, obtaining the probability that the Nth predicted label in the predicted label sequence corresponding to the text sample is predicted as each label, and taking the label corresponding to the maximum probability as the Nth predicted label, so as to predict each predicted label in the predicted label sequence step by step; N represents the label order in the sample label sequence or the predicted label sequence; Based on the sample label sequence and the predicted label sequence, the parameters of the label prediction model are adjusted.

2. The method according to claim 1, wherein obtaining the Nth predicted label in the predicted label sequence corresponding to the text sample based on the text sample and at least part of the sample labels in the first N-1 sample labels in the sample label sequence comprises: Based on the attention mechanism, obtaining weights corresponding to at least some of the first N-1 sample labels in the sample label sequence; Based on the text sample, the partial sample labels, and the weights corresponding to the partial sample labels, an Nth predicted label in the predicted label sequence corresponding to the text sample is obtained.

3. The method according to claim 2, before adjusting the parameters of the label prediction model based on the sample label sequence and the predicted label sequence, further comprising: Maintaining a correspondence between each sample label and its ancestor label according to an ancestor-descendant relationship between each sample label included in the sample label sequence; Based on the corresponding relationship, the weights corresponding to the ancestor labels of each predicted label in the predicted label sequence are screened out; The weights corresponding to the ancestral labels of the predicted labels are respectively input into the second loss function to obtain second loss information; the second loss function is used to increase the weights corresponding to the ancestral labels of the predicted labels during the training phase of the label prediction model.

4. The method according to claim 3, wherein adjusting the parameters of the label prediction model based on the sample label sequence and the predicted label sequence comprises: Inputting the sample label sequence and the predicted label sequence into a first loss function to obtain first loss information; The first loss function indicates an error between the predicted label sequence and the sample label sequence; Acquire the second loss information; and adjust the parameters of the label prediction model based on the first loss information and the second loss information.

5. According to any one of the methods of claims 1 to 4, the method for obtaining the sample tag sequence comprises: Get the tag tree; The label tree includes a plurality of label nodes corresponding to the plurality of labels respectively; The hierarchical relationship between the multiple label nodes indicates the hierarchical relationship between the multiple labels; The multiple label nodes are traversed level by level to obtain a target label node corresponding to the text sample to obtain the sample label sequence. The method according to claim 5 , wherein the level-by-level traversal comprises a breadth-first traversal.

7. The method according to claim 1, wherein the label prediction model comprises a conversion model.

8. The method of claim 7, wherein the conversion model comprises a text-to-text transfer conversion model.

9. The method according to claim 1, further comprising: Get the target text; The target text is input into the trained label prediction model to obtain a predicted label sequence corresponding to the target text.

10. A text processing device, comprising: An acquisition module, used to acquire a text sample and a sample label sequence corresponding to the text sample; the sample label sequence includes a plurality of sample labels having a hierarchical relationship; The arrangement order of the multiple sample labels indicates the hierarchical relationship; A first prediction module is used to take the text sample and the sample label sequence as input, and use a label prediction model to obtain the probability of the Nth predicted label in the predicted label sequence corresponding to the text sample being predicted as each label based on the text sample and at least part of the sample labels in the first N-1 sample labels in the sample label sequence, and use the label corresponding to the maximum probability as the Nth predicted label to predict each predicted label in the predicted label sequence step by step; N represents the label sequence in the sample label sequence or the predicted label sequence; The adjustment module adjusts the parameters of the label prediction model based on the sample label sequence and the predicted label sequence.

11. The device according to claim 10, wherein the first prediction module is specifically configured to: Based on the attention mechanism, obtaining weights corresponding to at least some of the first N-1 sample labels in the sample label sequence; Based on the text sample, the partial sample labels, and the weights corresponding to the partial sample labels, an Nth predicted label in the predicted label sequence corresponding to the text sample is obtained.

12. The device according to claim 11, further comprising: A second loss information determination module, configured to maintain a correspondence between each sample label and its ancestor label according to an ancestor-descendant relationship between each sample label included in the sample label sequence; Based on the corresponding relationship, the weights corresponding to the ancestor labels of each predicted label in the predicted label sequence are screened out; The weights corresponding to the ancestral labels of the predicted labels are respectively input into the second loss function to obtain second loss information; the second loss function is used to increase the weights corresponding to the ancestral labels of the predicted labels during the training phase of the label prediction model.

13. The device according to claim 12, wherein the adjustment module is specifically configured to: Inputting the sample label sequence and the predicted label sequence into a first loss function to obtain first loss information; the first loss function indicates the error between the predicted label sequence and the sample label sequence; Acquire the second loss information; and adjust the parameters of the label prediction model based on the first loss information and the second loss information.

14. The device according to any one of claims 10 to 13, further comprising: A sample label sequence acquisition module is used to acquire a label tree; the label tree includes a plurality of label nodes corresponding to the plurality of labels respectively; The hierarchical relationship between the multiple label nodes indicates the hierarchical relationship between the multiple labels; The multiple label nodes are traversed level by level to obtain a target label node corresponding to the text sample to obtain the sample label sequence. The apparatus according to claim 14 , wherein the level-by-level traversal comprises a breadth-first traversal.

16. The apparatus according to claim 10, wherein the label prediction model comprises a conversion model.

17. The apparatus of claim 16, wherein the conversion model comprises a text-to-text transfer conversion model.

18. The apparatus according to claim 10, further comprising: The second acquisition module is used to acquire the target text; The second prediction module is used to input the target text into the trained label prediction model to obtain a predicted label sequence corresponding to the target text.

19. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor implements the text processing method as described in any one of claims 1-9 by running the executable instructions.

20. A computer-readable storage medium storing a computer program, wherein the computer program is used to enable a processor to execute the text processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and device for automatically verifying commodity categories

    CN111353838A

  • Hierarchical classification method and device, electronic equipment and storage medium

    CN112528658A