Model construction method, classification method, device, storage medium and electronic device

By constructing a model construction method, using the data to be trained and the preset relationship dependency matrix to be trained, the first-level label and second-level label prediction matrix are generated, which solves the problem of low prediction accuracy in the existing technology and realizes more efficient information flow data push.

CN114175017BActive Publication Date: 2025-06-10SHENZHEN HEYTAP TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980098836.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-30
Publication Date
2025-06-10
Estimated Expiration
2039-10-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the accuracy of joint prediction of primary and secondary tags, especially in the push of information flow data, the hierarchical relationship of the tags is complex, which affects the accuracy of the recommendation algorithm.

Method used

By constructing a model construction method, obtain the data to be trained and the preset relationship dependency matrix, input the model to be trained, generate a first-level label prediction matrix and a second-level label prediction matrix, and combine the target relationship dependency matrix and loss value for model training until the model converges, and improve the prediction accuracy.

Benefits of technology

It significantly improves the accuracy of joint prediction of primary and secondary tags, and improves the accuracy and user experience of information flow data push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175017B_ABST
    Figure CN114175017B_ABST
Patent Text Reader

Abstract

A model construction method, classification method, device, storage medium, and electronic device. The model construction method includes: determining a target loss value based on a first-level label prediction matrix, a first-level label reference matrix, a target matrix, and a second-level label reference matrix; obtaining a target loss value every time a batch of data to be trained is completed, and every time a target loss value is obtained, backpropagating the target loss value into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic technology, and particularly relates to a model construction method, a classification method, a device, a storage medium, and an electronic device. Background Art

[0002] With the continuous development of electronic technology, the number of application programs (APPs) installed in electronic devices such as smart phones or tablet computers is increasing. During the process of users using the above APPs, information flow services are often involved. An information flow service refers to a service that pushes information flow data to the above electronic devices. Among them, the information flow data is used to form pages such as a home page, a list page, or a content page.

[0003] Taking the information flow data as an article as an example, before pushing the article to an electronic device, it is necessary to mark the article with a first-level label, a second-level label, a third-level label, etc. for use by the information flow recommendation algorithm, so as to push different articles to different electronic devices. Summary of the Invention

[0004] The embodiments of this application provide a model construction method, a classification method, a device, a storage medium, and an electronic device, which can improve the accuracy of jointly predicting the first-level label and the second-level label.

[0005] In a first aspect, the embodiments of this application provide a model construction method, including:

[0006] Obtain training data to be trained, where the training data to be trained includes a token matrix composed of encodings corresponding to tokens of each text in multiple texts, a first-level label reference matrix composed of encodings corresponding to the first-level labels of each text, and a second-level label reference matrix composed of encodings corresponding to the second-level labels of each text;

[0007] Obtain a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between the first-level label and the second-level label;

[0008] Input the training data to be trained and the preset relationship dependency matrix into a model to be trained to obtain a first-level label prediction matrix and a second-level label prediction matrix;

[0009] Determine a target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix;

[0010] Determine a target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix, where the target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix;

[0011] Whenever a batch of data to be trained is trained and a target loss value is obtained, each time a target loss value is obtained, the target loss value is fed back into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirming that the model training is completed, and obtaining the trained model.

[0012] In a second aspect, an embodiment of the present application provides a classification method, including:

[0013] Obtain the text to be classified;

[0014] Input the text to be classified into the trained model to obtain a first-level label probability matrix and a second-level label prediction probability matrix. Each element in the first-level label probability matrix corresponds to a first-level label, and all elements in the first-level label probability matrix are real numbers. Each element in the second-level label prediction probability matrix corresponds to a second-level label, and all elements in the second-level label prediction probability matrix are real numbers;

[0015] According to the first-level label probability matrix, determine the first-level label corresponding to the text to be classified. The first-level label corresponding to the element with the largest value in the first-level label probability matrix is the first-level label corresponding to the text to be classified;

[0016] Perform integerization processing on the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, obtaining a first-level label integerization matrix, and the value of the element in the first-level label integerization matrix is 0 or 1;

[0017] According to the first-level label integerization matrix and a preset relationship dependency matrix, determine a first relationship dependency matrix;

[0018] According to the first relationship dependency matrix and the second-level label prediction probability matrix, determine a second-level label probability matrix. Each element in the second-level label probability matrix corresponds to a second-level label;

[0019] According to the second-level label probability matrix, determine the second-level label corresponding to the text to be classified. The second-level label corresponding to the element with the largest value in the second-level label probability matrix is the second-level label corresponding to the text to be classified.

[0020] In a third aspect, an embodiment of the present application provides a model construction device, including:

[0021] A first acquisition module, configured to acquire data to be trained, where the data to be trained includes a token matrix composed of encodings corresponding to tokens of each text in multiple texts, a first-level label reference matrix composed of encodings corresponding to the first-level labels of each text, and a second-level label reference matrix composed of encodings corresponding to the second-level labels of each text;

[0022] A second acquisition module, configured to acquire a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between first-level tags and second-level tags;

[0023] A first training module, configured to input the data to be trained and the preset relationship dependency matrix into a model to be trained, so as to obtain a first-level tag prediction matrix and a second-level tag prediction matrix;

[0024] A first determination module, configured to determine a target relationship dependency matrix according to the first-level tag prediction matrix and the preset relationship dependency matrix;

[0025] A second determination module, configured to determine a target loss value according to the first-level tag prediction matrix, the first-level tag reference matrix, a target matrix, and the second-level tag reference matrix, where the target matrix is determined according to the second-level tag prediction matrix and the target relationship dependency matrix;

[0026] A second training module, configured to obtain a target loss value every time a batch of data to be trained is trained. Every time a target loss value is obtained, the target loss value is fed back into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirm that the model training is completed, and obtain a trained model.

[0027] In a fourth aspect, an embodiment of the present application provides a classification device, including:

[0028] A third acquisition module, configured to acquire a text to be classified;

[0029] A prediction module, configured to input the text to be classified into the trained model to obtain a first-level tag probability matrix and a second-level tag prediction probability matrix, where each element in the first-level tag probability matrix corresponds to a first-level tag, and all elements in the first-level tag probability matrix are real numbers; each element in the second-level tag prediction probability matrix corresponds to a second-level tag, and all elements in the second-level tag prediction probability matrix are real numbers;

[0030] A third determination module, configured to determine the first-level tag corresponding to the text to be classified according to the first-level tag probability matrix, where the first-level tag corresponding to the element with the largest value in the first-level tag probability matrix is the first-level tag corresponding to the text to be classified;

[0031] An integer conversion module, configured to perform integer conversion on the first-level tag probability matrix so that each element in the first-level tag probability matrix changes from a real number to an integer, obtaining a first-level tag integer matrix, where the value of an element in the first-level tag integer matrix is 0 or 1;

[0032] A fourth determination module, configured to determine a first relationship dependency matrix according to the first-level tag integer matrix and the preset relationship dependency matrix;

[0033] A fifth determination module, configured to determine a secondary label probability matrix according to the first relationship dependency matrix and the secondary label prediction probability matrix, where each element in the secondary label probability matrix corresponds to a secondary label;

[0034] A sixth determination module, configured to determine a secondary label corresponding to the text to be classified according to the secondary label probability matrix, where the secondary label corresponding to the element with the largest value in the secondary label probability matrix is the secondary label corresponding to the text to be classified.

[0035] In a fifth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is enabled to execute the model construction method or classification method provided in this embodiment.

[0036] In a sixth aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute by calling the computer program stored in the memory:

[0037] Obtain training data to be trained, where the training data to be trained includes a token matrix composed of encodings corresponding to tokens of each text in multiple texts, a primary label reference matrix composed of encodings corresponding to primary labels of each text, and a secondary label reference matrix composed of encodings corresponding to secondary labels of each text;

[0038] Obtain a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between primary labels and secondary labels;

[0039] Input the training data to be trained and the preset relationship dependency matrix into a model to be trained, so as to obtain a primary label prediction matrix and a secondary label prediction matrix;

[0040] Determine a target relationship dependency matrix according to the primary label prediction matrix and the preset relationship dependency matrix;

[0041] Determine a target loss value according to the primary label prediction matrix, the primary label reference matrix, a target matrix, and the secondary label reference matrix, where the target matrix is determined according to the secondary label prediction matrix and the target relationship dependency matrix;

[0042] Whenever a batch of training data to be trained is trained to obtain a target loss value, each time a target loss value is obtained, the target loss value is fed back into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirm that the model training is completed, and obtain the trained model.

[0043] Seventh aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the following by calling the computer program stored in the memory:

[0044] Obtain the text to be classified;

[0045] Input the text to be classified into the trained model to obtain a first-level label probability matrix and a second-level label prediction probability matrix. Each element in the first-level label probability matrix corresponds to a first-level label, and all elements in the first-level label probability matrix are real numbers. Each element in the second-level label prediction probability matrix corresponds to a second-level label, and all elements in the second-level label prediction probability matrix are real numbers;

[0046] Determine the first-level label corresponding to the text to be classified according to the first-level label probability matrix. The first-level label corresponding to the element with the largest value in the first-level label probability matrix is the first-level label corresponding to the text to be classified;

[0047] Perform integerization processing on the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, obtaining a first-level label integerization matrix. The value of the element in the first-level label integerization matrix is 0 or 1;

[0048] Determine a first relationship dependency matrix according to the first-level label integerization matrix and a preset relationship dependency matrix;

[0049] Determine a second-level label probability matrix according to the first relationship dependency matrix and the second-level label prediction probability matrix. Each element in the second-level label probability matrix corresponds to a second-level label;

[0050] Determine the second-level label corresponding to the text to be classified according to the second-level label probability matrix. The second-level label corresponding to the element with the largest value in the second-level label probability matrix is the second-level label corresponding to the text to be classified. Description of the Drawings

[0051] The following will clearly show the technical solutions and their beneficial effects of the present application by describing the specific embodiments of the present application in detail with reference to the accompanying drawings.

[0052] Figure 1 is the first flowchart of the model construction method provided by the embodiment of the present application.

[0053] Figure 2 is the second flowchart of the model construction method provided by the embodiment of the present application.

[0054] Figure 3 is a schematic diagram of the preset relationship dependency matrix M0 provided by the embodiment of the present application.

[0055] Figure 4 It is a schematic diagram of the first-level label reference matrix y1 provided by an embodiment of the present application.

[0056] Figure 5 It is a schematic diagram of the second-level label reference matrix y2 provided by an embodiment of the present application.

[0057] Figure 6 It is a schematic diagram of the first-level label prediction matrix P1 provided by an embodiment of the present application.

[0058] Figure 7 It is a schematic diagram of the second-level label prediction matrix P2 provided by an embodiment of the present application.

[0059] Figure 8 It is a schematic diagram of the 0-1 integer matrix P1-1 provided by an embodiment of the present application.

[0060] Figure 9 It is a schematic diagram of the target relationship dependency matrix M provided by an embodiment of the present application.

[0061] Figure 10 It is a schematic diagram of the dictionary provided by an embodiment of the present application.

[0062] Figure 11 It is a schematic diagram of the process of the classification method provided by an embodiment of the present application.

[0063] Figure 12 It is a schematic diagram of the scenario of the classification method provided by an embodiment of the present application.

[0064] Figure 13 It is a schematic diagram of the structure of the model construction device provided by an embodiment of the present application.

[0065] Figure 14 It is a schematic diagram of the structure of the classification device provided by an embodiment of the present application.

[0066] Figure 15 It is a schematic diagram of the first structure of the electronic device provided by an embodiment of the present application.

[0067] Figure 16 It is a schematic diagram of the second structure of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0068] Please refer to the drawings, where the same component symbols represent the same components. The principle of the present application is illustrated by being implemented in a suitable computing environment. The following description is based on the specific embodiments of the present application illustrated, and it should not be regarded as limiting other specific embodiments of the present application not described in detail herein.

[0069] Please refer to Figure 1 , Figure 1It is the first process schematic diagram of the model construction method provided by the embodiments of the present application. The process of the model construction method may include:

[0070] 101. Obtain the data to be trained, where the data to be trained includes a token matrix composed of the encodings corresponding to the tokens of each text in multiple texts, a primary label reference matrix composed of the encodings corresponding to the primary labels of each text, and a secondary label reference matrix composed of the encodings corresponding to the secondary labels of each text.

[0071] For example, first, the electronic device can obtain multiple texts, and each text in the multiple texts is marked with a primary label and a secondary label. Then, the electronic device can tokenize each text and determine the encoding corresponding to the token of each text. Finally, the electronic device can form a token matrix according to the encoding corresponding to the token of each text. Among them, the first dimension (row) of the token matrix represents each text, and the second dimension (column) represents the encoding corresponding to the token of each text. For example, the element in the i-th row and j-th column of the token matrix represents the encoding corresponding to the j-th token of the i-th text.

[0072] Subsequently, the electronic device can encode the primary label corresponding to each text and then form a primary label reference matrix. Among them, the first dimension (row) of the primary label reference matrix represents each text, and the second dimension (column) represents the encoding corresponding to the primary label of each text. For example, the i-th row of the primary label reference matrix represents the encoding corresponding to the primary label of the i-th text. The electronic device can encode the secondary label corresponding to each text and then form a secondary label reference matrix. Among them, the first dimension (row) of the secondary label reference matrix is each text, and the second dimension (column) is the encoding corresponding to the secondary label of each text. For example, the i-th row of the secondary label reference matrix represents the encoding corresponding to the secondary label of the i-th text.

[0073] The token matrix, the primary label reference matrix, and the secondary label reference matrix constitute the data to be trained.

[0074] 102. Obtain a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between the primary label and the secondary label.

[0075] For example, multiple first-level tags and multiple second-level tags can be obtained from a database or pre-collected by a user, and the hierarchical relationship between each first-level tag and each second-level tag can be determined. That is, determine which second-level tags correspond to each first-level tag. For example, assume the first-level tag is: TV drama, and the second-level tags under it can be: ancient costume, fantasy, modern, etc. Another example, assume the first-level tag is: sports, and the second-level tags under it can be: football, basketball, volleyball, table tennis, etc. Then, the electronic device can establish a preset relationship dependency matrix according to the hierarchical relationship between the first-level tags and the second-level tags. Among them, the first dimension (row) of the preset relationship dependency matrix represents each first-level tag, and the second dimension (column) represents each second-level tag. If the j-th second-level tag is included under the i-th first-level tag, then the element in the i-th row and j-th column of the preset relationship dependency matrix is 1, otherwise it is 0.

[0076] It should be noted that in the embodiments of the present application, after the preset relationship dependency matrix is established, when the electronic device needs to use the preset relationship dependency matrix, it can directly obtain the preset relationship dependency matrix for use, without having to establish the preset relationship dependency matrix again before using. That is, establish once and use multiple times.

[0077] In the embodiments of the present application, the electronic device can obtain the preset relationship dependency matrix.

[0078] It should be noted that when obtaining multiple texts, the user can also obtain multiple texts according to the multiple first-level tags and multiple second-level tags collected, and input them into the electronic device, and the electronic device will then obtain the multiple texts. That is, among the multiple texts obtained by the electronic device, the first-level tag corresponding to each text is one of the multiple first-level tags collected by the user; the second-level tag corresponding to each text is one of the multiple second-level tags collected by the user.

[0079] It can be understood that to improve the accuracy of model prediction, the first-level tag and the second-level tag corresponding to the text obtained by the electronic device can both fully represent the text. For example, the artificial annotation method can be used to label the first-level tag and the second-level tag for the text. Then, the electronic device can obtain these texts with artificially labeled tags.

[0080] 103. Input the data to be trained and the preset relationship dependency matrix into the model to be trained to obtain a first-level tag prediction matrix and a second-level tag prediction matrix.

[0081] For example, after obtaining the data to be trained and the preset relationship dependency matrix, the electronic device can input the data to be trained and the preset relationship dependency matrix into the model to be trained to obtain a first-level tag prediction matrix and a second-level tag prediction matrix.

[0082] Among them, in the first-level label prediction matrix, the number of rows of the matrix is the number of texts included in the data to be trained, and the number of columns of the matrix is the number of first-level labels. For example, if the number of texts is 64 and the number of first-level labels is 10, then the first-level label prediction matrix is a 64×10 matrix. The first dimension (row) of the first-level label prediction matrix represents each text, and the second dimension (column) represents the probability that the first-level label corresponding to each text predicted by the model to be trained is each of the multiple first-level labels. For example, the element in the i-th row and j-th column of the first-level label prediction matrix represents the probability that the first-level label corresponding to the i-th text is the j-th first-level label among the multiple first-level labels.

[0083] In the second-level label prediction matrix, the number of rows of the matrix is the number of texts included in the data to be trained, and the number of columns of the matrix is the number of second-level labels. For example, if the number of texts is 64 and the number of second-level labels is 200, then the second-level label prediction matrix is a 64×200 matrix. The first dimension (row) of the second-level label prediction matrix represents each text, and the second dimension (column) represents the probability that the second-level label corresponding to each text predicted by the model to be trained is each of the multiple second-level labels. For example, the element in the i-th row and j-th column of the second-level label prediction matrix represents the probability that the second-level label corresponding to the i-th text is the j-th second-level label among the multiple second-level labels.

[0084] In the embodiments of the present application, the model to be trained can first perform parameter initialization. Then, the electronic device can input the data to be trained and the preset relationship dependency matrix into the model to be trained, and obtain an output result through forward propagation of each layer such as the convolutional layer, downsampling layer, and fully connected layer, that is, obtain the first-level label prediction matrix and the second-level label prediction matrix.

[0085] 104. Determine the target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix.

[0086] 105. Determine the target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix, where the target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix.

[0087] For example, after obtaining the first-level label prediction matrix, the electronic device can determine the target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix. After obtaining the target relationship dependency matrix, the electronic device can determine the target matrix according to the target relationship dependency matrix and the second-level label prediction matrix. Subsequently, the electronic device can determine the target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix.

[0088] 106. Whenever a batch of data to be trained is completed, a target loss value is obtained. Whenever a target loss value is obtained, the target loss value is passed back to the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirming that the model training is completed, and obtaining the trained model.

[0089] It can be understood that after a target loss value is obtained, the target loss value can be passed back to each layer of the model to be trained, so as to adjust the parameters of the model to be trained. Subsequently, the electronic device can continue to obtain the data to be trained and input it into the model to be trained with adjusted parameters to continue training the model to be trained. Among them, the data to be trained and the data to be trained obtained in process 101 are two different batches of data, and the acquisition process of the data to be trained can refer to the acquisition process of process 101. After a target loss value is obtained this time, the target loss value can still be passed back to the model to be trained to adjust the parameters again until the model to be trained converges, confirming that the model training is completed, and obtaining the trained model. Among them, when the target loss value gradually approaches a certain value, or fluctuates around a certain value, and the loss change is less than a very small positive number, it can be confirmed that the model to be trained converges.

[0090] It can be understood that in this embodiment, the target relationship dependency matrix determined by using the preset relationship dependency matrix and the first-level label prediction matrix can perform enhanced training on the second-level labels, and the target loss value determined according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix can make the model to be trained reach the overall optimum, so as to improve the accuracy of the trained model for jointly predicting the first-level labels and the second-level labels.

[0091] Please refer to Figure 2 , Figure 2 which is the second process schematic diagram of the model construction method provided by the embodiment of the present application. The model construction method may include:

[0092] 201. The electronic device obtains a plurality of first-level labels.

[0093] 202. The electronic device obtains a plurality of second-level labels.

[0094] 203. The electronic device determines the hierarchical relationship between each first-level label and each second-level label.

[0095] 204. The electronic device establishes a preset relationship dependency matrix according to the hierarchical relationship.

[0096] For example, 201, 202, 203, and 204 may be:

[0097] A user can pre-collect multiple first-level tags and multiple second-level tags in advance. Then, the user can input the multiple first-level tags and multiple second-level tags collected into an electronic device, and the electronic device can obtain the multiple first-level tags and multiple second-level tags.

[0098] Next, the electronic device can determine the hierarchical relationship between the multiple first-level tags and second-level tags obtained, that is, which second-level tags are respectively under each first-level tag. It can be understood that this process can be analyzed and classified by the user. For example, the user can determine which second-level tags are under each first-level tag, and then mark the second-level tags with the marks of their respective first-level tags. The user can input these first-level tags and second-level tags marked with their respective first-level tags into the electronic device, and the electronic device can determine which second-level tag belongs to which first-level tag according to the marks on the second-level tags.

[0099] After determining the hierarchical relationship between each first-level tag and each second-level tag, the electronic device can establish a preset relationship dependency matrix according to the hierarchical relationship between each first-level tag and each second-level tag. Among them, the first dimension (row) of the preset relationship dependency matrix represents each first-level tag among the multiple first-level tags, and the second dimension (column) represents each second-level tag among the multiple second-level tags. If the i-th first-level tag contains the j-th second-level tag, then the element in the i-th row and j-th column of the preset relationship dependency matrix is 1, otherwise it is 0.

[0100] For example, assume there are 5 first-level tags, namely: L1, L2, L3, L4, L5, and 15 second-level tags, namely: S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, S11, S12, S13, S14, S15. Among them, the first-level tag L1 contains the second-level tags S1, S2, S4; the first-level tag L2 contains the second-level tags S3, S5; the first-level tag L3 contains the second-level tags S5, S6, S7; the first-level tag L4 contains the second-level tags S9, S11, S12, S15; the first-level tag L5 contains the second-level tags S10, S13, S14. Then, according to the 5 first-level tags L1, L2, L3, L4, L5 and 15 second-level tags S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, S11, S12, S13, S14, S15, the elements in the 0th row of the established preset relationship dependency matrix M0 at the 0th, 1st, and 3rd columns are 1, and the others are 0; the elements in the 1st row at the 2nd and 4th columns are 1, and the others are 0; the elements in the 2nd row at the 5th, 6th, and 7th columns are 1, and the others are 0; the elements in the 3rd row at the 8th, 10th, 11th, and 14th columns are 1, and the others are 0; the elements in the 4th row at the 9th, 12th, and 13th columns are 1, and the others are 0; that is, the preset relationship dependency matrix M0 is as Figure 3 shown.

[0101] 205. The electronic device obtains multiple texts, the corresponding first-level tags of each text, and the corresponding second-level tags of each text.

[0102] 206. The electronic device performs word segmentation processing on each text to obtain the word segmentation corresponding to each text.

[0103] 207. The electronic device determines the encoding corresponding to the word segmentation of each text.

[0104] 208. The electronic device determines a word segmentation matrix according to the encoding corresponding to the word segmentation of each text.

[0105] 209. The electronic device performs one-hot encoding processing on the first-level tags corresponding to each text to obtain the encoding corresponding to the first-level tags of each text.

[0106] 210. The electronic device determines a first-level tag reference matrix according to the encoding corresponding to the first-level tags of each text.

[0107] 211. The electronic device performs one-hot encoding processing on the second-level tags corresponding to each text to obtain the encoding corresponding to the second-level tags of each text.

[0108] 212. The electronic device determines a second-level tag reference matrix according to the encoding corresponding to the second-level tags of each text.

[0109] 213. The electronic device determines the data to be trained according to the word segmentation matrix, the first-level tag reference matrix, and the second-level tag reference matrix.

[0110] For example, 205 to 213 may be:

[0111] The user can collect multiple texts marked with first-level tags and second-level tags according to multiple first-level tags and multiple second-level tags collected by the user. Among them, the first-level tag marked by each text is one of the multiple first-level tags collected by the user; the second-level tag marked by each text is one of the multiple second-level tags collected by the user. The user can input the multiple texts marked with first-level tags and second-level tags collected into the electronic device, and the electronic device then obtains multiple texts, the corresponding first-level tags of each text, that is, the first-level tags marked by each text, and the corresponding second-level tags of each text, that is, the second-level tags marked by each text.

[0112] In some embodiments, in order to improve the accuracy of the prediction of the trained model, among the multiple texts marked with first-level labels and second-level labels, the second-level label marked for each text belongs to the first-level label marked for each text. That is to say, the second-level label marked for each text is included under the first-level label marked for each text. That is, if the second-level label marked for a certain text does not belong to the first-level label marked for that text, that text may not be collected when collecting texts.

[0113] Subsequently, the electronic device can perform word segmentation processing on each text in the multiple texts to obtain the word segmentation corresponding to each text. For example, the electronic device can use the jieba word segmentation tool to perform word segmentation processing on each text.

[0114] In some embodiments, the electronic device can filter out invalid and infrequently used special characters in each text, and then use the jieba word segmentation tool to perform word segmentation processing on each text.

[0115] After obtaining the word segmentation corresponding to each text, the electronic device can determine the encoding corresponding to the word segmentation of each text. Then, the electronic device can determine a word segmentation matrix according to the encoding corresponding to the word segmentation of each text. Among them, the first dimension (row) of the word segmentation matrix represents each text in the multiple texts, and the second dimension (column) represents the encoding corresponding to the word segmentation of each text. For example, the j-th column of the i-th row of the word segmentation matrix represents the encoding corresponding to the j-th word segmentation of the i-th text.

[0116] Next, the electronic device can perform one-hot encoding processing on the first-level label corresponding to each text to obtain the encoding corresponding to the first-level label of each text. After obtaining the encoding corresponding to the first-level label of each text, the electronic device can determine a first-level label reference matrix according to the encoding corresponding to the first-level label of each text. Among them, the first dimension (row) of the first-level label matrix represents each text in the multiple texts, and the second dimension (column) represents the encoding corresponding to the first-level label of each text in the multiple texts. For example, the i-th row of the first-level label reference matrix represents the encoding corresponding to the first-level label of the i-th text.

[0117] For example, assume there are 5 first-level labels, namely L1, L2, L3, L4, L5. When a certain text contains L1, the encoding corresponding to the first-level label of this text is: 10000; when a certain text contains L2, the encoding corresponding to the first-level label of this text is: 01000; when a certain text contains L3, the encoding corresponding to the first-level label of this text is: 00100; when a certain text contains L4, the encoding corresponding to the first-level label of this text is: 00010; when a certain text contains L5, the encoding corresponding to the first-level label of this text is: 00001.

[0118] For example, assume there are 15 pieces of text. The first-level label corresponding to the first piece of text is L2, the first-level label corresponding to the second piece of text is L3, the first-level label corresponding to the third piece of text is L1, the first-level label corresponding to the fourth piece of text is L2, the first-level label corresponding to the fifth piece of text is L5, the first-level label corresponding to the sixth piece of text is L4, the first-level label corresponding to the seventh piece of text is L1, the first-level label corresponding to the eighth piece of text is L5, the first-level label corresponding to the ninth piece of text is L1, the first-level label corresponding to the tenth piece of text is L3, the first-level label corresponding to the eleventh piece of text is L4, the first-level label corresponding to the twelfth piece of text is L4, the first-level label corresponding to the thirteenth piece of text is L5, the first-level label corresponding to the fourteenth piece of text is L3, and the first-level label corresponding to the fifteenth piece of text is L4.

[0119] It can be determined that the code corresponding to the first-level label of the first piece of text is: 01000, the code corresponding to the first-level label of the second piece of text is: 00100, the code corresponding to the first-level label of the third piece of text is: 10000, the code corresponding to the first-level label of the fourth piece of text is: 01000, the code corresponding to the first-level label of the fifth piece of text is: 00001, the code corresponding to the first-level label of the sixth piece of text is: 00010, the code corresponding to the first-level label of the seventh piece of text is: 10000, the code corresponding to the first-level label of the eighth piece of text is: 00001, the code corresponding to the first-level label of the ninth piece of text is: 10000, the code corresponding to the first-level label of the tenth piece of text is: 00100, the code corresponding to the first-level label of the eleventh piece of text is: 00010, the code corresponding to the first-level label of the twelfth piece of text is: 00010, the code corresponding to the first-level label of the thirteenth piece of text is: 00001, the code corresponding to the first-level label of the fourteenth piece of text is: 00100, and the code corresponding to the first-level label of the fifteenth piece of text is: 00010. Then, according to the codes corresponding to the first-level labels of each piece of text in the above 15 pieces of text, the first-level label reference matrix y1 is as Figure 4 shown.

[0120] For example, the electronic device can perform one-hot encoding processing on the second-level labels corresponding to each piece of text to obtain the codes corresponding to the second-level labels of each piece of text. After obtaining the codes corresponding to the second-level labels of each piece of text, the electronic device can determine the second-level label reference matrix according to the codes corresponding to the second-level labels of each piece of text. Among them, the first dimension (row) of the second-level label reference matrix is each piece of text in the multiple pieces of text, and the second dimension (column) is the codes corresponding to the second-level labels of each piece of text in the multiple pieces of text. For example, the i-th row of the second-level label reference matrix represents the code corresponding to the second-level label of the i-th piece of text.

[0121] For example, assume there are 15 secondary labels, namely S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, S11, S12, S13, S14, S15. When a piece of text contains S1, the code corresponding to the secondary label of this text is: 100000000000000; when a piece of text contains S2, the code corresponding to the secondary label of this text is: 010000000000000; when a piece of text contains S3, the code corresponding to the secondary label of this text is: 001000000000000; when a piece of text contains S4, the code corresponding to the secondary label of this text is: 000100000000000; when a piece of text contains S5, the code corresponding to the secondary label of this text is: 000010000000000; when a piece of text contains S6, the code corresponding to the secondary label of this text is: 000001000000000; when a piece of text contains S7, the code corresponding to the secondary label of this text is: 000000100000000; when a piece of text contains S8, the code corresponding to the secondary label of this text is: 000000010000000; when a piece of text contains S9, the code corresponding to the secondary label of this text is: 000000001000000; when a piece of text contains S10, the code corresponding to the secondary label of this text is: 000000000100000; when a piece of text contains S11, the code corresponding to the secondary label of this text is: 000000000010000; when a piece of text contains S12, the code corresponding to the secondary label of this text is: 000000000001000; when a piece of text contains S13, the code corresponding to the secondary label of this text is: 000000000000100; when a piece of text contains S14, the code corresponding to the secondary label of this text is: 000000000000010; when a piece of text contains S15, the code corresponding to the secondary label of this text is: 000000000000001.

[0122] For example, assume there are 15 pieces of text. The secondary label corresponding to the 1st piece of text is S3, the secondary label corresponding to the 2nd piece of text is S7, the secondary label corresponding to the 3rd piece of text is S4, the secondary label corresponding to the 4th piece of text is S5, the secondary label corresponding to the 5th piece of text is S10, the secondary label corresponding to the 6th piece of text is S12, the secondary label corresponding to the 7th piece of text is S1, the secondary label corresponding to the 8th piece of text is S14, the secondary label corresponding to the 9th piece of text is S2, the secondary label corresponding to the 10th piece of text is S6, the secondary label corresponding to the 11th piece of text is S11, the secondary label corresponding to the 12th piece of text is S15, the secondary label corresponding to the 13th piece of text is S13, the secondary label corresponding to the 14th piece of text is S8, and the secondary label corresponding to the 15th piece of text is S9.

[0123] It can be determined that the code corresponding to the secondary label of the 1st piece of text is: 001000000000000, the code corresponding to the secondary label of the 2nd piece of text is: 000000100000000, the code corresponding to the secondary label of the 3rd piece of text is: 000100000000000, the code corresponding to the secondary label of the 4th piece of text is: 000010000000000, the code corresponding to the secondary label of the 5th piece of text is: 000000000100000, the code corresponding to the secondary label of the 6th piece of text is: 000000000001000, the code corresponding to the secondary label of the 7th piece of text is: 100000000000000, the code corresponding to the secondary label of the 8th piece of text is: 000000000000010, the code corresponding to the secondary label of the 9th piece of text is: 010000000000000, the code corresponding to the secondary label of the 10th piece of text is: 000001000000000, the code corresponding to the secondary label of the 11th piece of text is: 000000000010000, the code corresponding to the secondary label of the 12th piece of text is: 000000000000001, the code corresponding to the secondary label of the 13th piece of text is: 000000000000100, the code corresponding to the secondary label of the 14th piece of text is: 000000010000000, and the code corresponding to the secondary label of the 15th piece of text is: 000000001000000. Then, the secondary label reference matrix y2 composed of the codes corresponding to the secondary labels of each piece of text in the above 15 pieces of text is as Figure 5 shown.

[0124] Finally, the electronic device can determine the data to be trained according to the word segmentation matrix, the primary label reference matrix y1, and the secondary label reference matrix y2.

[0125] For example, if the form of the data to be trained in the input model is (x, y), then, in the embodiments of the present application, the word segmentation matrix can be used as x to be input into the model, and the first-level label reference matrix y1 and the second-level label reference matrix y2 can be used as y to be input into the model.

[0126] 214. The electronic device obtains a preset relationship dependency matrix, which is used to represent the hierarchical relationship between the first-level labels and the second-level labels.

[0127] For example, the electronic device can obtain, such as Figure 3 the preset relationship dependency matrix M0 shown.

[0128] 215. The electronic device inputs the data to be trained and the preset relationship dependency matrix into the model to be trained, so as to obtain a first-level label prediction matrix and a second-level label prediction matrix, and each element in the first-level label prediction matrix is a real number.

[0129] For example, the electronic device can input the data to be trained (x, y) and the preset relationship dependency matrix M0 into the model to be trained for training, so as to obtain a first-level label prediction matrix and a second-level label prediction matrix. Wherein, x is the word segmentation matrix, and y is the first-level label reference matrix y1 and the second-level label reference matrix y2.

[0130] In the embodiments of the present application, the model to be trained can first initialize its parameters. Then, the electronic device can input the data to be trained and the preset relationship dependency matrix into the model to be trained, and obtain the output result through the forward propagation of each layer such as the convolutional layer, the downsampling layer, and the fully connected layer, that is, obtain the first-level label prediction matrix and the second-level label prediction matrix.

[0131] It can be understood that the network structure of the model to be trained can adopt any one of CNN and RNN, such as LSTM, GRU, Bi-LSTM, etc.

[0132] After being trained by the model to be trained, each element of the first-level label prediction matrix and the second-level label prediction matrix is a real number. Among them, the first dimension (row) of the first-level label prediction matrix represents each text in multiple texts, and the second dimension (column) represents the probability that each text corresponding to the multiple texts predicted by the model to be trained is each of the multiple first-level labels. For example, the element in the i-th row and the j-th column represents the probability that the first-level label corresponding to the i-th text is the j-th first-level label. For example, the first-level label prediction matrix P1 can be as Figure 6 shown.

[0133] In this Figure 6Among them, the 0th row and 0th column represent the probability that the first-level label corresponding to the first piece of text is L1. The 1st row and 1st column represent the probability that the first-level label corresponding to the second piece of text is L2. The 2nd row and 2nd column represent the probability that the first-level label corresponding to the third piece of text is L3. The 3rd row and 3rd column represent the probability that the first-level label corresponding to the fourth piece of text is L4. The 4th row and 4th column represent the probability that the first-level label corresponding to the fifth piece of text is L5... The 10th row and 0th column represent the probability that the first-level label corresponding to the 11th piece of text is L1. The 11th row and 1st column represent the probability that the first-level label corresponding to the 12th piece of text is L2. The 12th row and 2nd column represent the probability that the first-level label corresponding to the 13th piece of text is L3. The 13th row and 3rd column represent the probability that the first-level label corresponding to the 14th piece of text is L4. The 14th row and 4th column represent the probability that the first-level label corresponding to the 15th piece of text is L5.

[0134] The first dimension (row) of the secondary label prediction matrix represents each piece of text among multiple pieces of text, and the second dimension (column) represents the probability that each secondary label corresponding to each piece of text predicted by the model to be trained is each secondary label among multiple secondary labels. For example, the i-th row and j-th column represent the probability that the secondary label corresponding to the i-th piece of text is the j-th secondary label. For example, the secondary label prediction matrix P2 can be as Figure 7 shown.

[0135] In this Figure 7 Among them, the 0th row and 0th column represent the probability that the secondary label corresponding to the first piece of text is S1. The 1st row and 1st column represent the probability that the secondary label corresponding to the second piece of text is S2. The 2nd row and 2nd column represent the probability that the secondary label corresponding to the third piece of text is S3. The 3rd row and 3rd column represent the probability that the secondary label corresponding to the fourth piece of text is S4. The 4th row and 4th column represent the probability that the secondary label corresponding to the fifth piece of text is S5. The 5th row and 5th column represent the probability that the secondary label corresponding to the sixth piece of text is S6. The 6th row and 6th column represent the probability that the secondary label corresponding to the seventh piece of text is S7. The 7th row and 7th column represent the probability that the secondary label corresponding to the eighth piece of text is S8. The 8th row and 8th column represent the probability that the secondary label corresponding to the ninth piece of text is S9. The 9th row and 9th column represent the probability that the secondary label corresponding to the tenth piece of text is S10. The 10th row and 10th column represent the probability that the secondary label corresponding to the 11th piece of text is S11. The 11th row and 11th column represent the probability that the secondary label corresponding to the 12th piece of text is S12. The 12th row and 12th column represent the probability that the secondary label corresponding to the 13th piece of text is S13. The 13th row and 13th column represent the probability that the secondary label corresponding to the 14th piece of text is S14. The 14th row and 14th column represent the probability that the secondary label corresponding to the 15th piece of text is S15, and so on.

[0136] 216. The electronic device performs integerization processing on the first-level label prediction matrix, so that each element in the first-level label prediction matrix changes from a real number to an integer, obtaining a first-level label integer matrix, and the value of each element in the first-level label integer matrix is 0 or 1.

[0137] For example, the electronic device can perform integerization processing on the first-level label prediction matrix, so that each element in the first-level label prediction matrix changes from a real number to an integer, obtaining a first-level label integer matrix, and the value of each element in the first-level label integer matrix is 0 or 1.

[0138] For instance, the electronic device can perform integerization processing on the first-level label prediction matrix P1, converting the first-level label prediction matrix P1 into a 0-1 integer matrix P1-1. In this 0-1 integer matrix P1-1, only the position with the highest probability is 1, and the rest are 0. That is, the 0-1 integer matrix P1-1 is as Figure 8 shown.

[0139] 217. The electronic device multiplies the first-level label integer matrix by a preset relationship dependency matrix to obtain a target relationship dependency matrix.

[0140] Among them, multiplying matrix A by matrix B means calculating the product of matrix A and matrix B.

[0141] For example, the electronic device can calculate the product of the first-level label integer matrix and the preset relationship dependency matrix, and determine the product as the target relationship dependency matrix. The target relationship dependency matrix M = 0-1 integer matrix P1-1 × M0. For instance, the target relationship dependency matrix M is as Figure 9 shown.

[0142] 218. The electronic device determines a first loss value according to the first-level label prediction matrix and the first-level label reference matrix.

[0143] For example, the electronic device can use the first-level label prediction matrix P1 and the first-level label reference matrix y1 as parameters and input them into a specified loss function, so as to calculate the loss value between the first-level label prediction matrix P1 and the first-level label reference matrix y1, and this loss value is the first loss value.

[0144] For instance, the specified loss function can be: H(P1, y1) = -∑ n y1(n)·logP1(n). Where n represents the nth column element of the matrix.

[0145] In the embodiments of the present application, the loss function is generally used to estimate the degree of inconsistency between the predicted value of the model (such as the first-level label prediction matrix P1) and the true value (such as the first-level label reference matrix). Generally, the smaller the loss function, the better the robustness of the model. The loss function can be set according to actual needs, and the present application does not limit it.

[0146] 219. The electronic device determines a target matrix according to the second-level label prediction matrix and the target relationship dependency matrix.

[0147] 220. The electronic device determines a second loss value according to the target matrix and the second-level label reference matrix.

[0148] For example, the electronic device can determine the target matrix P3 according to the second-level label prediction matrix P2 and the target relationship dependency matrix M.

[0149] The electronic device can use the target matrix P3 and the second-level label reference matrix y2 as parameters and input them into a specified loss function, so as to calculate the loss value between the target matrix P3 and the second-level label reference matrix y2, and this loss value is the second loss value.

[0150] For example, the specified loss function can be: H(P3, y2) = -∑ n y2(n)·logP3(n). Where n represents the nth column element of the matrix.

[0151] Among them, in order to avoid the situation of log(0) when taking the logarithm, the electronic device can add the target relationship dependency matrix M to a positive number e approaching 0 to obtain the first matrix M1. Then multiply the second-level label prediction matrix P2 by the first matrix M1 to obtain the target matrix P3. Where multiplying matrix A by matrix B means calculating the Hadamard product of matrix A and matrix B.

[0152] 221. The electronic device determines a target loss value according to the first loss value and the second loss value.

[0153] For example, after obtaining the first loss value and the second loss value, the electronic device can determine the target loss value according to the first loss value and the second loss value. This target loss value is the loss value corresponding to the model to be trained, and this target loss value is used to characterize whether the model to be trained reaches the optimum.

[0154] 222. Whenever a target loss value is obtained after a batch of data to be trained is trained, each time a target loss value is obtained, the target loss value is fed back to the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirm that the model training is completed, and obtain the trained model.

[0155] It can be understood that after obtaining a target loss value, the target loss value can be backpropagated to each layer of the model to be trained, so as to adjust the parameters of the model to be trained. Subsequently, the electronic device can continue to obtain another batch of texts labeled with first-level labels and second-level labels to obtain another data to be trained, and input the data to be trained with adjusted parameters into the model to be trained to continue training the model to be trained. Among them, the other data to be trained and the data to be trained obtained in process 213 are two different batches of data. The determination process of the other data to be trained can refer to processes 205 to 213. After obtaining a target loss value this time, the target loss value can still be backpropagated to each layer of the model to be trained, so as to adjust the parameters again until the model to be trained converges, confirm that the model training is completed, and obtain the trained model. Among them, when the target loss value gradually approaches a certain value, or fluctuates around a certain value, and the loss change is less than a very small positive number, it can be confirmed that the model to be trained converges.

[0156] In some other embodiments, after each batch of data to be trained is input to train the model to be trained, a trained model can be obtained. The electronic device can obtain a batch of verification data from the verification set and input it into the trained model to verify the accuracy of the trained model. When the accuracy obtained this time is greater than the accuracy obtained last time, the electronic device can save the trained model this time. When the accuracy obtained this time is less than the accuracy obtained last time, the electronic device can not save the trained model this time. When the accuracy of the trained model obtained multiple times does not increase, for example, when the accuracies of the trained models obtained multiple times are 87%, 86.9%, 86.7%, and 86.8% respectively, the electronic device can confirm that the model training is completed.

[0157] In some embodiments, process 221 may include:

[0158] The electronic device multiplies the first loss value by the first weight value to obtain a third loss value;

[0159] The electronic device multiplies the second loss value by the second weight value to obtain a fourth loss value, and the second weight value is less than the first weight value;

[0160] The electronic device determines the target loss value according to the third loss value and the fourth loss value.

[0161] To improve the accuracy of the model in predicting the first-level labels, after obtaining the first loss value and the second loss value, the electronic device may multiply the first loss value by a relatively large weight value to obtain a third loss value; multiply the second loss value by a relatively small weight value to obtain a fourth loss value. Then, the electronic device may add the third loss value and the fourth loss value to obtain a target loss value. Herein, the first weight value and the second weight value may be set according to actual situations. For example, the first weight value may be 0.6 and the second weight value may be 0.4.

[0162] In some embodiments, process 219 may include:

[0163] The electronic device adds the target relationship dependency matrix to a preset value to obtain a first matrix;

[0164] The electronic device multiplies the second-level label prediction matrix by the first matrix to obtain a target matrix.

[0165] To avoid the situation of log(0) when taking the logarithm, the electronic device may add the target relationship dependency matrix to a preset value, such as adding a positive number e approaching 0, to obtain a first matrix M1. Then, multiply the second-level label prediction matrix P2 by the first matrix M1 to obtain a target matrix P3. Herein, multiplying matrix A by matrix B means calculating the Hadamard product of matrix A and matrix B.

[0166] In some embodiments, process 207 may include:

[0167] The electronic device constructs a dictionary according to the word segmentation corresponding to each text, and the dictionary includes multiple word segmentations and their corresponding encodings;

[0168] The electronic device determines the encodings corresponding to the word segmentations of each text according to the word segmentations corresponding to each text and the dictionary.

[0169] For example, the electronic device may sort out the word segmentations corresponding to multiple texts respectively, and find out the different word segmentations among the word segmentations corresponding to these multiple texts respectively. Then, the electronic device may encode these different word segmentations and construct a dictionary according to these different word segmentations and the encodings corresponding to these different word segmentations respectively.

[0170] For example, assume that the word segmentation corresponding to a certain text is: we, of, motherland, is, China. The word segmentation corresponding to another text is: we, of, motherland, is, South Korea. Then, the dictionary constructed according to the word segmentation corresponding to this text and the word segmentation corresponding to another text may be as Figure 10 shown.

[0171] Then, the electronic device may determine the encodings corresponding to the word segmentations of each text according to the word segmentations corresponding to each text and the dictionary.

[0172] For example, if a text is: Our motherland is China, then the codes corresponding to the word segmentation of this text are: 1, 2, 3, 4, 5.

[0173] Please refer to Figure 11 , Figure 11 which is a schematic flowchart of the classification method provided by the embodiments of the present application. The process of this classification method may include:

[0174] 301. Obtain the text to be classified.

[0175] For example, an electronic device may obtain a text that needs to be marked with a first-level label and a second-level label, that is, the text to be classified.

[0176] 302. Input the text to be classified into the trained model to obtain a first-level label probability matrix and a second-level label prediction probability matrix. Each element in the first-level label probability matrix corresponds to a first-level label, and each element in the first-level label probability matrix is a real number. Each element in the second-level label prediction probability matrix corresponds to a second-level label, and each element in the second-level label prediction probability matrix is a real number.

[0177] 303. Determine the first-level label corresponding to the text to be classified according to the first-level label probability matrix. The first-level label corresponding to the element with the largest value in the first-level label probability matrix is the first-level label corresponding to the text to be classified.

[0178] After obtaining the text to be classified, the electronic device may input the text to be classified into the trained model to obtain a first-level label probability matrix. Among them, each element in the first-level label probability matrix corresponds to a first-level label. The value of each element in the first-level label probability matrix represents the probability that the first-level label corresponding to the text to be classified is each of the first-level labels among multiple first-level labels. For example, as Figure 12 shown, assume that multiple first-level labels are: L1, L2, L3, L4, L5, and the first-level label probability matrix may be P4. In the first-level label probability matrix P4, the 0th column represents the probability that the first-level label corresponding to the text to be classified is L1, the 1st column represents the probability that the first-level label corresponding to the text to be classified is L2... the 4th column represents the probability that the first-level label corresponding to the text to be classified is L5. It can be seen that in the first-level label probability matrix P4, the probability of the 0th column is the largest. Therefore, the first-level label corresponding to the text to be classified is L1.

[0179] In the embodiments of the present application, the trained model may be generated by using the model construction method described in the above embodiments. The specific generation process may refer to the relevant descriptions in the above embodiments and will not be elaborated here.

[0180] 304. Integerize the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, obtaining a first-level label integerization matrix, where the value of each element in this first-level label integerization matrix is 0 or 1.

[0181] For example, after obtaining the first-level label probability matrix, the electronic device can integerize the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, obtaining a first-level label integerization matrix. Among them, the value of each element in this first-level label integerization matrix is 0 or 1.

[0182] For example, as Figure 12 shown, the electronic device can integerize the first-level label probability matrix P4 and convert this first-level label probability matrix P4 into a 0-1 integer matrix P5. In this 0-1 integer matrix P5, only the place with the maximum probability is 1, and the rest are 0.

[0183] 305. Determine the first relationship dependency matrix according to the first-level label integerization matrix and the preset relationship dependency matrix.

[0184] For example, as Figure 12 shown, the electronic device can cross-multiply this first-level label integerization matrix P5 by the preset relationship dependency matrix M0 as Figure 3 shown to obtain the first relationship dependency matrix M2. Among them, matrix A cross-multiplying matrix B means calculating the product of matrix A and matrix B.

[0185] 306. Determine the second-level label probability matrix according to the first relationship dependency matrix and the second-level label prediction probability matrix, where each element in this second-level label probability matrix corresponds to a second-level label.

[0186] 307. Determine the second-level label corresponding to the text to be classified according to the second-level label probability matrix. The second-level label corresponding to the element with the largest value in this second-level label probability matrix is the second-level label corresponding to the text to be classified.

[0187] For example, after obtaining the text to be classified, the electronic device can input this text to be classified into the trained model to obtain the second-level label prediction probability matrix. Among them, each element in this second-level label prediction probability matrix corresponds to a second-level label. The value of each element in this second-level label prediction probability matrix represents the probability that the second-level label corresponding to this text to be classified is each of the multiple second-level labels. For example, as Figure 12As shown, assume that multiple secondary labels are respectively: S1, S2, S3, S4, S5, S6, S7, S8, S9, S10, S11, S12, S13, S14, S15, and the secondary label prediction probability matrix can be P6. In the secondary label prediction probability matrix P6, the 0th column represents the probability that the secondary label corresponding to the text to be classified is S1, the 1st column represents the probability that the secondary label corresponding to the text to be classified is S2... the 14th column represents the probability that the secondary label corresponding to the text to be classified is S15.

[0188] Then, the electronic device can multiply the secondary label prediction probability matrix P6 by the first relationship dependence matrix M2 to obtain the secondary label probability matrix P7. It can be seen that in the first relationship dependence matrix M2, only the 0th column, the 1st column, and the 3rd column are 1, and the others are 0. Therefore, only the 0th column, the 1st column, and the 3rd column in the secondary label probability matrix P7 have a non-zero value, and the others are 0. And the value of each element in the secondary label probability matrix P7 represents the probability that the secondary label corresponding to the text to be classified is each secondary label among the multiple secondary labels. Then, it only needs to determine the maximum value from the 0th column, the 1st column, and the 2nd column in the secondary label probability matrix P7, and determine the secondary label corresponding to the maximum value as the secondary label corresponding to the text to be trained. Compared with the solution in the related art that needs to determine the maximum value from the values corresponding to all elements in the secondary label probability matrix and determine the secondary label corresponding to the maximum value as the secondary label corresponding to the text to be trained, the classification method provided by the embodiment of the present application has a higher accuracy rate.

[0189] As Figure 12 shown, in the secondary label probability matrix P7, the probability of the element in the 1st column is the largest, so the secondary label corresponding to the text to be classified is S2.

[0190] After determining the primary label and the secondary label corresponding to the text to be classified, the electronic device can mark the primary label and the secondary label on the text to be classified. Subsequently, the electronic device can classify the text to be classified into the corresponding category. For example, if the primary label is "TV drama" and the secondary label is "costume", then the electronic device can classify the text to be classified into the costume category under the TV drama category.

[0191] It can be understood that in the embodiments of the present application, the electronic device can sequentially execute processes 301 to 307 to mark different texts to be classified with corresponding first-level tags and second-level tags. When it is necessary to recommend texts to a certain user, the electronic device can obtain the tags corresponding to the user, and then select the corresponding texts from the texts marked with first-level tags and second-level tags according to the tags corresponding to the user, and recommend the texts to the user. Among them, the tags corresponding to the user can be determined by the electronic device according to the user's browsing preference for articles. For example, if the articles that a certain user often browses carry the first-level tag L1 and the second-level tag S2, then the tags corresponding to the user are L1 and S2. When pushing articles to this user, articles marked with L1 and S2 can be pushed to him.

[0192] It can be understood that if it is necessary to enable the trained model to also provide services such as marking third-level tags and fourth-level tags, the electronic device can determine the hierarchical relationship between the third-level tags and the second-level tags, and establish a relationship dependence matrix between the third-level tags and the second-level tags according to this hierarchical relationship. The establishment method can refer to the establishment method of the relationship dependence matrix between the first-level tags and the second-level tags. Similarly, the electronic device can also determine the hierarchical relationship between the fourth-level tags and the third-level tags, and establish a relationship dependence matrix between the fourth-level tags and the third-level tags according to this hierarchical relationship. The establishment method can also refer to the establishment method of the relationship dependence matrix between the first-level tags and the second-level tags. The electronic device can encode the third-level tags and fourth-level tags of multiple texts according to the same encoding method as the first-level tags, so as to form a third-level tag reference matrix and a fourth-level tag reference matrix. Then, the electronic device can also input the relationship dependence matrix between the third-level tags and the second-level tags, the relationship dependence matrix between the fourth-level tags and the third-level tags, the third-level tag reference matrix, and the fourth-level tag reference matrix into the model to be trained, so as to finally train a model that can provide services for marking first-level tags, second-level tags, third-level tags, and fourth-level tags.

[0193] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of the model construction device provided by the embodiments of the present application. The model construction device may include: a first acquisition module 401, a second acquisition module 402, a first training module 403, a first determination module 404, a second determination module 405, and a second training module 406.

[0194] The first acquisition module 401 is used to acquire training data to be trained, where the training data to be trained includes a token matrix composed of encodings corresponding to tokens of each text in multiple texts, a first-level tag reference matrix composed of encodings corresponding to first-level tags of each text, and a second-level tag reference matrix composed of encodings corresponding to second-level tags of each text;

[0195] The second acquisition module 402 is configured to acquire a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between the first-level labels and the second-level labels;

[0196] The first training module 403 is configured to input the data to be trained and the preset relationship dependency matrix into the model to be trained, so as to obtain a first-level label prediction matrix and a second-level label prediction matrix;

[0197] The first determination module 404 is configured to determine a target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix;

[0198] The second determination module 405 is configured to determine a target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix, where the target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix;

[0199] The second training module 406 is configured to obtain a target loss value every time a batch of data to be trained is trained. Each time a target loss value is obtained, the target loss value is fed back to the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirm that the model training is completed, and obtain the trained model.

[0200] In some embodiments, the second determination module 405 may be configured to: determine a first loss value according to the first-level label prediction matrix and the first-level label reference matrix; determine a target matrix according to the second-level label prediction matrix and the target relationship dependency matrix; determine a second loss value according to the target matrix and the second-level label reference matrix; and determine a target loss value according to the first loss value and the second loss value.

[0201] In some embodiments, the second determination module 405 may be configured to: multiply the first loss value by a first weight value to obtain a third loss value; multiply the second loss value by a second weight value to obtain a fourth loss value, where the second weight value is less than the first weight value; and determine a target loss value according to the third loss value and the fourth loss value.

[0202] In some embodiments, the second determination module 405 may be configured to: add the target relationship dependency matrix to a preset value to obtain a first matrix; and multiply the second-level label prediction matrix by the first matrix to obtain a target matrix.

[0203] In some embodiments, the first determination module 404 may be configured to: perform integerization processing on the first-level label prediction matrix to change each element in the first-level label prediction matrix from a real number to an integer, obtaining a first-level label integer matrix, where the value of each element in the first-level label integer matrix is 0 or 1; multiply the first-level label integer matrix by the preset relationship dependency matrix to obtain a target relationship dependency matrix.

[0204] In some embodiments, the first acquisition module 401 may be configured to: acquire multiple texts, the first-level labels corresponding to each text, and the second-level labels corresponding to each text; perform word segmentation processing on each text to obtain the word segmentation corresponding to each text; determine the encoding corresponding to the word segmentation of each text; determine a word segmentation matrix according to the encoding corresponding to the word segmentation of each text; perform one-hot encoding processing on the first-level labels corresponding to each text to obtain the encoding corresponding to the first-level labels of each text; determine a first-level label reference matrix according to the encoding corresponding to the first-level labels of each text; perform one-hot encoding processing on the second-level labels corresponding to each text to obtain the encoding corresponding to the second-level labels of each text; determine a second-level label reference matrix according to the encoding corresponding to the second-level labels of each text; determine the data to be trained according to the word segmentation matrix, the first-level label reference matrix, and the second-level label reference matrix.

[0205] In some embodiments, the first acquisition module 401 may be configured to: construct a dictionary according to the word segmentation corresponding to each text, where the dictionary includes multiple word segmentations and their corresponding encodings; determine the encoding corresponding to the word segmentation of each text according to the word segmentation corresponding to each text and the dictionary.

[0206] In some embodiments, the first acquisition module 401 may be configured to: acquire multiple first-level labels; acquire multiple second-level labels; determine the hierarchical relationship between each first-level label and each second-level label; establish a preset relationship dependency matrix according to the hierarchical relationship.

[0207] Please refer to Figure 14 , Figure 14 , which is a schematic structural diagram of the classification device provided by the embodiment of the present application. The classification device may include: a third acquisition module 501, a prediction module 502, a third determination module 503, an integerization module 504, a fourth determination module 505, a fifth determination module 506, and a sixth determination module 507.

[0208] The third acquisition module 501 is configured to acquire the text to be classified;

[0209] A prediction module 502, configured to input the text to be classified into a trained model to obtain a first-level label probability matrix and a second-level label prediction probability matrix. Each element in the first-level label probability matrix corresponds to a first-level label, and all elements in the first-level label probability matrix are real numbers. Each element in the second-level label prediction probability matrix corresponds to a second-level label, and all elements in the second-level label prediction probability matrix are real numbers;

[0210] A third determination module 503, configured to determine the first-level label corresponding to the text to be classified according to the first-level label probability matrix. The first-level label corresponding to the element with the largest value in the first-level label probability matrix is the first-level label corresponding to the text to be classified;

[0211] An integerization module 504, configured to perform integerization processing on the first-level label probability matrix, so that each element in the first-level label probability matrix changes from a real number to an integer, to obtain a first-level label integerization matrix. The value of each element in the first-level label integerization matrix is 0 or 1;

[0212] A fourth determination module 505, configured to determine a first relationship dependency matrix according to the first-level label integerization matrix and a preset relationship dependency matrix;

[0213] A fifth determination module 506, configured to determine a second-level label probability matrix according to the first relationship dependency matrix and the second-level label prediction probability matrix. Each element in the second-level label probability matrix corresponds to a second-level label;

[0214] A sixth determination module 507, configured to determine the second-level label corresponding to the text to be classified according to the second-level label probability matrix. The second-level label corresponding to the element with the largest value in the second-level label probability matrix is the second-level label corresponding to the text to be classified.

[0215] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is enabled to execute the processes in the model construction method or classification method provided in this embodiment.

[0216] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the processes in the model construction method or classification method provided in this embodiment by calling the computer program stored in the memory.

[0217] For example, the above-mentioned electronic device may be a mobile terminal such as a tablet computer or a smart phone. Please refer to Figure 15 , Figure 15 which is the first structural schematic diagram of the electronic device provided by the embodiment of the present application.

[0218] The electronic device 600 may include components such as a memory 601 and a processor 602. Those skilled in the art can understand that Figure 15 the structure of the electronic device shown in

[0219] does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0220] The memory 601 can be used to store application programs and data. The application programs stored in the memory 601 contain executable code. The application programs can form various functional modules. The processor 602 executes various functional applications and data processing by running the application programs stored in the memory 601.

[0221] In this embodiment, the processor 602 in the electronic device will load the executable code corresponding to the processes of one or more application programs into the memory 601 according to the following instructions, and the processor 601 will run the application programs stored in the memory 601, thereby implementing the process:

[0222] Obtain the data to be trained, where the data to be trained includes a token matrix composed of the encodings corresponding to the tokens of each text in multiple texts, a first-level label reference matrix composed of the encodings corresponding to the first-level labels of each text, and a second-level label reference matrix composed of the encodings corresponding to the second-level labels of each text;

[0223] Obtain a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between the first-level labels and the second-level labels;

[0224] Input the data to be trained and the preset relationship dependency matrix into the model to be trained to obtain a first-level label prediction matrix and a second-level label prediction matrix;

[0225] Determine a target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix;

[0226] Determine a target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix, where the target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix;

[0227] Whenever a batch of data to be trained is trained to obtain a target loss value, each time a target loss value is obtained, the target loss value is fed back into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirming that the model training is completed, and obtaining the trained model.

[0228] In this embodiment, the processor 602 in the electronic device will load the executable code corresponding to the processes of one or more application programs into the memory 601 according to the following instructions, and the processor 601 will run the application programs stored in the memory 601, thereby implementing the process:

[0229] Obtain the text to be classified;

[0230] Input the text to be classified into the trained model to obtain a first-level label probability matrix and a second-level label prediction probability matrix. Each element in the first-level label probability matrix corresponds to a first-level label, and all elements in the first-level label probability matrix are real numbers. Each element in the second-level label prediction probability matrix corresponds to a second-level label, and all elements in the second-level label prediction probability matrix are real numbers;

[0231] According to the first-level label probability matrix, determine the first-level label corresponding to the text to be classified. The first-level label corresponding to the element with the largest value in the first-level label probability matrix is the first-level label corresponding to the text to be classified;

[0232] Perform integerization processing on the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, obtaining a first-level label integerization matrix, and the value of the element in the first-level label integerization matrix is 0 or 1;

[0233] According to the first-level label integerization matrix and the preset relationship dependency matrix, determine the first relationship dependency matrix;

[0234] According to the first relationship dependency matrix and the second-level label prediction probability matrix, determine the second-level label probability matrix. Each element in the second-level label probability matrix corresponds to a second-level label;

[0235] According to the second-level label probability matrix, determine the second-level label corresponding to the text to be classified. The second-level label corresponding to the element with the largest value in the second-level label probability matrix is the second-level label corresponding to the text to be classified.

[0236] Please refer to Figure 16 , Figure 16 which is the second structural schematic diagram of the electronic device provided by the embodiment of the present application.

[0237] The electronic device 600 may include components such as a memory 601, a processor 602, an input unit 603, an output unit 604, and a display screen 605.

[0238] The memory 601 can be used to store application programs and data. The application programs stored in the memory 601 contain executable code. The application programs can form various functional modules. The processor 602 executes various functional applications and data processing by running the application programs stored in the memory 601.

[0239] The processor 602 is the control center of the electronic device, connects various parts of the entire electronic device using various interfaces and lines, executes various functions of the electronic device and processes data by running or executing the application programs stored in the memory 601 and calling the data stored in the memory 601, thereby monitoring the electronic device as a whole.

[0240] The input unit 603 can be used to receive input digital, character information or user characteristic information (such as fingerprints), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0241] The output unit 604 can be used to display the information input by the user or the information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The output unit may include a display panel.

[0242] The display screen 605 can be used to display information such as text and pictures.

[0243] In this embodiment, the processor 602 in the electronic device will, according to the following instructions, load the executable code corresponding to the processes of one or more application programs into the memory 601, and the processor 602 will run the application programs stored in the memory 601 to implement the process:

[0244] Obtain the data to be trained, where the data to be trained includes a token matrix composed of the encodings corresponding to the tokens of each text in multiple texts, a first-level label reference matrix composed of the encodings corresponding to the first-level labels of each text, and a second-level label reference matrix composed of the encodings corresponding to the second-level labels of each text;

[0245] Obtain a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between the first-level labels and the second-level labels;

[0246] Input the data to be trained and the preset relationship dependency matrix into the model to be trained to obtain a first-level label prediction matrix and a second-level label prediction matrix;

[0247] Determine a target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix;

[0248] Determine a target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix, where the target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix;

[0249] Whenever a batch of data to be trained is trained to obtain a target loss value, each time a target loss value is obtained, the target loss value is fed back into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirming that the model training is completed, and obtaining the trained model.

[0250] In some embodiments, when the processor 602 executes determining the target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix, where the target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix, it may execute: determining a first loss value according to the first-level label prediction matrix and the first-level label reference matrix; determining a target matrix according to the second-level label prediction matrix and the target relationship dependency matrix; determining a second loss value according to the target matrix and the second-level label reference matrix; and determining the target loss value according to the first loss value and the second loss value.

[0251] In some embodiments, when the processor 602 executes determining the target loss value according to the first loss value and the second loss value, it may execute: multiplying the first loss value by a first weight value to obtain a third loss value; multiplying the second loss value by a second weight value to obtain a fourth loss value, where the second weight value is less than the first weight value; and determining the target loss value according to the third loss value and the fourth loss value.

[0252] In some embodiments, when the processor 602 executes determining the target matrix according to the second-level label prediction matrix and the target relationship dependency matrix, it may execute: adding the target relationship dependency matrix to a preset value to obtain a first matrix; and multiplying the second-level label prediction matrix by the first matrix to obtain the target matrix.

[0253] In some embodiments, each element in the first-level label prediction matrix is a real number. When the processor 602 determines the target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix, it may perform: performing an integerization process on the first-level label prediction matrix to change each element in the first-level label prediction matrix from a real number to an integer, obtaining a first-level label integer matrix, where the value of each element in the first-level label integer matrix is 0 or 1; multiplying the first-level label integer matrix by the preset relationship dependency matrix to obtain the target relationship dependency matrix.

[0254] In some embodiments, when the processor 602 obtains the data to be trained, where the data to be trained includes a token matrix composed of encodings corresponding to tokens of each text in multiple texts, a first-level label reference matrix composed of encodings corresponding to the first-level labels of each text, and a second-level label reference matrix composed of encodings corresponding to the second-level labels of each text, it may perform: obtaining multiple texts, the first-level label corresponding to each text, and the second-level label corresponding to each text; performing a tokenization process on each text to obtain the tokens corresponding to each text; determining the encodings corresponding to the tokens of each text; determining the token matrix according to the encodings corresponding to the tokens of each text; performing one-hot encoding on the first-level labels corresponding to each text to obtain the encodings corresponding to the first-level labels of each text; determining the first-level label reference matrix according to the encodings corresponding to the first-level labels of each text; performing one-hot encoding on the second-level labels corresponding to each text to obtain the encodings corresponding to the second-level labels of each text; determining the second-level label reference matrix according to the encodings corresponding to the second-level labels of each text; and determining the data to be trained according to the token matrix, the first-level label reference matrix, and the second-level label reference matrix.

[0255] In some embodiments, when the processor 602 determines the encodings corresponding to the tokens of each text, it may perform: constructing a dictionary according to the tokens of each text, where the dictionary includes multiple tokens and their corresponding encodings; and determining the encodings corresponding to the tokens of each text according to the tokens of each text and the dictionary.

[0256] In some embodiments, before the processor 602 obtains the data to be trained, it may perform: obtaining multiple first-level labels; obtaining multiple second-level labels; determining the hierarchical relationship between each first-level label and each second-level label; and establishing a preset relationship dependency matrix according to the hierarchical relationship.

[0257] In this embodiment, the processor 602 in the electronic device will load the executable code corresponding to the processes of one or more application programs into the memory 601 according to the following instructions, and the processor 602 will run the application programs stored in the memory 601, thereby implementing the process:

[0258] Obtain the text to be classified;

[0259] Input the text to be classified into the trained model to obtain a first-level label probability matrix and a second-level label prediction probability matrix. Each element in the first-level label probability matrix corresponds to a first-level label, and all elements in the first-level label probability matrix are real numbers. Each element in the second-level label prediction probability matrix corresponds to a second-level label, and all elements in the second-level label prediction probability matrix are real numbers;

[0260] Determine the first-level label corresponding to the text to be classified according to the first-level label probability matrix. The first-level label corresponding to the element with the largest value in the first-level label probability matrix is the first-level label corresponding to the text to be classified;

[0261] Perform integerization processing on the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, obtaining a first-level label integerization matrix. The value of the element in the first-level label integerization matrix is 0 or 1;

[0262] Determine a first relationship dependency matrix according to the first-level label integerization matrix and a preset relationship dependency matrix;

[0263] Determine a second-level label probability matrix according to the first relationship dependency matrix and the second-level label prediction probability matrix. Each element in the second-level label probability matrix corresponds to a second-level label;

[0264] Determine the second-level label corresponding to the text to be classified according to the second-level label probability matrix. The second-level label corresponding to the element with the largest value in the second-level label probability matrix is the second-level label corresponding to the text to be classified.

[0265] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the detailed descriptions of the model construction method / classification method above, and details will not be repeated here.

[0266] The model construction method / classification method device provided in the embodiments of the present application and the model construction method / classification method in the above embodiments belong to the same concept. Any method provided in the embodiments of the model construction method / classification method can be run on the model construction method / classification method device. The specific implementation process is detailed in the embodiments of the model construction method / classification method, and details will not be repeated here.

[0267] It should be noted that for the model construction method / classification method described in the embodiments of the present application, those of ordinary skill in the art can understand that all or part of the processes for implementing the model construction method / classification method described in the embodiments of the present application can be completed by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, such as stored in a memory and executed by at least one processor. During the execution process, it can include the processes of the embodiments of the model construction method / classification method. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0268] For the model construction method / classification method device described in the embodiments of the present application, its various functional modules can be integrated in a processing chip, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk or an optical disk, etc.

[0269] The above has introduced in detail a model construction method, a classification method, a device, a storage medium and an electronic device provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A model construction method, wherein, it includes: Obtain the data to be trained, where the data to be trained includes a token matrix composed of the encodings corresponding to the tokens of each text in multiple texts, a primary label reference matrix composed of the encodings corresponding to the primary labels of each text, and a secondary label reference matrix composed of the encodings corresponding to the secondary labels of each text; Obtain a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between the primary label and the secondary label; Input the data to be trained and the preset relationship dependency matrix into the model to be trained to obtain a primary label prediction matrix and a secondary label prediction matrix; According to the primary label prediction matrix and the preset relationship dependency matrix, determine the target relationship dependency matrix. Among them, each element in the primary label prediction matrix is a real number. Perform integerization processing on the primary label prediction matrix so that each element in the primary label prediction matrix changes from a real number to an integer, obtaining a primary label integer matrix. Each element value in the primary label integer matrix is 0 or 1. Multiply the primary label integer matrix by the preset relationship dependency matrix to obtain the target relationship dependency matrix; According to the primary label prediction matrix, the primary label reference matrix, the target matrix, and the secondary label reference matrix, determine the target loss value. The target matrix is determined according to the secondary label prediction matrix and the target relationship dependency matrix. Among them, add the target relationship dependency matrix and a preset value to obtain a first matrix, multiply the secondary label prediction matrix by the first matrix to obtain the target matrix; determine the first loss value according to the primary label prediction matrix and the primary label reference matrix; determine the second loss value according to the target matrix and the secondary label reference matrix; determine the target loss value according to the first loss value and the second loss value; Whenever a batch of data to be trained is trained to obtain a target loss value, every time a target loss value is obtained, the target loss value is backpropagated into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirm that the model training is completed, and obtain the trained model.

2. The model construction method according to claim 1, wherein, the determining the target loss value according to the first loss value and the second loss value includes: Multiply the first loss value by a first weight value to obtain a third loss value; Multiply the second loss value by a second weight value to obtain a fourth loss value, where the second weight value is less than the first weight value; Determine the target loss value according to the third loss value and the fourth loss value.

3. The model construction method according to claim 1, wherein, the obtaining the data to be trained, where the data to be trained includes a token matrix composed of the encodings corresponding to the tokens of each text in multiple texts, a primary label reference matrix composed of the encodings corresponding to the primary labels of each text, and a secondary label reference matrix composed of the encodings corresponding to the secondary labels of each text, includes: Obtain multiple texts, the primary label corresponding to each text, and the secondary label corresponding to each text; Perform word segmentation on each text to obtain the word segmentation corresponding to each text; Determine the encoding corresponding to the word segmentation of each text; Determine a word segmentation matrix according to the encoding corresponding to the word segmentation of each text; Perform one-hot encoding on the first-level labels corresponding to each text to obtain the encoding corresponding to the first-level labels of each text; Determine a first-level label reference matrix according to the encoding corresponding to the first-level labels of each text; Perform one-hot encoding on the second-level labels corresponding to each text to obtain the encoding corresponding to the second-level labels of each text; Determine a second-level label reference matrix according to the encoding corresponding to the second-level labels of each text; Determine the data to be trained according to the word segmentation matrix, the first-level label reference matrix, and the second-level label reference matrix.

4. The model construction method according to claim 3, wherein, the determining the encoding corresponding to the word segmentation of each text includes: Construct a dictionary according to the word segmentation corresponding to each text, the dictionary including a plurality of word segmentations and their corresponding encodings; Determine the encoding corresponding to the word segmentation of each text according to the word segmentation corresponding to each text and the dictionary.

5. The model construction method according to claim 1, wherein, before obtaining the data to be trained, it further includes: Obtain a plurality of first-level labels; Obtain a plurality of second-level labels; Determine the hierarchical relationship between each first-level label and each second-level label; Establish a preset relationship dependency matrix according to the hierarchical relationship.

6. A classification method, wherein, it includes: Obtain the text to be classified; Input the text to be classified into the trained model, the trained model being trained according to any one of the model construction methods of claims 1 to 5, to obtain a first-level label probability matrix and a second-level label prediction probability matrix, each element in the first-level label probability matrix corresponding to a first-level label, each element in the first-level label probability matrix being a real number, each element in the second-level label prediction probability matrix corresponding to a second-level label, and each element in the second-level label prediction probability matrix being a real number; Determine the first-level label corresponding to the text to be classified according to the first-level label probability matrix, the first-level label corresponding to the element with the largest value in the first-level label probability matrix being the first-level label corresponding to the text to be classified; Perform integerization processing on the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, to obtain a first-level label integerization matrix, the value of the element in the first-level label integerization matrix being 0 or 1; Determine a first relationship dependency matrix according to the first-level label integerization matrix and the preset relationship dependency matrix; Determine a second-level label probability matrix according to the first relationship dependency matrix and the second-level label prediction probability matrix, each element in the second-level label probability matrix corresponding to a second-level label; Determine the second-level label corresponding to the text to be classified according to the second-level label probability matrix, the second-level label corresponding to the element with the largest value in the second-level label probability matrix being the second-level label corresponding to the text to be classified.

7. A model construction device, wherein, it includes: A first acquisition module, configured to acquire training data to be trained, where the training data to be trained includes a token matrix composed of encodings corresponding to tokens of each text in multiple texts, a first-level label reference matrix composed of encodings corresponding to first-level labels of each text, and a second-level label reference matrix composed of encodings corresponding to second-level labels of each text; A second acquisition module, configured to acquire a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between first-level labels and second-level labels; A first training module, configured to input the training data to be trained and the preset relationship dependency matrix into a model to be trained, so as to obtain a first-level label prediction matrix and a second-level label prediction matrix; A first determination module, configured to determine a target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix, where each element in the first-level label prediction matrix is a real number, and the first-level label prediction matrix is integerized so that each element in the first-level label prediction matrix changes from a real number to an integer, obtaining a first-level label integer matrix, where each value of the elements in the first-level label integer matrix is 0 or 1, and the first-level label integer matrix is cross-multiplied by the preset relationship dependency matrix to obtain the target relationship dependency matrix; A second determination module, configured to determine a target loss value according to the first-level label prediction matrix, the first-level label reference matrix, a target matrix, and the second-level label reference matrix, where the target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix, where the target relationship dependency matrix is added to a preset value to obtain a first matrix, and the second-level label prediction matrix is dot-multiplied by the first matrix to obtain the target matrix; determining a first loss value according to the first-level label prediction matrix and the first-level label reference matrix; determining a second loss value according to the target matrix and the second-level label reference matrix; determining a target loss value according to the first loss value and the second loss value; A second training module, configured to obtain a target loss value every time a batch of training data to be trained is trained, and every time a target loss value is obtained, the target loss value is fed back to the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirming that the model training ends, and obtaining a trained model.

8. A classification device wherein it includes A third acquisition module, configured to acquire a text to be classified; A prediction module, configured to input the text to be classified into the trained model, where the trained model is trained according to the model construction method described in any one of claims 1 to 5, to obtain a first-level label probability matrix and a second-level label prediction probability matrix, where each element in the first-level label probability matrix corresponds to a first-level label, and each element in the first-level label probability matrix is a real number, and each element in the second-level label prediction probability matrix corresponds to a second-level label, and each element in the second-level label prediction probability matrix is a real number; A third determination module, configured to determine a first-level label corresponding to the text to be classified according to the first-level label probability matrix, where the first-level label corresponding to the element with the largest value in the first-level label probability matrix is the first-level label corresponding to the text to be classified; An integer conversion module, configured to perform integer conversion on the first-level label probability matrix, so that each element in the first-level label probability matrix changes from a real number to an integer, to obtain a first-level label integer matrix, where the value of each element in the first-level label integer matrix is 0 or 1; A fourth determination module, configured to determine a first relationship dependency matrix according to the first-level label integer matrix and a preset relationship dependency matrix; A fifth determination module, configured to determine a second-level label probability matrix according to the first relationship dependency matrix and the second-level label prediction probability matrix, where each element in the second-level label probability matrix corresponds to a second-level label; A sixth determination module, configured to determine a second-level label corresponding to the text to be classified according to the second-level label probability matrix, where the second-level label corresponding to the element with the largest value in the second-level label probability matrix is the second-level label corresponding to the text to be classified.

9. A storage medium, wherein, a computer program is stored in the storage medium, and when the computer program runs on a computer, the computer is caused to execute the model construction method according to any one of claims 1 to 5 or the classification method according to claim 6.

10. An electronic device, wherein, the electronic device includes a processor and a memory, a computer program is stored in the memory, and the processor is configured to execute, by calling the computer program stored in the memory: obtain training data to be trained, where the training data to be trained includes a token matrix composed of encodings corresponding to tokens obtained by segmenting each text in multiple texts, a first-level label reference matrix composed of encodings corresponding to first-level labels corresponding to each text, and a second-level label reference matrix composed of encodings corresponding to second-level labels corresponding to each text; obtain a preset relationship dependency matrix, where the preset relationship dependency matrix is used to represent the hierarchical relationship between first-level labels and second-level labels; input the training data to be trained and the preset relationship dependency matrix into a model to be trained, so as to obtain a first-level label prediction matrix and a second-level label prediction matrix; determine a target relationship dependency matrix according to the first-level label prediction matrix and the preset relationship dependency matrix, where each element in the first-level label prediction matrix is a real number, perform integer conversion on the first-level label prediction matrix, so that each element in the first-level label prediction matrix changes from a real number to an integer, to obtain a first-level label integer matrix, where the value of each element in the first-level label integer matrix is 0 or 1, and multiply the first-level label integer matrix by the preset relationship dependency matrix to obtain the target relationship dependency matrix; Determine a target loss value according to the first-level label prediction matrix, the first-level label reference matrix, the target matrix, and the second-level label reference matrix. The target matrix is determined according to the second-level label prediction matrix and the target relationship dependency matrix. Specifically, add the target relationship dependency matrix to a preset value to obtain a first matrix, and multiply the second-level label prediction matrix by the first matrix to obtain the target matrix; determine a first loss value according to the first-level label prediction matrix and the first-level label reference matrix; determine a second loss value according to the target matrix and the second-level label reference matrix; determine the target loss value according to the first loss value and the second loss value. Whenever a batch of data to be trained is trained to obtain a target loss value, each time a target loss value is obtained, the target loss value is fed back into the model to be trained to adjust the parameters of the model to be trained until the model to be trained converges, confirming that the model training is completed, and obtaining the trained model.

11. The electronic device according to claim 10, wherein, the processor is configured to execute: Determine a first loss value according to the first-level label prediction matrix and the first-level label reference matrix; Determine the target matrix according to the second-level label prediction matrix and the target relationship dependency matrix; Determine a second loss value according to the target matrix and the second-level label reference matrix; Determine the target loss value according to the first loss value and the second loss value.

12. The electronic device according to claim 11, wherein, the processor is configured to execute: Multiply the first loss value by a first weight value to obtain a third loss value; Multiply the second loss value by a second weight value to obtain a fourth loss value, where the second weight value is less than the first weight value; Determine the target loss value according to the third loss value and the fourth loss value.

13. The electronic device according to claim 11, wherein, the processor is configured to execute: Add the target relationship dependency matrix to a preset value to obtain a first matrix; Multiply the second-level label prediction matrix by the first matrix to obtain the target matrix.

14. The electronic device according to claim 10, wherein, each element in the first-level label prediction matrix is a real number, and the processor is configured to execute: Perform integerization processing on the first-level label prediction matrix so that each element in the first-level label prediction matrix changes from a real number to an integer, obtaining a first-level label integer matrix, where the value of each element in the first-level label integer matrix is 0 or 1; Multiply the first-level label integer matrix by the preset relationship dependency matrix to obtain the target relationship dependency matrix.

15. The electronic device according to claim 10, wherein, the processor is configured to execute: Obtain multiple texts, the first-level labels corresponding to each text, and the second-level labels corresponding to each text; Perform word segmentation processing on each text to obtain the word segmentation corresponding to each text; Determine the encoding corresponding to the word segmentation of each text; Determine a word segmentation matrix according to the encoding corresponding to the word segmentation of each text; Perform one-hot encoding processing on the first-level labels corresponding to each text to obtain the encoding corresponding to the first-level labels of each text; Determine the first-level label reference matrix according to the codes corresponding to the first-level labels of each text; Perform one-hot encoding on the second-level labels corresponding to each text to obtain the codes corresponding to the second-level labels of each text; Determine the second-level label reference matrix according to the codes corresponding to the second-level labels of each text; Determine the data to be trained according to the word segmentation matrix, the first-level label reference matrix, and the second-level label reference matrix.

16. The electronic device according to claim 15, wherein, the processor is configured to execute: Construct a dictionary according to the word segmentation of each text, the dictionary including a plurality of word segments and their corresponding codes; Determine the codes corresponding to the word segments of each text according to the word segmentation of each text and the dictionary.

17. An electronic device, wherein, the electronic device includes a processor and a memory, a computer program is stored in the memory, and the processor is configured to execute, by calling the computer program stored in the memory: Obtain the text to be classified; Input the text to be classified into the trained model, the trained model being trained according to the model construction method described in any one of claims 1 to 5, to obtain a first-level label probability matrix and a second-level label prediction probability matrix, each element in the first-level label probability matrix corresponding to a first-level label, all elements in the first-level label probability matrix being real numbers, each element in the second-level label prediction probability matrix corresponding to a second-level label, all elements in the second-level label prediction probability matrix being real numbers; Determine the first-level label corresponding to the text to be classified according to the first-level label probability matrix, the first-level label corresponding to the element with the largest value in the first-level label probability matrix being the first-level label corresponding to the text to be classified; Perform integerization processing on the first-level label probability matrix so that each element in the first-level label probability matrix changes from a real number to an integer, to obtain a first-level label integerization matrix, the value of the element in the first-level label integerization matrix being 0 or 1; Determine a first relationship dependency matrix according to the first-level label integerization matrix and a preset relationship dependency matrix; Determine a second-level label probability matrix according to the first relationship dependency matrix and the second-level label prediction probability matrix, each element in the second-level label probability matrix corresponding to a second-level label; Determine the second-level label corresponding to the text to be classified according to the second-level label probability matrix, the second-level label corresponding to the element with the largest value in the second-level label probability matrix being the second-level label corresponding to the text to be classified.

Citation Information

Patent Citations

  • Text classification model construction method and device, terminal and storage medium

    CN109960726A

  • Short text labeling method, system and device for large-scale classification system

    CN110059181A