Text classification method, apparatus, device, storage medium, and program product

By generating text vectors and combining the correlation and citation relationships between texts, the problem of low classification accuracy in existing text classification methods is solved, and more accurate text classification is achieved.

CN116010604BActive Publication Date: 2026-06-02SHENYANG NEUSOFT INTELLIGENT MEDICAL TECH RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENYANG NEUSOFT INTELLIGENT MEDICAL TECH RES INST
Filing Date
2023-01-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing text classification methods suffer from low classification accuracy and are unable to effectively distinguish texts with low relevance.

Method used

By acquiring textual information from multiple texts and the categories of some texts, a machine learning model is used to generate text vectors. Combining the correlation and citation relationships between texts, the category of the text is determined.

Benefits of technology

It improves the accuracy of text classification, ensuring that highly relevant texts are classified into the same category, thus enhancing classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116010604B_ABST
    Figure CN116010604B_ABST
Patent Text Reader

Abstract

The application provides a text classification method, device, equipment, storage medium and program product. The method comprises the following steps: obtaining text information of N texts respectively and categories of part of the N texts, N being an integer greater than 1; determining a first vector of each of the N texts based on the text information of each of the N texts, wherein the first vector of a target text is used to represent the correlation between the target text and other texts in the N texts except the target text, and the target text is any one of the N texts; and determining categories of the remaining texts in the N texts except the part of the texts based on the first vector of each of the N texts and the categories of the part of the texts. Through the above technical solution, the accuracy of text classification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a text classification method, apparatus, device, storage medium, and program product. Background Technology

[0002] Currently, papers are often categorized based on keywords, so that researchers can retrieve relevant papers based on keywords when studying a certain type of problem.

[0003] However, the above classification method has the problem of low classification accuracy. For example, according to the above classification method, the three papers "Research Progress in Clinical Diagnosis and Treatment of Gastric Cancer", "Application of Artificial Intelligence in Gastric Cancer Imaging" and "Impact of Payment on Hospitalized Gastric Cancer Patients" can be classified into the same category based on the keyword "gastric cancer". However, "Research Progress in Clinical Diagnosis and Treatment of Gastric Cancer" mainly studies the treatment of gastric cancer, "Application of Artificial Intelligence in Gastric Cancer Imaging" mainly studies the application of artificial intelligence in gastric cancer, and "Research on Hospitalization Costs and Influencing Factors of Gastric Cancer Patients" mainly studies the hospitalization costs of gastric cancer. Obviously, these three papers mainly study different issues, that is, the correlation between the three papers is not high. In other words, the classification accuracy of the above classification method is low. Summary of the Invention

[0004] This application provides a text classification method, apparatus, device, storage medium, and program product that can improve the accuracy of text classification.

[0005] Firstly, a text classification method is provided, comprising: obtaining text information of N texts and categories of some texts among the N texts, where N is an integer greater than 1; determining a first vector for each of the N texts based on the text information of each of the N texts, wherein the first vector of the target text is used to characterize the correlation between the target text and other texts among the N texts excluding the target text, and the target text is any text among the N texts; and determining the categories of the remaining texts among the N texts excluding the remaining texts based on the first vectors of each of the N texts and the categories of some texts.

[0006] Secondly, a text classification device is provided, comprising: an acquisition module, a first determination module, and a second determination module, wherein the acquisition module is used to acquire text information of N texts and the categories of some texts among the N texts, where N is an integer greater than 1; the first determination module is used to determine a first vector of each of the N texts based on the text information of each of the N texts, wherein the first vector of the target text is used to characterize the correlation between the target text and other texts among the N texts excluding the target text, and the target text is any text among the N texts; the second determination module is used to determine the category of the remaining texts among the N texts excluding some texts based on the first vectors of each of the N texts and the categories of some texts.

[0007] Thirdly, an electronic device is provided, comprising: a processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory, and performing the methods as described in the first aspect or its various implementations.

[0008] Fourthly, a computer-readable storage medium is provided for storing a computer program that causes a computer to perform the methods described in the first aspect or its various implementations.

[0009] Fifthly, a computer program product is provided, including computer program instructions that cause a computer to perform the methods as described in the first aspect or its various implementations.

[0010] Sixthly, a computer program is provided that causes a computer to perform the methods described in the first aspect or its various implementations.

[0011] Through the technical solution of this application, an electronic device can acquire the text information of N texts and the categories of some texts among the N texts, where N is an integer greater than 1. Then, the electronic device can determine the first vector of each of the N texts based on their respective text information. The first vector of the target text is used to characterize the correlation between the target text and any of the other texts among the N texts, where the target text is any one of the N texts. Finally, the electronic device can determine the category of the remaining texts among the N texts, excluding some texts, based on their respective first vectors and the categories of some texts. In other words, this application considers the correlation between two texts and the known categories of some texts when performing text classification, classifying two highly correlated texts into the same category. This ensures that two texts belonging to the same category have a high correlation, thus making the determined text categories more accurate and improving the precision of text classification. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 An application scenario diagram provided for an embodiment of this application;

[0014] Figure 2 A flowchart illustrating a text classification method provided in this application embodiment;

[0015] Figure 3 A flowchart illustrating another text classification method provided in this application embodiment;

[0016] Figure 4 A flowchart illustrating another text classification method provided in this application embodiment;

[0017] Figure 5 A schematic diagram illustrating a text classification method provided in an embodiment of this application;

[0018] Figure 6 A schematic diagram illustrating another text classification method provided in an embodiment of this application;

[0019] Figure 7 A schematic diagram illustrating another text classification method provided in an embodiment of this application;

[0020] Figure 8 A flowchart illustrating yet another text classification method provided in this application embodiment;

[0021] Figure 9 A schematic diagram of a system architecture provided for an embodiment of this application;

[0022] Figure 10 A schematic diagram of a text classification device 1000 provided in an embodiment of this application;

[0023] Figure 11 This is a schematic block diagram of the electronic device 1100 provided in the embodiments of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0026] As mentioned above, current text classification methods often classify texts based on keywords. This method may result in low relevance among multiple texts classified into the same category, meaning that the classification accuracy of this method is low.

[0027] To solve the above-mentioned technical problems, the inventive concept of this application is: an electronic device can determine the category of the remaining texts in multiple texts, excluding some texts, based on the category of some texts in multiple texts and the correlation between two texts in multiple texts.

[0028] It should be understood that the technical solution of this application can be applied to the following scenarios, but is not limited to:

[0029] In some possible ways, Figure 1 An application scenario diagram provided for an embodiment of this application, such as... Figure 1 As shown, this application scenario may include electronic device 110 and network device 120. Electronic device 110 can establish a connection with network device 120 through a wired network or a wireless network.

[0030] For example, electronic device 110 may be a desktop computer, laptop computer, tablet computer, etc., but is not limited to these. Network device 120 may be a terminal device or a server, but is not limited to these. In this embodiment of the application, electronic device 110 may send a request message to network device 120, the request message being used to request the acquisition of text information of each of N texts and the category of some texts among the N texts, where N is an integer greater than 1. Further, electronic device 110 may receive a response message sent by network device 120, the response message including the text information of each of the N texts and the category of some texts.

[0031] also, Figure 1 An electronic device and a network device are given as examples, but in practice, other numbers of electronic devices and network devices may be included, and this application does not limit this.

[0032] In other possible implementations, the technical solution of this application may also be executed by the aforementioned electronic device 110, or by the aforementioned network device 120, and this application does not impose any restrictions on this.

[0033] After introducing the application scenarios of the embodiments of this application, the technical solution of this application will be described in detail below:

[0034] Figure 2 A flowchart illustrating a text classification method provided in this application embodiment, the method can be derived from, for example... Figure 1 The electronic device 110 shown performs, but is not limited to, its functions. For example... Figure 2 As shown, the method may include the following steps:

[0035] S210: Obtain the text information of each of the N texts and the category of some of the texts in the N texts, where N is an integer greater than 1;

[0036] S220: Determine the first vector of each of the N texts based on their respective text information, wherein the first vector of the target text is used to characterize the correlation between the target text and the other texts in the N texts excluding the target text, and the target text is any one of the N texts;

[0037] S230: Based on the first vector of each of the N texts and the category of the partial text, determine the category of the remaining texts among the N texts excluding the partial texts.

[0038] In some possible implementations, the electronic device can first obtain N texts from a terminal device or a network device such as a server, and then the electronic device can identify and extract the N texts respectively to obtain the text information of each of the N texts; or, the network device can first identify and extract the N texts respectively to obtain the text information of each of the N texts, and then send the text information of each of the N texts to the electronic device, so that the electronic device can obtain the text information of each of the N texts. This application does not limit this.

[0039] For example, suppose N texts are N papers, each including an abstract and references. After the electronic device obtains the N papers from the terminal device, it can first identify the abstracts in the N papers to determine the location of each abstract. Then, it can extract the abstracts of each paper based on their locations. Similarly, the electronic device can also identify the references in the N papers to determine the location of each reference. Then, it can extract the references of each paper based on their locations. Thus, the electronic device can obtain the abstracts and references of each of the N papers, and can use the abstracts and references of each of the N papers as the text information of each paper.

[0040] In some possible implementations, the user of the electronic device can input the categories of some texts out of N texts into the electronic device, so that the electronic device can obtain the categories of some texts; or, the electronic device can send a request message to a terminal device or a network device such as a server to obtain the categories of some texts out of N texts. After receiving the request message, the network device can respond to the request message and send the categories of the aforementioned partial texts to the electronic device. The categories of the partial texts in the network device can be obtained based on manual classification, and this application does not impose any restrictions on this.

[0041] For example, a user of an electronic device can classify three of the five texts into categories related to gastric cancer treatment, health economics related to gastric cancer, and artificial intelligence applications related to gastric cancer, based on manual classification. Then, the user can input the categories of the three texts into the electronic device, so that the electronic device can obtain the categories of the three texts.

[0042] In some implementations, the text information may include: citation information and text content. Figure 3 A flowchart illustrating another text classification method provided in an embodiment of this application. Based on... Figure 2 ,like Figure 3 As shown, the above S220 includes:

[0043] S310: Vectorize the text content of each of the two texts included in the target text pair to obtain the second vector of each of the two texts. The target text pair is any text pair among N texts.

[0044] S320: Determine the first probability of the target text pair based on the initial vectors corresponding to the second vectors of the two texts and the first vectors of the two texts, respectively;

[0045] S330: Determine the citation weight of the target text pair based on the citation information of each of the two texts;

[0046] S340: Determine the similarity of the target text pair based on the second vector of each of the two texts;

[0047] S350: Determine the second probability corresponding to the target text pair based on the citation weight of the target text pair, the similarity of the target text pair, the citation weight of all text pairs in N texts, and the similarity of all text pairs in N texts;

[0048] S360: Determine the first vector of each of the N texts based on the first and second probabilities of all text pairs in the N texts.

[0049] The citation information of the text can be used to indicate how the text references other texts. For example, suppose text 1 references text 2 and text 3, and text 1 references text 2 and text 3 5 times and 1 time respectively. The identification information of text 2 is the name of text 2: text 2, and the identification information of text 3 is the name of text 3: text 3. Then the citation information of text 1 can be: text 1, text 2, or it can be: text 1, 5, text 2, 1. This application does not limit this.

[0050] The text content may include at least one of the following, but is not limited to: a summary of the text, a title of the text, and the body of the text.

[0051] In some implementations, electronic devices can use machine learning models to vectorize the text content of each of the two texts in a target text pair, resulting in a second vector for each text. The machine learning model can be a word-to-vector (word2vec) model or a document-to-vector (doc2vec) model, but is not limited to these.

[0052] In the following embodiments, this application will describe the first probability of an electronic device determining a target text pair:

[0053] It is understood that the initial vector corresponding to the first vector is the initial form of the first vector. Electronic devices can randomly set the initial vector; for example, electronic devices can randomly set the initial vector to a low-dimensional vector, and this application does not impose any restrictions on this.

[0054] In some possible implementations, S320 may include: the electronic device calculates the product of the initial vectors corresponding to the two texts to obtain a first product, calculates the product of the second vectors corresponding to the two texts to obtain a second product, calculates the maximum value of the modulus of the second vectors corresponding to the two texts, and finally, the electronic device may determine the first probability corresponding to the target text pair based on the first product, the second product, and the maximum value.

[0055] For example, when determining the first probability corresponding to a target text pair based on a first product, a second product, and a maximum value, the electronic device can first calculate the hyperbolic tangent of the first product and then calculate the ratio of the second product to the maximum value to obtain a first numerical value. Finally, the electronic device can determine the first probability corresponding to the target text pair based on the hyperbolic tangent and the first numerical value. For instance, the electronic device can first normalize the hyperbolic tangent to obtain a second numerical value, and then determine the first probability corresponding to the target text pair based on the first and second numerical values. For example, the electronic device can obtain the first probability corresponding to the target text pair by calculating the product of the first and second numerical values. The normalization of the hyperbolic tangent means adjusting its range to between 0 and 1. It can be understood that the probability range is 0-1, so normalizing the hyperbolic tangent makes the determined first probability more reasonable.

[0056] It's understandable that the dependent variable of the hyperbolic tangent function increases with the increase of the independent variable. Therefore, when using the product of the initial vectors corresponding to the two texts in the target text pair (i.e., the first product) as the independent variable of the hyperbolic tangent function, a larger first product results in a larger hyperbolic tangent value for the dependent variable (the first product). Furthermore, a larger product of two vectors indicates higher similarity. Thus, the higher the similarity between the initial vectors corresponding to the two texts in the target text pair, the larger the corresponding hyperbolic tangent value. The initial vector is the initial form of the first vector, which can be used to characterize the correlation between the target text and other texts among N texts. In other words, the initial vector can also be used to characterize the correlation between the target text and other texts among N texts. Therefore, for two texts in the target text pair, the closer their correlation with each other among the N texts, the more similar their initial vectors are, and the larger the corresponding hyperbolic tangent value of the target text pair. Thus, the first probability determined based on the hyperbolic tangent value can characterize the correlation between the two texts in the target text pair.

[0057] The second vector of a text is the vectorized representation of the text content by an electronic device. In other words, the second vector of a text can be used to characterize the text content. The larger the product of the two vectors, the higher their similarity. Therefore, the larger the product of the second vectors of the two texts in the target text pair, the more similar the second vectors of the two texts are, that is, the more similar the text content of the two texts are. So the first probability determined based on the second product can characterize the correlation between the two texts in the target text pair.

[0058] Since the ratio of the second product to the aforementioned maximum value (i.e., the size of the first value) is smaller than the second product, the first probability determined based on the hyperbolic tangent and the first value will be smaller. This reduces the computational load on the electronic device when determining the first vector based on the first probability, thus improving the processing efficiency of the electronic device. Alternatively, the electronic device can first calculate the minimum modulus of the second vectors of the two texts, and then obtain the first value by calculating the ratio of the second product to the minimum value; or, the electronic device can first calculate the product of the moduli of the second vectors of the two texts, and then obtain the first value by calculating the ratio of the second product to the aforementioned product. This application does not impose any restrictions on this approach.

[0059] For example, suppose vi is the i-th text in N texts, and vj is the j-th text in N texts. It is the initial vector of the first vector of the i-th text. It is the initial vector of the first vector of the j-th text. It is the second vector of the i-th text. If is the second vector of the j-th text, then the first probability corresponding to the target text pair composed of vi and vj can be expressed as shown in formula (1):

[0060]

[0061] In formula (1), Yes This refers to the normalization of the hyperbolic tangent of the first product. It can be understood that the range of the hyperbolic tangent is -1 to 1, and the range of the sum of the hyperbolic tangent and 1 is 0-2. This sum is then normalized... Multiplying them together yields a second value in the range of 0-1.

[0062] In other possible implementations, the electronic device can directly determine the hyperbolic tangent of the first product as the first probability; or, the electronic device can first set a target coefficient and determine the product of the target coefficient and the hyperbolic tangent as the first probability; or, the electronic device can also determine the first probability based on other function values ​​of the first product. For example, the electronic device can use the first product as the exponent of an exponential function with base e to obtain the exponent value of the target text pair, and then determine the first probability based on the exponent value of the target text pair. The specific determination method is similar to that described above based on the hyperbolic tangent of the first product, and will not be elaborated here.

[0063] It should be noted that this application does not limit the method for determining the first probability of an electronic device.

[0064] In the following embodiments, this application will describe the second probability for an electronic device to determine the target text pair:

[0065] In some possible implementations, the electronic device can first determine the number of citations between the two texts in the target text pair based on their respective citation information, and then determine the citation weight of the target text pair based on the number of citations between the two texts. The number of citations between the two texts can be the first number of times the first text cites the second text, the second number of times the second text cites the first text, the average, maximum, or minimum of the first and second citations. Alternatively, when there is a citation relationship between the two texts (i.e., the first text cites the second text and / or the second text cites the first text), the electronic device can determine the number of citations between the two texts as a first preset value. When there is no citation relationship between the two texts (i.e., neither the first nor the second text cites the first text), the electronic device can determine the number of citations between the two texts as a second preset value. The first preset value is greater than or equal to the second preset value. The first preset value can be 1, and the second preset value can be 0. The electronic device can determine the citation weight of the target text pair by the number of citations between the two texts, or by the normalized result of the number of citations between the two texts. This application does not impose any restrictions on this.

[0066] In some possible implementations, the electronic device can determine the similarity of the target text pair by the product of the second vectors of the two texts, or by the distance between the second vectors of the two texts, such as the Euclidean distance. This application does not limit this.

[0067] In some possible implementations, S350 may include: the electronic device calculating the sum of the reference weight of the target text pair and the similarity of the target text pair to obtain a third value, and calculating the sum of the reference weight of all text pairs in N texts and the similarity of all text pairs in N texts to obtain a fourth value. Then, the electronic device may determine the second probability corresponding to the target text pair based on the third value and the fourth value. For example, the electronic device may calculate the ratio of the third value to the fourth value to obtain the second probability corresponding to the target text pair.

[0068] It is understandable that the citation weight of a target text pair can characterize the citation relationship between the two texts in the target text pair, and the citation relationship between the two texts can characterize the relevance between the two texts. For example, the greater the number of citations between two texts, the greater the relevance between the two texts. Therefore, the citation weight of a target text pair can characterize the relevance between the two texts in the target text pair. When the similarity between two texts is higher, their relevance is higher, so the similarity of a target text pair can characterize the relevance between the two texts in the target text pair. Therefore, the second probability corresponding to the target text pair determined by the above method can characterize the relevance between the two texts in the target text pair.

[0069] For example, suppose vi is the i-th text in N texts, vj is the j-th text in N texts, the i-th text and the j-th text can form a target text pair, ,j is the reference weight between the i-th text and the j-th text, i.e. the reference weight of the target text pair, and ti,j is the similarity between the i-th text and the j-th text, i.e. the similarity of the target text pair. Then the second probability corresponding to the target text pair formed by vi and vj can be as shown in formula (2):

[0070]

[0071] In other possible implementations, the electronic device can determine a second probability based on the citation weight of the target text pair and the citation weight of all text pairs in the N texts. For example, the electronic device can calculate the ratio of the citation weight of the target text pair to the citation weight of all text pairs in the N texts to obtain the second probability corresponding to the target text pair; alternatively, the electronic device can also determine the second probability based on the similarity of the target text pair and the similarity of all text pairs in the N texts. For example, the electronic device can calculate the ratio of the similarity of the target text pair to the similarity of all text pairs in the N texts to obtain the second probability corresponding to the target text pair; alternatively, the electronic device can also perform a weighted summation of the citation weight of the target text pair and the similarity of the target text pair to obtain a third value, perform a weighted summation of the citation weight of all text pairs in the N texts and the similarity of all text pairs in the N texts to obtain a fourth value, and then calculate the ratio of the third value to the fourth value to obtain the second probability corresponding to the target text pair.

[0072] It should be noted that this application does not limit the method for determining the second probability of electronic devices.

[0073] In the following embodiments, this application will describe S360:

[0074] In some implementations, S360 may include: the electronic device can first determine a first probability distribution corresponding to the first probability of all text pairs in N texts, and determine a second probability distribution corresponding to the second probability of all text pairs in N texts. Then, the electronic device can calculate the difference between the first probability distribution and the second probability distribution. If the difference between the first probability distribution and the second probability distribution is less than a preset difference, then the initial vector of each of the N texts is determined as the first vector of each of the N texts. If the difference between the first probability distribution and the second probability distribution is greater than or equal to the preset difference, then the initial vector of each of the N texts is adjusted, and the process of determining the first probability corresponding to the target text pair based on the second vector of each of the two texts and the initial vector corresponding to the first vector of each of the two texts continues until the difference between the first probability distribution and the second probability distribution is less than the preset difference, and the first vector of each of the N texts is obtained.

[0075] The first probability distribution consists of the first probabilities of all text pairs in N texts, and the second probability distribution consists of the second probabilities of all text pairs in N texts. Electronic devices can use metrics that measure the similarity, or the degree of difference, between the two probability distributions to represent the difference between the first and second probability distributions. Metrics that measure the degree of difference between two probability distributions can include, but are not limited to, KL divergence (KLD), JS divergence (JSD), cross entropy, and Wasserstein distance.

[0076] For example, when the difference between the first probability distribution and the second probability distribution is greater than or equal to a preset difference, the electronic device can adjust the initial vectors of each of the N texts based on the stochastic gradient descent method. For instance, the electronic device can calculate the gradient of the difference between the first probability distribution and the second probability distribution, and adjust the initial vectors of each of the N texts according to the direction of gradient descent.

[0077] It is understandable that since both the second and first probabilities corresponding to a target text pair can characterize the correlation between the two texts in the target text pair, the first and second probability distributions can also characterize the correlation between any two texts in N texts. The second probability distribution, determined based on the citation weights and similarities of the text pairs, can more accurately characterize the correlation between any two texts in N texts. Therefore, when the difference between the first and second probability distributions is less than a preset difference, the electronic device can determine that the first probability distribution, determined based on the initial vectors of each of the N texts, can more accurately characterize the correlation between any two texts in N texts. Thus, it can be determined that the initial vector of each text in N texts can more accurately characterize the correlation between that text and the other texts in N texts. The electronic device determines the initial vectors of each of the N texts as their respective first vectors. When the difference between the first probability distribution and the second probability distribution is greater than or equal to a preset difference, the electronic device determines that the first probability distribution determined based on the initial vectors of the N texts cannot accurately represent the correlation between any two texts in the N texts. Therefore, it can be determined that the initial vector of each text in the N texts cannot accurately represent the correlation between that text and the other texts in the N texts. The electronic device then adjusts the initial vectors of the N texts based on the difference between the first and second probability distributions and redetermines the first probability distribution until the difference between the first and second probability distributions is less than the preset difference, thus obtaining the first vectors of each of the N texts. In this way, it can be ensured that the first vector of each text in the N texts can accurately represent the correlation between that text and the other texts in the N texts, thereby accurately determining the category of the remaining texts in the N texts based on their respective first vectors, further improving the text classification accuracy.

[0078] In some implementations, when determining the first vector of each of the N texts based on their respective textual information, the electronic device can also determine the first vector of each of the N texts based on their respective citation weights and / or similarities. Here, the citation weight of the target text is the citation weight of the text pair including the target text, and the similarity of the target text is the similarity of the text pair including the target text. The target text is any one of the N texts. The method for determining the citation weight or similarity of the text pair is similar to that described in the above embodiments, and will not be elaborated upon here.

[0079] For example, suppose there are N texts, specifically three texts: text 1, text 2, and text 3. An electronic device determines the reference weights of the text pairs consisting of text 1 and text 2, text 1 and text 3, and text 2 and text 3 as w12, w13, and w23, respectively. It also determines the similarities of these text pairs as t12, t13, and t23, respectively. Therefore, the electronic device can determine that the reference weight of text 1 is w12 and w13, the similarity of text 1 is t12 and t13, and the reference weight of text 2 is w12 and w13. 2. The similarity of text 2 is t12, t23, and the reference weight of text 3 is w13, w23. The similarity of text 3 is t13, t23. Then, the electronic device can determine that the first vectors of text 1, text 2, and text 3 are (w12, w13, t12, t13), (w12, w23, t12, t23), (w13, w23, t13, t23), or (w12, w13), (w12, w23), (w13, w23), or (t12, t13), (t12, t23), (t13, t23).

[0080] It should be noted that this application does not limit the method by which an electronic device determines the first vector of each of the N texts.

[0081] Figure 4 A flowchart illustrating yet another text classification method provided in an embodiment of this application. Based on Figure 2 ,like Figure 4 As shown, the above S230 includes:

[0082] S410: Generate a first matrix based on the first vectors of N texts;

[0083] S420: Input the first matrix and the categories of a portion of the text into the first neural network to obtain the categories of the remaining text.

[0084] In some possible implementations, the electronic device can first generate a graph structure of N texts, which includes N nodes. The N nodes correspond one-to-one with the N texts. When any two texts have a reference relationship, the two nodes corresponding to the two texts are connected. Then, for any node among the N nodes, the electronic device can perform a random walk starting from that node to obtain a node sequence. Finally, the electronic device can combine the first vectors of the N texts according to the order of the node sequence to obtain a first matrix.

[0085] For example, suppose there are N texts, specifically four texts: text1, text2, text3, and text4. These four texts reference each other. The graph structure used by the electronic device to generate these four texts could be as follows: Figure 5 The graph structure is shown in the image. This application uses node 1 corresponding to text 1 as an example to describe the process of obtaining a node sequence by randomly selecting node 1 as the initial node. The methods for obtaining other node sequences using node 1 as the initial node and other node sequences as initial nodes are similar and will not be elaborated upon here. The electronic device can first determine the length of the node sequence, such as 6, and then the electronic device can randomly determine a direction, such as... Figure 5 The electronic device can then traverse along the directions indicated by dashed arrows 1-5, i.e., the edges corresponding to dashed arrows 1-5, to obtain the node sequence (v1, v4, v3, v1, v2, v4). Assuming the first vectors of text 1, text 2, text 3, and text 4 are (N11, ..., N16), (N21, ..., N26), (N31, ..., N36), and (N41, ..., N46) respectively, then according to the node order in the node sequence (v1, v4, v3, v1, v2, v4), the electronic device combines the first vectors of each of the four texts to obtain the first matrix:

[0086]

[0087] In other possible implementations, the electronic device can also directly generate a first matrix based on the first vectors of each of the N texts. For example, assuming the first vectors of text 1, text 2, text 3, and text 4 are (N11, ..., N16), (N21, ..., N26), (N31, ..., N36), and (N41, ..., N46) respectively, the electronic device can directly combine the first vectors of the four texts to obtain the first matrix:

[0088]

[0089] In some implementations, the first neural network can be a multiscale convolutional neural network (MCNN), which comprises M convolutional neural networks (CNNs). At least two CNNs have at least one pair of convolutional layers with corresponding kernel sizes that differ in size, where M is an integer greater than 1. For example, if M is 3... Figure 6As shown, the architecture of each of the three CNNs is as follows: convolutional layer 1, pooling layer 1, convolutional layer 2, and pooling layer 2. Among them, convolutional layer 1 in the first CNN and convolutional layer 1 in the second CNN are a pair of convolutional layers at corresponding positions. Convolutional layer 2 in the first CNN and convolutional layer 2 in the second CNN are a pair of convolutional layers at corresponding positions. Convolutional layer 1 in the first CNN and convolutional layer 1 in the third CNN are a pair of convolutional layers at corresponding positions. Other pairs of convolutional layers at corresponding positions are similar to the above pairs of convolutional layers at corresponding positions, and will not be described in detail here.

[0090] It is understandable that larger convolutional kernels can extract global features of the first matrix, while smaller convolutional kernels can extract local features of the first matrix. When at least two CNNs have at least one pair of convolutional layers with different kernel sizes at corresponding positions, the features of the first matrix extracted by MCNN can include both local and global features. This allows the features of the first matrix to be represented from multiple perspectives, making the extracted features of the first matrix more accurate and thus improving the accuracy of text classification.

[0091] In some implementations, the dimension of the first vector of each of the N texts is the same as the width of any convolutional kernel in the CNN. In this way, when the convolutional layer of the CNN extracts features from the first matrix, it can extract all the features of the first vector of at least one text. That is, each convolution operation can be performed on at least one complete first vector, so that the extracted features can more accurately represent the relevant information between the texts corresponding to at least one text, thereby improving the accuracy of text classification.

[0092] For example, assuming the dimension of the first vector of each of the N texts is s, then the size of each convolution kernel of each CNN can be h*s, where h is the height of the convolution kernel, s is the width of the convolution kernel, s is a positive integer, and h is a positive integer. This application does not impose any restrictions on the size of h and s.

[0093] In some possible implementations, such as Figure 7 As shown, assuming there are N texts, there are 4 texts: text1, text2, text3, and text4. The node sequence and the first vector of the 4 texts are as follows: Figure 7As shown, each first vector has a dimension of 6. The electronic device can first generate a first matrix based on the node sequence and four texts. Then, the electronic device can input the first matrix into an MCNN, which includes three CNNs and one fully connected layer. Each of the three CNNs includes one convolutional layer and one pooling layer. The kernel sizes of the three convolutional layers are 3*6, 4*6, and 5*6, respectively. It can be seen that the width of the three convolutional kernels is an integer multiple of the dimension of the first vector, and the kernel sizes are different. The electronic device can extract features from the first matrix based on the three CNNs to obtain three features of the first matrix. Then, the electronic device can perform pooling operations on the above three features based on the pooling layers of the three CNNs. The pooling operation can be max pooling, but is not limited to this. Then, the electronic device can redistribute the weights of the pooled results based on the fully connected layer to obtain the final features of the four text pairs. The final features can represent the correlation between each pair of texts in the four text pairs. Finally, the electronic device can determine the category of the remaining texts based on the final features and the categories of some texts using the classification module. For example, suppose that some texts are categorized into categories: the category of text 1 and the category of text 2. The electronic device determines, based on the final features, that text 3 has the highest relevance to text 1 and text 4 has the highest relevance to text 2. Then the classification module can assign the category of text 1 to the category of text 3 and the category of text 2 to the category of text 4. Thus, the electronic device can determine the category of the remaining texts.

[0094] It is understood that in the above embodiments, this application uses the method of determining the remaining text categories based on a first matrix generated by an electronic device based on a node sequence as an example to describe the method of determining the remaining text categories by an electronic device. The method of determining the text categories based on a first matrix generated by an electronic device based on any other node sequence is similar to this method, and will not be described in detail here. If the electronic device determines multiple node sequences, the electronic device can determine the first matrix corresponding to each node sequence through the above method, and input each first matrix into the above MCNN respectively. Then, the electronic device can obtain the features corresponding to each first matrix based on the MCNN. Then, the electronic device can fuse the features corresponding to each first matrix based on the MCNN to obtain fused features. Finally, the electronic device can determine the categories of the remaining text based on the MCNN, according to the fused features and the categories of some texts. The specific determination method is similar to the above method of determining the remaining text categories based on the final features and the categories of some texts, and will not be described in detail here.

[0095] Figure 8 A flowchart illustrating yet another text classification method provided in an embodiment of this application. Based on... Figure 2 ,like Figure 8 As shown, the above S230 includes:

[0096] S810: Generate a graph structure containing N texts. The graph structure includes N nodes, and each of the N nodes corresponds one-to-one with one of the N texts. When any two texts have a reference relationship, the two nodes corresponding to the two texts are connected. N is a positive integer.

[0097] S820: Uses a graph neural network to learn the connection relationships of N nodes and obtain the node vectors of each of the N nodes;

[0098] S830: Merge the node vectors of N nodes and the first vectors of N texts to obtain the merged vectors of N texts;

[0099] S840: Generate a second matrix based on the fusion vectors of N texts;

[0100] S850: Input the second matrix and the categories of a portion of the text into the second neural network to obtain the categories of the remaining text.

[0101] S810 is similar to the above embodiments, and will not be described again here.

[0102] In some implementations, the graph neural network can be a Node2Vec-based neural network, and this application does not limit this to such implementations.

[0103] In some possible implementations, the electronic device can use concat to fuse the node vectors of N nodes and the first vectors of N texts to obtain the fused vectors of N texts. This application does not limit this.

[0104] It should be noted that S840 and S850 are similar to S410 and S420 as described above, and their specific processes and effects can be referred to the above embodiments. This application will not elaborate on them here.

[0105] In summary, the technical solution provided by the above embodiments brings at least the following beneficial effects: The electronic device can acquire the text information of each of N texts and the categories of some texts among the N texts, where N is an integer greater than 1. Then, the electronic device can determine the first vector of each of the N texts based on the text information of each of the N texts. The first vector of the target text is used to characterize the correlation between the target text and the other texts among the N texts excluding the target text. The target text is any one of the N texts. Finally, the electronic device can determine the category of the remaining texts among the N texts excluding the remaining texts based on the first vectors of each of the N texts and the categories of some texts. In other words, when performing text classification, this application considers the correlation between two texts and the known categories of some texts to classify two highly correlated texts into the same category. That is, this application can ensure that two texts belonging to the same category have a high correlation, thereby making the determined text categories more accurate and improving the precision of text classification.

[0106] Furthermore, both the first probability distribution and the second probability distribution can characterize the correlation between any two texts in the N texts. The second probability distribution, determined based on the citation weights and similarities of text pairs, can more accurately characterize the correlation between any two texts in the N texts. Therefore, when the difference between the first probability distribution and the second probability distribution is less than a preset difference, the electronic device can determine that the first probability distribution, determined based on the initial vectors of each of the N texts, can more accurately characterize the correlation between any two texts in the N texts. This allows the electronic device to determine that the initial vector of each text in the N texts can more accurately characterize the correlation between that text and the other texts in the N texts. Thus, the electronic device can assign the initial vector of each of the N texts to the initial vector of each text. The vectors are determined as the first vectors of each of the N texts. When the difference between the first probability distribution and the second probability distribution is greater than or equal to a preset difference, the electronic device can determine that the first probability distribution determined based on the initial vectors of each of the N texts cannot accurately represent the correlation between any two texts in the N texts. Therefore, it can be determined that the initial vector of each text in the N texts cannot accurately represent the correlation between that text and the other texts in the N texts. Thus, the electronic device can adjust the initial vectors of each of the N texts according to the difference between the first probability distribution and the second probability distribution, and redetermine the first probability distribution until the difference between the first probability distribution and the second probability distribution is less than the preset difference, obtaining the first vectors of each of the N texts. In this way, it can be ensured that the first vector of each text in the N texts can accurately represent the correlation between that text and the other texts in the N texts, thereby accurately determining the category of the remaining texts in the N texts based on the first vectors of each text in the N texts, further improving the text classification accuracy.

[0107] Furthermore, since the node sequence is generated by the electronic device based on a random walk, the order of the nodes in the node sequence is random. This results in each pair of vectors in the first matrix generated based on the node order being the first vector of each of the two random texts. As a result, the features extracted by the neural network can better represent the correlation between any two texts, thus improving the accuracy of text classification.

[0108] The following will describe a system architecture involved in an embodiment of this application:

[0109] Figure 9 This is a schematic diagram of a system architecture provided for an embodiment of this application. Figure 9 As shown, the system architecture may include: user equipment 901, data acquisition equipment 902, training equipment 903, execution equipment 904, database 905, and content library 906.

[0110] The data acquisition device 902 is used to read training samples from the content library 906 and store the read training samples in the database 905. The training samples include: the first vector of each of the N texts, the category of some of the N texts, and the category of the remaining texts of the N texts excluding the part of the text. N is an integer greater than 1. The method for determining the first vector of each of the N texts can be referred to the description in the above embodiments, and will not be repeated here.

[0111] Training device 903 trains the first neural network and the second neural network based on training samples maintained in database 905. Specifically, training device 903 can train the first neural network based on the first vectors of each of the N texts, the categories of some texts, and the categories of the remaining texts, so that the trained first neural network can accurately output the categories of the remaining texts; alternatively, training device 903 can also train the second neural network based on the first vectors of each of the N texts, the categories of some texts, and the categories of the remaining texts, so that the trained second neural network can accurately output the categories of the remaining texts. The first neural network and the second neural network obtained by training device 903 can be applied to different systems or devices. The specific training process will be described in the following embodiments, and will not be repeated here.

[0112] In addition, such as Figure 9As shown, the execution device 904 is configured with an I / O interface 907 for data interaction with external devices. For example, it receives first vectors for each of the N texts and categories of a portion of the N texts from the user device 901 via the I / O interface. The computation module 909 in the execution device 904 can use a trained first neural network to process the first vectors for each of the N texts and the categories of a portion of the N texts, outputting the categories of the remaining texts excluding the first vectors, and then sending the output categories of the remaining texts to the user device 901 via the I / O interface. Alternatively, the computation module 909 in the execution device 904 can use a trained second neural network to process the first vectors for each of the N texts and the categories of a portion of the N texts, outputting the categories of the remaining texts excluding the first vectors, and then sending the output categories of the remaining texts to the user device 901 via the I / O interface. The execution device 904 and the user device 901 can be the same device, such as the electronic device in the above embodiments. The specific processing procedures of the execution device 904 and the user device 901 can be referred to the description in the above embodiments, and will not be repeated here.

[0113] User equipment 901 may include mobile phones, tablets, laptops, handheld computers, mobile internet devices (MIDs), or other terminal devices with browser installation capabilities.

[0114] The execution device 904 can be a server. Optionally, the server can be a rack server, blade server, tower server, or cabinet server, etc. The server can be a standalone test server or a test server cluster composed of multiple test servers.

[0115] The execution device 904 can connect to the user equipment 901 via a network. This network can be an intranet, the Internet, Global System for Mobile Communications (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, voice communication network, or other wireless or wired networks.

[0116] It should be noted that, Figure 9The positional relationships between the devices, components, modules, etc. shown are not restrictive. Optionally, the data acquisition device 902, user device 901, training device 903, and execution device 904 may be the same device; the database 905 may be distributed on one server or multiple servers; and the content library 906 may be distributed on one server or multiple servers. This application does not impose any restrictions on these aspects.

[0117] The training process of the first neural network is explained below:

[0118] In some implementations, the training device can first input the first vectors of N texts and the categories of a portion of the N texts into a first neural network to obtain the categories of the remaining texts excluding the portion of texts. Then, a loss function can be used to calculate the loss value between the output categories of the remaining texts and the predetermined categories of the remaining texts. After obtaining the loss value, the parameters of the first neural network can be updated based on the loss value. The training device can complete the training of the first neural network when the loss value reaches a preset value; or, it can complete the training when the number of training iterations of the first neural network reaches a preset number. Of course, this application does not limit the training method of the first neural network.

[0119] In some implementations, the first neural network can be an MCNN, which includes M CNNs, where at least two CNNs have at least one pair of convolutional layers with different kernel sizes at corresponding positions, and M is an integer greater than 1.

[0120] In some implementations, the dimension of the first vector of each of the N texts is the same as the width of any convolutional kernel in the CNN.

[0121] It should be understood that the first neural network can be used to implement the methods described in the text classification section above. The content and effects can be referenced in the text classification section above, and this application will not elaborate on its content and effects further. The training process of the second neural network is similar to that of the first neural network, and its content and effects can be referenced in the training method described above for the first neural network. This application will not elaborate on its details here.

[0122] Figure 10 A schematic diagram of a text classification device 1000 provided in an embodiment of this application is shown below. Figure 10 As shown, the device 1000 includes:

[0123] The acquisition module 1010 is used to acquire the text information of each of the N texts and the category of some of the texts in the N texts, where N is an integer greater than 1;

[0124] The first determining module 1020 is used to determine the first vector of each of the N texts based on their respective text information. The first vector of the target text is used to characterize the correlation between the target text and the other texts in the N texts, except for the target text. The target text is any one of the N texts.

[0125] The second determining module 1030 is used to determine the category of the remaining texts (excluding the partial texts) among the N texts based on the first vector of each of the N texts and the category of the partial text.

[0126] In some implementations, the text information includes: citation information and text content. The first determining module 1020 is specifically used to: vectorize the text content of each of the two texts included in the target text pair to obtain the second vector of each of the two texts; the target text pair is any text pair among N texts; determine the first probability corresponding to the target text pair based on the second vector of each of the two texts and the initial vector corresponding to the first vector of each of the two texts; determine the citation weight of the target text pair based on the citation information of each of the two texts; determine the similarity of the target text pair based on the second vector of each of the two texts; determine the second probability corresponding to the target text pair based on the citation weight of the target text pair, the similarity of the target text pair, the citation weight of all text pairs among the N texts, and the similarity of all text pairs among the N texts; and determine the first vector of each of the N texts based on the first probability and the second probability of all text pairs among the N texts.

[0127] In some possible implementations, the first determining module 1020 is specifically used to: calculate the product of the initial vectors corresponding to the two texts to obtain a first product; calculate the product of the second vectors of the two texts to obtain a second product; calculate the maximum value of the modulus of the second vectors of the two texts; and determine the first probability corresponding to the target text pair based on the first product, the second product, and the maximum value.

[0128] In some implementations, the first determining module 1020 is specifically used to: calculate the hyperbolic tangent of the first product; calculate the ratio of the second product to the maximum value to obtain a first value; and determine the first probability corresponding to the target text pair based on the hyperbolic tangent and the first value.

[0129] In some implementations, the first determining module 1020 is specifically used to: normalize the hyperbolic tangent value to obtain a second value; and determine the first probability corresponding to the target text pair based on the first and second values.

[0130] In some implementations, the first determining module 1020 is specifically used to: calculate the product of the first value and the second value to obtain the first probability corresponding to the target text pair.

[0131] In some possible implementations, the first determining module 1020 is specifically used to: calculate the sum of the citation weight of the target text pair and the similarity of the target text pair to obtain a third value; calculate the sum of the citation weight of all text pairs in N texts and the similarity of all text pairs in N texts to obtain a fourth value; and determine the second probability corresponding to the target text pair based on the third value and the fourth value.

[0132] In some implementations, the first determining module 1020 is specifically used to: calculate the ratio of the third value to the fourth value to obtain the second probability corresponding to the target text pair.

[0133] In some possible implementations, the first determining module 1020 is specifically used to: determine the first probability distribution corresponding to the first probability of all text pairs in N texts; determine the second probability distribution corresponding to the second probability of all text pairs in N texts; calculate the difference between the first probability distribution and the second probability distribution; if the difference between the first probability distribution and the second probability distribution is less than a preset difference, then determine the initial vector of each of the N texts as the first vector of each of the N texts; if the difference between the first probability distribution and the second probability distribution is greater than or equal to the preset difference, then adjust the initial vector of each of the N texts, and continue to execute the determination of the first probability corresponding to the target text pair based on the second vector of each of the two texts and the initial vector corresponding to the first vector of each of the two texts, until the difference between the first probability distribution and the second probability distribution is less than the preset difference, and the first vector of each of the N texts is obtained.

[0134] In some implementations, the second determining module 1030 is specifically used to: generate a first matrix based on the first vectors of each of the N texts; input the first matrix and the categories of some texts into a first neural network to obtain the categories of the remaining texts.

[0135] In some possible implementations, the second determining module 1030 is specifically used to: generate a graph structure of N texts, the graph structure including N nodes, the N nodes and N texts corresponding one-to-one, and when any two texts in the N texts have a reference relationship, the two nodes corresponding to the two texts are connected; for any node in the N nodes, perform a random walk starting from the node to obtain a node sequence; and combine the first vectors of the N texts according to the order of the node sequence to obtain a first matrix.

[0136] In some implementations, the first neural network is a multi-scale convolutional neural network (MCNN), which includes M convolutional neural networks (CNNs). At least two of the CNNs have at least one pair of convolutional layers with different kernel sizes at corresponding positions, and M is an integer greater than 1.

[0137] In some implementations, the dimension of the first vector of each of the N texts is the same as the width of any convolutional kernel in the CNN.

[0138] In some implementations, the second determining module 1030 is specifically used to: generate a graph structure of N texts, the graph structure including N nodes, with each of the N nodes corresponding to one of the N texts, and the two nodes corresponding to any two texts being connected when there is a reference relationship between them; use a graph neural network to learn the connection relationship of the N nodes to obtain the node vectors of the N nodes; fuse the node vectors of the N nodes with the first vectors of the N texts to obtain the fused vectors of the N texts; generate a second matrix based on the fused vectors of the N texts; and input the second matrix and the categories of some texts into the second neural network to obtain the categories of the remaining texts.

[0139] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here. Specifically, Figure 10 The apparatus 1000 shown can execute the above-described method embodiments, and the aforementioned and other operations and / or functions of each module in the apparatus 1000 are respectively for implementing the corresponding processes in the above-described methods. For the sake of brevity, they will not be described in detail here.

[0140] The apparatus 1000 of this application embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this application embodiment can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.

[0141] Figure 11 This is a schematic block diagram of the electronic device 1100 provided in the embodiments of this application.

[0142] like Figure 11 As shown, the electronic device 1100 may include:

[0143] The system includes a memory 1110 and a processor 1120. The memory 1110 stores computer programs and transfers the program code to the processor 1120. In other words, the processor 1120 can retrieve and run the computer program from the memory 1110 to implement the methods described in the embodiments of this application.

[0144] For example, the processor 1120 can be used to execute the above-described method embodiments according to instructions in the computer program.

[0145] In some embodiments of this application, the processor 1120 may include, but is not limited to:

[0146] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0147] In some embodiments of this application, the memory 1110 includes, but is not limited to:

[0148] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0149] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 1110 and executed by the processor 1120 to complete the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device.

[0150] like Figure 11 As shown, the electronic device may also include:

[0151] Transceiver 1130, which can be connected to processor 1120 or memory 1110.

[0152] The processor 1120 can control the transceiver 1130 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 1130 may include a transmitter and a receiver. The transceiver 1130 may further include antennas, and the number of antennas may be one or more.

[0153] It should be understood that the various components in the electronic device are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.

[0154] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.

[0155] When implemented using software, it can be implemented entirely or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0156] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0157] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0158] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0159] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A text classification method, characterized in that, include: Obtain the text information of each of N texts and the category of some texts among the N texts, where N is an integer greater than 1. The text information includes: citation information and text content; For any target text pair among the N texts, the text content of each of the two texts is vectorized to obtain the second vector of each of the two texts; the first probability corresponding to the target text pair is determined based on the second vector of each of the two texts and the initial vector corresponding to the first vector of each of the two texts; the reference weight of the target text pair is determined based on the reference information of each of the two texts. The similarity of the target text pair is determined based on the second vectors of the two texts respectively; the second probability corresponding to the target text pair is determined based on the citation weight of the target text pair, the similarity of the target text pair, the citation weight and similarity of all text pairs in the N texts; the first vector of each of the N texts is determined based on the first probability and the second probability of all text pairs in the N texts, wherein the first vector of any target text in the N texts is used to characterize the relevance between the target text and other texts in the N texts other than the target text itself; Based on the first vector of each of the N texts and the category of the partial text, determine the category of the remaining texts among the N texts excluding the partial text; The step of determining the first probability corresponding to the target text pair based on the second vector of each of the two texts and the initial vector corresponding to the first vector of each of the two texts includes: Calculate the product of the initial vectors corresponding to the two texts to obtain the first product; Calculate the product of the second vectors of the two texts to obtain the second product; Calculate the maximum value of the modulus of the second vector of each of the two texts; The first probability corresponding to the target text pair is determined based on the first product, the second product, and the maximum value; The determination of the second probability corresponding to the target text pair based on the citation weight of the target text pair, the similarity of the target text pair, the citation weight of all text pairs in the N texts, and the similarity of all text pairs in the N texts includes: The third value is obtained by summing the citation weight of the target text pair with the similarity of the target text pair. The fourth value is obtained by calculating the sum of the citation weights of all text pairs in the N texts and the similarity of all text pairs in the N texts. The second probability corresponding to the target text pair is determined based on the third and fourth values.

2. The method according to claim 1, characterized in that, The step of determining the first probability corresponding to the target text pair based on the first product, the second product, and the maximum value includes: Calculate the hyperbolic tangent of the first product; Calculate the ratio of the second product to the maximum value to obtain the first value; The first probability corresponding to the target text pair is determined based on the hyperbolic tangent value and the first numerical value.

3. The method according to claim 2, characterized in that, Determining the first probability corresponding to the target text pair based on the hyperbolic tangent value and the first numerical value includes: The hyperbolic tangent value is normalized to obtain a second value; The first probability corresponding to the target text pair is determined based on the first value and the second value.

4. The method according to any one of claims 1-3, characterized in that, The process of determining the first vector for each of the N texts based on the first and second probabilities of all text pairs in the N texts includes: Determine the first probability distribution corresponding to the first probability of all text pairs in the N texts; Determine the second probability distribution corresponding to the second probability of all text pairs in the N texts; Calculate the difference between the first probability distribution and the second probability distribution; If the difference between the first probability distribution and the second probability distribution is less than a preset difference, then the initial vector of each of the N texts is determined as the first vector of each of the N texts; If the difference between the first probability distribution and the second probability distribution is greater than or equal to the preset difference, then the initial vectors of the N texts are adjusted, and the process of determining the first probability of the target text pair based on the initial vectors corresponding to the second vectors of the two texts and the first vectors of the two texts continues until the difference between the first probability distribution and the second probability distribution is less than the preset difference, and the first vectors of the N texts are obtained.

5. The method according to any one of claims 1-3, characterized in that, The step of determining the category of the remaining texts (excluding the partial texts) among the N texts based on their respective first vectors and the category of the partial texts includes: A first matrix is ​​generated based on the first vector of each of the N texts; The first matrix and the categories of the partial text are input into the first neural network to obtain the categories of the remaining text.

6. The method according to claim 5, characterized in that, The step of generating a first matrix based on the first vectors of the N texts includes: Generate a graph structure for the N texts, the graph structure including N nodes, the N nodes corresponding one-to-one with the N texts, and when any two texts have a reference relationship, the two nodes corresponding to the two texts are connected; For any one of the N nodes, a random walk is performed starting from that node to obtain a node sequence; The first matrix is ​​obtained by combining the first vectors of the N texts according to the order of the node sequence.

7. The method according to claim 5, characterized in that, The first neural network is a multi-scale convolutional neural network (MCNN), which includes M convolutional neural networks (CNNs). At least two of the CNNs have at least one pair of convolutional layers with different kernel sizes at corresponding positions, and M is an integer greater than 1.

8. The method according to any one of claims 1-3, characterized in that, The first vector based on each of the N texts and the The classification of a portion of the text, determining the classification of the remaining texts among the N texts excluding the portion of the text, includes: Generate a graph structure for the N texts, the graph structure including N nodes, the N nodes corresponding one-to-one with the N texts, and when any two texts have a reference relationship, the two nodes corresponding to the two texts are connected; The connection relationships among the N nodes are learned using a graph neural network, and the node vectors of each of the N nodes are obtained. The node vectors of the N nodes and the first vectors of the N texts are fused together to obtain the fused vectors of the N texts. A second matrix is ​​generated based on the fusion vectors of the N texts; The second matrix and the categories of the partial text are input into the second neural network to obtain the categories of the remaining text.

9. A text classification device, characterized in that, include: The acquisition module is used to acquire the text information of each of N texts and the category of a portion of the texts among the N texts, where N is an integer greater than 1. The text information includes: citation information and text content; The first determining module is configured to: vectorize the text content of each of the two texts in any target text pair from the N texts to obtain a second vector for each of the two texts; determine a first probability corresponding to the target text pair based on the second vectors of each of the two texts and the initial vectors corresponding to the first vectors of each of the two texts; determine the citation weight of the target text pair based on the citation information of each of the two texts; determine the similarity of the target text pair based on the second vectors of each of the two texts; determine a second probability corresponding to the target text pair based on the citation weight of the target text pair, the similarity of the target text pair, the citation weights and similarities of all text pairs in the N texts; and determine a first vector for each of the N texts based on the first probability and the second probability of all text pairs in the N texts, wherein the first vector of any target text in the N texts is used to characterize the relevance between the target text and other texts in the N texts besides the target text itself. The second determining module is used to determine the category of the remaining texts (excluding the partial texts) among the N texts based on the first vector of each of the N texts and the category of the partial texts; The first determining module is specifically used to: calculate the initial values ​​corresponding to each of the two texts. The product of the initial vectors is used to obtain the first product; the product of the second vectors of the two texts is used to obtain the second product; the maximum value of the modulus of the second vectors of the two texts is calculated; the first probability of the target text pair is determined based on the first product, the second product, and the maximum value. The first determining module is also specifically used for: calculating the sum of the citation weight of the target text pair and the similarity of the target text pair to obtain a third value; calculating the sum of the citation weight of all text pairs in N texts and the similarity of all text pairs in N texts to obtain a fourth value; and determining the second probability corresponding to the target text pair based on the third value and the fourth value.

10. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the method of any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes computer instructions that cause an electronic device to perform the method of any one of claims 1 to 8.