Text clustering method and device, equipment, medium and program product

The method uses a classification model to filter and cluster text based on keyword networks and vectors, addressing the challenge of varying user expressions to enhance clustering accuracy and efficiency in identifying risk events.

CN120316264APending Publication Date: 2025-07-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510481959.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-15

Smart Images

  • Figure CN120316264A_ABST
    Figure CN120316264A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text clustering method and device, equipment, a medium and a program product. The method comprises the steps that a first text set is input into a preset classification model, a second text set in the first text set is determined based on the preset classification model, and the second text set is a set of texts of a preset category. Constructing a keyword network according to the keywords in the second text set, generating a plurality of keyword sequences according to the keyword network, determining text vectors of the texts in the second text set based on the plurality of keyword sequences, and performing text clustering based on the text vectors. According to the embodiment of the invention, the second text set of the preset category is screened out from the first text set through the preset classification model, and the texts are clustered based on the keywords of the texts in the second text set, so that the to-be-clustered texts can be reduced, the clustering efficiency is improved, the interference of irrelevant information can be reduced, the clustering accuracy is improved, and the user experience is improved. Similar contents can be better clustered into the same cluster, and the clustering effect is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to Internet technologies, and in particular, to a text clustering method, apparatus, device, medium, and program product. Background Art

[0002] With the development of Internet technologies, more and more users conduct communication, learning, work, entertainment and other behaviors through the Internet. As the scale of Internet users gradually increases, the risks of network behaviors become more and more diverse, threatening the security of users' network behaviors. Users' feedback is an important information source for discovering risk events that threaten network behaviors. However, due to differences in people's expression habits, there are significant differences in the descriptions of the same risk event, resulting in the possibility that the solutions of related technologies may cluster different description texts of the same risk event into different clusters, with relatively low clustering accuracy. Summary of the Invention

[0003] Embodiments of the present disclosure provide a text clustering method, apparatus, device, medium, and program product, which can accurately perform text clustering and improve clustering accuracy.

[0004] In a first aspect, embodiments of the present disclosure provide a text clustering method, including:

[0005] Inputting a first text set into a preset classification model, and determining a second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set;

[0006] Constructing a keyword network according to keywords in the second text set, generating a plurality of keyword sequences according to the keyword network, determining text vectors of texts in the second text set based on the plurality of keyword sequences, and performing text clustering based on the text vectors.

[0007] In a second aspect, embodiments of the present disclosure further provide a text clustering apparatus, including:

[0008] A text screening module, configured to input a first text set into a preset classification model, and determine a second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set;

[0009] A clustering module, configured to construct a keyword network according to keywords in the second text set, generate a plurality of keyword sequences according to the keyword network, determine text vectors of texts in the second text set based on the plurality of keyword sequences, and perform text clustering based on the text vectors.

[0010] In a third aspect, embodiments of the present disclosure further provide an electronic device, which includes:

[0011] One or more processors;

[0012] A storage device for storing one or more programs,

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the text clustering method as described in any embodiment of the present disclosure.

[0014] In a fourth aspect, embodiments of the present disclosure further provide a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the text clustering method as described in any embodiment of the present disclosure when executed by a computer processor.

[0015] In a fifth aspect, embodiments of the present disclosure further provide a computer program product, including a computer program, and the computer program implements the text clustering method as described in any embodiment of the present disclosure when executed by a processor.

[0016] Embodiments of the present disclosure provide a text clustering method. By inputting a first text set into a preset classification model, a second text set in the first text set is determined based on the preset classification model, and the second text set is a set of texts of a preset category. A keyword network is constructed according to the keywords in the second text set, multiple keyword sequences are generated according to the keyword network, text vectors of the texts in the second text set are determined based on the multiple keyword sequences, and text clustering is performed based on the text vectors. Embodiments of the present disclosure screen out a second text set of a preset category from the first text set through a preset classification model, and cluster texts based on the keywords of the texts in the second text set, which can reduce the texts to be clustered, improve the clustering efficiency, reduce the interference of irrelevant information, improve the clustering accuracy, better cluster similar contents into the same cluster, and optimize the clustering effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.

[0018] Figure 1 It is a flowchart of a text clustering method provided by an embodiment of the present disclosure;

[0019] Figure 2 It is a schematic diagram of a keyword network provided by an embodiment of the present disclosure;

[0020] Figure 3 Schematic flowchart of another text clustering method provided by an embodiment of the present disclosure;

[0021] Figure 4 Schematic diagram of a process for generating word vectors of keywords provided by an embodiment of the present disclosure;

[0022] Figure 5 Schematic structural diagram of a text clustering device provided by an embodiment of the present disclosure;

[0023] Figure 6 Schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0025] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0026] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0027] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.

[0028] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0030] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0031] Figure 1 FIG. is a schematic flow chart of a text clustering method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of text mining. This method can be executed by a text clustering device, which can be implemented in the form of software and / or hardware. Optionally, it is implemented by an electronic device, which can be a mobile terminal, a PC or a server, etc.

[0032] As Figure 1 shown, the method includes:

[0033] S110. Input a first text set into a preset classification model, and determine a second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set.

[0034] Among them, the first text set includes feedback texts collected through a preset feedback channel. Optionally, the first text set is feedback texts collected through a preset feedback channel within a preset time period. The preset time period represents a time range determined according to a risk control strategy. The preset time period can be set according to actual needs. For example, the first text generated on the previous day can be collected at a preset time every day to form the first text set. Among them, the first text includes complaint texts or feedback texts collected through various feedback channels. The feedback channels include at least one of telephone complaints, online customer service, intelligent customer service, etc.

[0035] The preset category can be a text screening condition set according to business requirements. For example, the preset category can include at least one of a preset risk type, a preset theme, a preset region, a preset time, etc. Correspondingly, the second text set includes second texts associated with the preset risk type. Among them, the preset risk type represents a violation determined according to a risk control strategy. At least one risk type can be set according to actual needs. For example, the preset risk type includes violations corresponding to user accounts, etc. The second text can refer to the text in the first text set associated with the preset risk type.

[0036] Among them, the preset classification model can be a natural language processing model with text classification function.

[0037] Exemplarily, the acquisition method of the second text set includes:

[0038] Input the first text set into a preset classification model, where the first text in the first text set includes feedback text; determine the category label of the first text through the preset classification model based on preset classification prompt information, where the preset classification prompt information includes category definitions and / or category examples; screen the first text set according to the category label, and determine the second text set according to the screening result.

[0039] The feedback text of users is an important information source for identifying risk behaviors. By analyzing the feedback text, new types of risk behaviors can be promptly discovered. Thus, based on the new risk behaviors, the security policy can be updated in a timely manner, thereby effectively preventing and controlling risk behaviors.

[0040] For example, obtain the first text within a preset time period to form a first text set. The first text includes the feedback text of users. Due to significant differences in users' expression habits, there are diverse ways to describe events or behaviors. Even if the feedback content involves the same event or behavior, the description methods are different. To screen out the text associated with the preset risk type from the feedback text of users, it is necessary to input the first text set into a preset classification model. By determining the category label of the first text through the preset classification model based on the preset classification prompt information, the accuracy of text classification can be improved, and the classification efficiency can also be effectively enhanced. The preset classification prompt information includes category definitions and / or category examples, etc.

[0041] In some embodiments, the preset classification prompt information is the prompt information preset in the model. For example, the preset classification prompt information includes category definitions. The category definition includes the description content of risk behaviors in the feedback text associated with the preset risk type. Optionally, the preset risk type and / or the description detail level of risk behaviors can be specified through the category definition. Based on the category definition, the second text whose description of risk behaviors in the text associated with the preset risk type meets the preset requirements can be screened out. The category example includes the description examples of risk behaviors. Optionally, the preset classification prompt information can also include output format examples, etc. The output format example is used to specify the format of the model output result.

[0042] The second text associated with at least one of the preset risk type, preset theme, preset region, etc. can be screened out from the first text set through the category label. The category label is the information reflecting the text category. For example, the category label includes the feedback related to the preset risk type with detailed description of risk behaviors, the feedback related to the preset risk type without description of risk behaviors, and other feedbacks, etc. Optionally, the second text with detailed description of risk behaviors in the feedback related to the preset risk type can be screened out from the first text set through the category label. Since this type of text contains a large amount of description content related to risk behaviors, thus, new risk behaviors can be automatically identified based on the above description content.

[0043] S120. Construct a keyword network based on the keywords in the second text set, generate multiple keyword sequences according to the keyword network, determine the text vectors of the texts in the second text set based on the multiple keyword sequences, and perform text clustering based on the text vectors.

[0044] Among them, a keyword network refers to a graph structure formed by connecting keywords with an associated relationship. The nodes of the keyword network represent keywords, and the edges of the keyword network represent the associated relationship between two keywords. The weight of the edge in the keyword network represents the associated strength between two keywords. The associated relationship can represent the relationship between keywords. For example, the associated relationship includes co-occurrence relationships, etc. The co-occurrence relationship represents that two keywords appear in the same text.

[0045] The keywords in the second text can be obtained by using a preset keyword extraction algorithm. Among them, the preset keyword extraction algorithm can include the term frequency-inverse document frequency algorithm or the graph-based ranking algorithm, etc. For example, each second text in the second text set is processed as follows: The second text is segmented to obtain a list of words corresponding to the second text. The weight of each word is determined according to the term frequency and inverse document frequency of each word in the word list. According to the weights of each word in the word list corresponding to the second text, the keywords in the second text are determined. For example, the words with weights greater than a preset weight threshold in the word list corresponding to the second text are determined as the keywords of the second text. Or, the words in the word list corresponding to the second text are sorted in descending order according to the weights, and the top N words are selected as the keywords of the second text. By extracting the keywords of the second text, the influence of interference information introduced by differences in description methods on text clustering can be reduced.

[0046] Exemplarily, a keyword network is constructed according to the associated relationship between the keywords in the second text set, where the weight of the edge in the keyword network represents the associated strength between the keywords. Multiple keyword sequences are generated according to the associated strength between the keywords in the keyword network, the word vectors of the keywords in the second text set are determined according to the multiple keyword sequences, the text vectors of the second texts in the second text set are determined according to the word vectors, and the second texts are clustered according to the text vectors.

[0047] In some embodiments, constructing a keyword network according to the associated relationship between the keywords in the second text set includes: constructing the keyword network according to the co-occurrence relationship between the keywords in the second text set. The weight of the edge is determined according to the associated strength of the keyword pairs connected by the edge in the keyword network. The keyword network is updated according to the weight of the edge.

[0048] Figure 2A schematic diagram of a keyword network provided by an embodiment of the present disclosure. As Figure 2 shown, the nodes of the keyword network represent keywords. The edges in the keyword network connect two keywords with an associated relationship. The thickness of the edge represents the association strength.

[0049] For any two keywords in the second text set, if these two keywords appear in the same text, it is determined that these two keywords have an associated relationship (or co-occurrence relationship), and these two keywords can be connected by an edge. The two keywords connected by the edge are determined as a keyword pair. According to the co-occurrence probability of the keyword pair in the second text set and the occurrence probability of each keyword in the keyword pair, the association strength of the keyword pair is determined. For example, for each pair of keyword pairs, according to the number of occurrences of the current keyword pair and the number of keyword pairs in the second text set, the co-occurrence probability of the current keyword pair is determined. According to the number of occurrences of each keyword in the current keyword pair and the number of keywords in the first document set, the occurrence probability of each keyword in the current keyword pair is determined.

[0050] The weight of the edge is determined according to the association strength of the keyword pair connected by the edge in the keyword network. The weight of the edge is compared with a preset weight threshold. If the weight of the edge meets the preset weight threshold, the edge is retained. If the weight of the edge does not meet the preset weight threshold, the edge is deleted. The keyword network is updated through the weight of the edge to disconnect the connections between keywords with lower association degrees, simplifying the complexity of the keyword network.

[0051] Among them, the keyword sequence includes a unidirectional non-closed network branch in the keyword network. The keyword network can be divided based on the association relationship and association strength between keywords to obtain multiple keyword sequences. A preset word vector generation model is trained based on the multiple keyword sequences. After the preset word vector generation model is trained, the keywords in the second text set are input into the preset word vector generation model, and the word vector output by the hidden layer of the preset word vector generation model is obtained.

[0052] In some embodiments, the preset word vector generation model may be a neural network model for natural language processing. The preset word vector generation model maps keywords to a low-dimensional dense vector space through unsupervised learning to obtain the word vectors of the keywords. For example, the preset word vector generation model includes an input layer, a hidden layer, and an output layer. Among them, the input layer is used to convert keywords into keyword encodings that can be processed by a computer. For example, the keyword encoding can be one-hot encoding, etc. The hidden layer is used to convert the sparse vector output by the input layer into a word vector through a weight matrix. The output layer is used to output the central word or context word and the corresponding probability distribution. The central word can be any keyword in the keyword sequence. The context word is a keyword adjacent to the central word. The context word is determined based on the sliding window size in the hyperparameters corresponding to the preset word vector generation model.

[0053] Among them, the text vector is a vector representation of the second text. The second text can be characterized based on the word vectors of each keyword in the second text.

[0054] Exemplarily, the text vector of the second text can be determined according to the word vectors of the keywords in the second text. For example, according to the word vectors of each keyword in the second text, the text vector of the second text is constructed. Optionally, the text vector of the second text can be determined according to the average vector of the word vectors of each keyword in the second text. Or, the text vector of the second text can be determined according to the weighted sum of the word vectors of each keyword in the second text. The weight of the word vector of the keyword is a preset value.

[0055] For example, according to the core distance and the mutual reachability distance between the text vectors of the second text, a weighted graph of the text vectors is constructed. The nodes of the weighted graph represent the text vectors. According to the mutual reachability distance, the minimum spanning tree of the weighted graph is constructed to retain the connected structure of the shortest mutual reachability distance between the nodes. The mutual reachability distance is used to represent the weight of the edges in the minimum spanning tree. The edges of the minimum spanning tree are sorted in descending order according to the edge weights, and the edges with the maximum weights are disconnected in turn to form sub-clusters. Through recursive segmentation, a hierarchical clustering tree is formed. By performing persistence analysis on the sub-clusters in the hierarchical clustering tree, stable clusters are retained, and unstable clusters are removed to complete the clustering. Through the above clustering method, the number of clusters can be automatically determined.

[0056] Using a density-based clustering algorithm, cluster the text vectors to cluster the second texts with similar contents in the second text set into the same cluster, obtaining multiple clusters. The content in each cluster is the description content of a risk behavior. Correspondingly, text clustering is performed based on the clustering result, including: determining the risk behavior based on the keywords of the second texts within the same cluster. Since a large number of users usually provide feedback on a new risk behavior after it appears, these feedback texts are likely to be clustered in one cluster. Therefore, by analyzing the description content of the risk behavior in the cluster, new types of risk behaviors can be quickly discovered, the risk control strategy can be adjusted in a timely manner, and the confrontation speed of the risk control strategy can be accelerated. The embodiments of the present disclosure do not need to use a large number of high-quality sample training text clustering models. When a new type of risk behavior appears, the risk behavior can be identified in a timely manner through the clustering result, facilitating blocking before the risk behavior spreads on a large scale.

[0057] In the technical solution of the embodiments of the present disclosure, by inputting the first text set into a preset classification model, the second text set in the first text set is determined based on the preset classification model, and the second text set is a set of texts of a preset category. A keyword network is constructed according to the keywords in the second text set, multiple keyword sequences are generated according to the keyword network, the text vectors of the texts in the second text set are determined based on the multiple keyword sequences, and text clustering is performed based on the text vectors. The embodiments of the present disclosure screen out the second text set of the preset category from the first text set through the preset classification model, and perform text clustering based on the keywords of the texts in the second text set, which can reduce the texts to be clustered, improve the clustering efficiency, reduce the interference of irrelevant information, improve the accuracy of clustering, and better cluster similar contents into the same cluster, optimizing the clustering effect.

[0058] Figure 3 It is a schematic flowchart of another text clustering method provided by the embodiments of the present disclosure. On the basis of the above embodiments, the embodiments of the present disclosure specifically define generating multiple keyword sequences according to the association strength between the keywords in the keyword network, and determining the word vectors of the keywords in the second text set according to the multiple keyword sequences.

[0059] As Figure 3 shown, the method includes:

[0060] S310: Input the first text set into a preset classification model, and determine the second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set.

[0061] S320: Construct the keyword network according to the co-occurrence relationship between the keywords in the second text set.

[0062] S330. Determine the weight of the edge according to the association strength of the keyword pairs connected by the edge in the keyword network.

[0063] S340. Update the keyword network according to the weight of the edge.

[0064] S350. For any keyword in the keyword network, determine at least one neighbor keyword of the current keyword according to the edge corresponding to the current keyword.

[0065] For two keywords connected by an edge in the keyword network, if one keyword is used as the current keyword, the other keyword can be used as the neighbor keyword of the current keyword. It should be noted that the current keyword can have at least one neighbor keyword.

[0066] S360. Determine the probability of selecting the neighbor keyword to generate the keyword sequence according to the association strength between the current keyword and each neighbor keyword.

[0067] Exemplarily, the probability of selecting a neighbor keyword and the current keyword to generate a keyword sequence can be determined according to the weight of the edge in the keyword network. That is, for the current keyword, the probability of selecting a neighbor keyword can be determined according to the weight of the edge between the current keyword and the neighbor keyword, as the adjacent keyword of the current keyword in the keyword sequence. Refer to Figure 2 , taking keyword A in the keyword network as the starting point of the sequence, then the keywords C, F, R, E, Y, and U connected to keyword A by an edge are used as the neighbor keywords of keyword A. Assume that the weight of the edge between keyword A and keyword C is 0.2. The weight of the edge between keyword A and keyword F is 0.2. The weight of the edge between keyword A and keyword R is 0.2. The weight of the edge between keyword A and keyword E is 0.1, the weight of the edge between keyword A and keyword Y is 0.2, and the weight of the edge between keyword A and keyword U is 0.1. Then the probabilities of selecting keyword C, keyword F, keyword R, keyword E, keyword Y, or keyword U as the adjacent keyword of keyword A to generate a keyword sequence are 20%, 20%, 20%, 10%, 20%, and 10% respectively.

[0068] S370. Select the neighbor keyword as the adjacent keyword of the current keyword according to the probability corresponding to each neighbor keyword of the current keyword to generate the keyword sequence.

[0069] Adjacent keywords are selected from at least one neighbor keyword of the current keyword according to the probability of selection of each neighbor keyword to construct a keyword sequence. Thus, the distribution of the keyword sequence in which the current keyword is adjacent to the neighbor keyword conforms to the probability of selection of the corresponding neighbor keyword.

[0070] For example, if it is preset to generate 10 keyword sequences starting from keyword A, then among these 10 keyword sequences, in 20% of the keyword sequences, the adjacent keyword of keyword A is keyword C. Among these 10 keyword sequences, in 20% of the keyword sequences, the adjacent keyword of keyword A is keyword F. Among these 10 keyword sequences, in 20% of the keyword sequences, the adjacent keyword of keyword A is keyword R. Among these 10 keyword sequences, in 10% of the keyword sequences, the adjacent keyword of keyword A is keyword E. Among these 10 keyword sequences, in 20% of the keyword sequences, the adjacent keyword of keyword A is keyword Y. Among these 10 keyword sequences, in 10% of the keyword sequences, the adjacent keyword of keyword A is keyword U. It should be noted that the number of keyword sequences generated starting from keyword A mentioned in the above example is only for illustration, and the number of keyword sequences generated can be set according to actual needs.

[0071] S380: If the length of the keyword sequence does not meet the preset length condition, use the neighbor keyword as the new current keyword, and return to execute S350 to determine the neighbor keyword of the current keyword according to the edge corresponding to the current keyword.

[0072] Among them, the preset length condition can be the preset length range of the keyword sequence. If the length of the keyword sequence formed by the current keyword and the neighbor keyword does not meet the preset length condition, then use the neighbor keyword as the new current keyword, and loop to execute the above steps to add new neighbor keywords to the keyword sequence.

[0073] S390: Obtain keyword pairs in the keyword sequence according to the preset sliding window.

[0074] Among them, a preset sliding window is used to determine the associated keywords of each keyword in the keyword sequence. The associated keyword can be the context of the current keyword in the keyword sequence. For example, for the keyword sequence: keyword A, keyword F, keyword B, keyword Z, keyword O. Assuming that the preset sliding window is 1, if keyword A is the current keyword, keyword F after keyword A is used as the associated keyword of keyword A. If keyword F is the current keyword, then keyword A in front of keyword F and keyword B after keyword F are used as the associated keywords of keyword F. In a similar way, the associated keyword corresponding to each keyword in each keyword sequence can be determined. Keyword pairs are formed according to the keyword and the associated keyword corresponding to the keyword.

[0075] S3100. Determine a training sample set according to the keyword pairs in the keyword sequence, and train a preset word vector generation model according to the training sample set.

[0076] For example, for any keyword in any keyword sequence, the keyword pair formed by the current keyword and the associated keyword corresponding to the current keyword is used as a training sample pair, and a training sample set is constructed according to the training sample pairs. Taking the current keyword as keyword F as an example, the training sample pairs can include {keyword F, keyword A} and {keyword F, keyword B}. A preset word vector generation model is trained according to the training sample set.

[0077] S3110. Determine the word vectors of the keywords in the second text set according to the trained preset word vector generation model.

[0078] Exemplarily, the keywords in the second text set are grouped and input into the trained preset word vector generation model, and the keyword encoding is multiplied by the weight matrix of the hidden layer of the preset word vector generation model to obtain the word vector of the keyword. Among them, the keyword encoding can be one-hot encoding, etc.

[0079] Figure 4 It is a schematic diagram of a process for generating word vectors of keywords provided by an embodiment of the present disclosure. As Figure 4As shown, multiple keyword sequences 420 are generated according to the keyword network 410. For example, the keyword sequences 420 include: keyword A-keyword F-keyword B-keyword Z-keyword O; keyword A-keyword F-keyword B-keyword K; keyword A-keyword C-keyword T-keyword O,.... For each keyword sequence 420, multiple training sample pairs 430 are determined. For example, taking keyword A-keyword F-keyword B-keyword Z-keyword O as an example, for the training sample pair 430 generated with keyword A as the current keyword, {keyword A, keyword F}, and for keyword F as the current keyword, {keyword F, keyword A} and {keyword F, keyword B}, for keyword B as the current keyword, {keyword B, keyword F} and {keyword B, keyword Z},.... The preset word vector generation model 440 is trained based on the training sample pairs 430. The word vectors 450 of the keywords in the second text set are determined according to the trained preset word vector generation model 440.

[0080] S3120. Determine the text vector of the second text according to the word vector, and cluster the second text according to the text vector.

[0081] The technical solution of the embodiment of the present disclosure screens out the second text set of the preset category from the first text set through the preset classification model, reducing the number of texts to be clustered and facilitating the improvement of the clustering efficiency. A keyword network is constructed based on the keywords in the second text set, and multiple keyword sequences are generated based on the keyword network. Then, the preset word vector generation model is trained based on the keyword sequences. The word vectors of the keywords can be determined according to the trained preset word vector generation model, so that the representation of the word vectors of each keyword is very sufficient, improving the vector representation quality of the keywords. Then, the text vector can be determined through the word vectors of the keywords in the second text, improving the vector representation quality of the text. Clustering the text vectors can avoid directly clustering the second text, and the situation where the second texts describing the same event or behavior are clustered into different clusters due to the differences in the user's expression methods, greatly reducing the influence of the user's expression method on the clustering effect and improving the clustering accuracy.

[0082] Figure 5 It is a schematic structural diagram of a text clustering device provided by an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware. Optionally, it is implemented through an electronic device, and the electronic device can be a mobile terminal, a PC, or a server, etc.

[0083] As Figure 5 shown, the device includes: a text screening module 510 and a clustering module 520.

[0084] A text screening module 510, configured to input a first text set into a preset classification model, and determine a second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set;

[0085] A clustering module 520, configured to construct a keyword network according to keywords in the second text set, generate a plurality of keyword sequences according to the keyword network, determine text vectors of texts in the second text set based on the plurality of keyword sequences, and perform text clustering based on the text vectors.

[0086] Optionally, the text screening module 510 is specifically configured to:

[0087] Input the first text set into a preset classification model, where the first texts in the first text set include feedback texts;

[0088] Determine a category label of the first text through the preset classification model based on preset classification prompt information, where the preset classification prompt information includes category definitions and / or category examples;

[0089] Screen the first text set according to the category label, and determine the second text set according to the screening result.

[0090] Optionally, the clustering module 520 is specifically configured to:

[0091] Construct a keyword network according to the association relationship between keywords in the second text set, where the weight of an edge in the keyword network represents the association strength between the keywords;

[0092] Generate a plurality of keyword sequences according to the association strength between the keywords in the keyword network, and determine word vectors of the keywords in the second text set according to the plurality of keyword sequences;

[0093] Determine text vectors of second texts in the second text set according to the word vectors, and perform clustering on the second texts according to the text vectors.

[0094] Optionally, constructing a keyword network according to the association relationship between keywords in the second text set includes:

[0095] Construct the keyword network according to the co-occurrence relationship between keywords in the second text set;

[0096] Determine the weight of an edge according to the association strength of a keyword pair connected by an edge in the keyword network;

[0097] Update the keyword network according to the weight of the edge;

[0098] Or,

[0099] generating a plurality of keyword sequences according to the association strength between the keywords in the keyword network, including:

[0100] For any keyword in the keyword network, determining at least one neighbor keyword of the current keyword according to the edge corresponding to the current keyword;

[0101] Determining the probability of selecting the neighbor keyword to generate the keyword sequence according to the association strength between the current keyword and each neighbor keyword;

[0102] Selecting the neighbor keyword as the adjacent keyword of the current keyword according to the probability corresponding to each neighbor keyword of the current keyword to generate the keyword sequence;

[0103] If the length of the keyword sequence does not meet the preset length condition, taking the neighbor keyword as the new current keyword, and returning to execute determining the neighbor keyword of the current keyword according to the edge corresponding to the current keyword.

[0104] Optionally, the following method is used to determine the association strength of the keyword pair:

[0105] Determining the association strength of the keyword pair according to the co-occurrence probability of the keyword pair in the second text set and the probability of each keyword in the keyword pair appearing.

[0106] Optionally, the determining the word vectors of the keywords in the second text set according to the plurality of keyword sequences includes:

[0107] Obtaining keyword pairs in the keyword sequence according to a preset sliding window;

[0108] Determining a training sample set according to the keyword pairs in the keyword sequence, and training a preset word vector generation model according to the training sample set;

[0109] Determining the word vectors of the keywords in the second text set according to the trained preset word vector generation model;

[0110] Or,

[0111] The determining the text vector of the second text according to the word vectors includes:

[0112] Determining the text vector of the second text according to the word vectors of the keywords in the second text.

[0113] The text clustering device provided by an embodiment of the present disclosure can execute the text clustering method provided by any embodiment of the present disclosure, and has function modules and beneficial effects corresponding to the execution of the method.

[0114] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0115] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. The following refers to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure (such as Figure 6 in the terminal device or server). The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0116] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The editing / output (I / O) interface 605 is also connected to the bus 604.

[0117] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6An electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0118] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0119] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0120] The electronic device provided in the embodiments of the present disclosure and the text clustering method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0121] The embodiments of the present disclosure provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the text clustering method provided in the above embodiments is implemented.

[0122] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0123] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0124] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0125] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to:

[0126] Input a first text set into a preset classification model, and determine a second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set;

[0127] Construct a keyword network according to the keywords in the second text set, generate a plurality of keyword sequences according to the keyword network, determine the text vectors of the texts in the second text set based on the plurality of keyword sequences, and perform text clustering based on the text vectors.

[0128] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0130] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0131] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on a Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0132] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0133] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the present disclosure.

[0134] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0135] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. A text clustering method, characterized in that, Including: Input the first text set into a preset classification model, and determine a second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set; Construct a keyword network according to the keywords in the second text set, generate a plurality of keyword sequences according to the keyword network, determine the text vectors of the texts in the second text set based on the plurality of keyword sequences, and perform text clustering based on the text vectors.

2. The method according to claim 1, wherein The step of inputting the first text set into a preset classification model and determining a second text set in the first text set based on the preset classification model includes: Input the first text set into a preset classification model, where the first texts in the first text set include feedback texts; Determine the category labels of the first texts through the preset classification model based on preset classification prompt information, where the preset classification prompt information includes category definitions and / or category examples; Filter the first text set according to the category labels, and determine the second text set according to the filtering results.

3. The method according to claim 2, wherein The step of constructing a keyword network according to the keywords in the second text set, generating a plurality of keyword sequences according to the keyword network, determining the text vectors of the texts in the second text set based on the plurality of keyword sequences, and performing text clustering based on the text vectors includes: Construct a keyword network according to the association relationships between the keywords in the second text set, where the weights of the edges in the keyword network represent the association strengths between the keywords; Generate a plurality of keyword sequences according to the association strengths between the keywords in the keyword network, and determine the word vectors of the keywords in the second text set according to the plurality of keyword sequences; Determine the text vectors of the second texts in the second text set according to the word vectors, and cluster the second texts according to the text vectors.

4. The method according to claim 3, wherein The step of constructing a keyword network according to the association relationships between the keywords in the second text set includes: Construct the keyword network according to the co-occurrence relationships between the keywords in the second text set; Determine the weights of the edges according to the association strengths of the keyword pairs connected by the edges in the keyword network; Update the keyword network according to the weights of the edges; Or, The step of generating a plurality of keyword sequences according to the association strengths between the keywords in the keyword network includes: For any keyword in the keyword network, determine at least one neighbor keyword of the current keyword according to the edge corresponding to the current keyword; Determine the probability of selecting the neighbor keyword to generate the keyword sequence according to the association strength between the current keyword and each neighbor keyword; Select the neighbor keyword as the adjacent keyword of the current keyword according to the probability corresponding to each neighbor keyword of the current keyword to generate the keyword sequence; If the length of the keyword sequence does not meet the preset length condition, use the neighbor keyword as the new current keyword, and return to execute the step of determining the neighbor keyword of the current keyword according to the edge corresponding to the current keyword.

5. The method according to claim 4, characterized in that, The association strength of the keyword pair is determined in the following manner: Based on the co-occurrence probability of the keyword pair in the second text set and the occurrence probability of each keyword in the keyword pair, the association strength of the keyword pair is determined.

6. The method according to any one of claims 3-5, characterized in that, The determining of the word vectors of the keywords in the second text set according to the multiple keyword sequences includes: Obtaining keyword pairs in the keyword sequence according to a preset sliding window; Determining a training sample set according to the keyword pairs in the keyword sequence, and training a preset word vector generation model according to the training sample set; Determining the word vectors of the keywords in the second text set according to the trained preset word vector generation model; Or The determining of the text vector of the second text according to the word vectors includes: Determining the text vector of the second text according to the word vectors of the keywords in the second text.

7. A text clustering device, characterized in that, Including: A text screening module, configured to input a first text set into a preset classification model, and determine a second text set in the first text set based on the preset classification model, where the second text set is a set of texts of a preset category in the first text set; A clustering module, configured to construct a keyword network according to the keywords in the second text set, generate multiple keyword sequences according to the keyword network, determine the text vectors of the texts in the second text set based on the multiple keyword sequences, and perform text clustering based on the text vectors.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the text clustering method according to any one of claims 1-6.

9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the text clustering method according to any one of claims 1-6 when executed by a computer processor.

10. A computer program product, comprising a computer program, characterized in that, The computer program implements the text clustering method according to any one of claims 1-6 when executed by a processor.