Clustering Method and Device for Message Text

By extracting word vector feature and clustering the message text, the problem of low clustering accuracy of message text in the prior art is solved, and more accurate and better message text clustering results are achieved.

CN113961701BActive Publication Date: 2025-05-27VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111192804.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-13
Publication Date
2025-05-27
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

The clustering method of existing message texts is low in unsupervised situations, making it difficult to effectively group message texts.

Method used

By obtaining the word vector of the target message text, performing feature extraction, obtaining the feature vector, and clustering it to obtain the clustering probability to obtain the label.

Benefits of technology

This method can more abundantly and accurately represent the context semantic information of the message text, and the obtained feature vectors more accurately represent the features of the message text, thereby obtaining more accurate and better message text clustering results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113961701B_ABST
    Figure CN113961701B_ABST
Patent Text Reader

Abstract

The present application discloses a clustering method and device for message texts, belonging to the field of artificial intelligence technology. Among them, the clustering method for message texts includes: obtaining word vectors corresponding to N target message texts; performing feature extraction on the word vectors to obtain feature vectors of the target message texts; performing clustering processing on the feature vectors of the N target message texts to obtain clustering probabilities of the N target message texts; and obtaining labels corresponding to the target message texts based on the clustering probabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a method and device for clustering message texts. Background Art

[0002] When the message text clustering function of a mobile terminal is enabled, the mobile terminal can upload the message text to a server; based on a trained model, the server can obtain the clustering result of the message text and return the clustering result to the mobile terminal. The mobile terminal can, based on the clustering result of the above message text, integrate and display the above message text, display the aggregation result and the relevance between the above message texts, etc., facilitating management operations such as viewing and editing by users.

[0003] However, although existing methods for clustering message texts can group them according to the semantic relevance of the message texts without supervision, the clustering accuracy is relatively low. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method and device for clustering message texts, which can solve the problem of relatively low accuracy in clustering message texts.

[0005] In a first aspect, the embodiments of this application provide a method for clustering message texts, the method including:

[0006] Obtaining word vectors corresponding to N target message texts;

[0007] Performing feature extraction on the word vectors to obtain feature vectors of the target message texts;

[0008] Performing clustering processing on the feature vectors of the N target message texts to obtain the clustering probabilities of the N target message texts;

[0009] Obtaining labels corresponding to the target message texts based on the clustering probabilities.

[0010] In a second aspect, the embodiments of this application provide a device for clustering message texts, the device including:

[0011] An obtaining module, configured to obtain word vectors corresponding to N target message texts;

[0012] An extraction module, configured to perform feature extraction on the word vectors to obtain feature vectors of the target message texts;

[0013] A clustering module, configured to perform clustering processing on the feature vectors of the N target message texts to obtain the clustering probabilities of the N target message texts;

[0014] Obtaining labels corresponding to the target message texts based on the clustering probabilities.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0017] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the method described in the first aspect.

[0018] In the embodiment of the present application, by performing feature extraction on the word vectors corresponding to the target message text to obtain the feature vectors of the target message text, performing clustering processing on the feature vectors of N target message texts to obtain the clustering probabilities of the N target message texts, and obtaining the labels corresponding to the target message text based on the clustering probabilities, the context semantic information of the target message text can be represented more richly and accurately. The obtained feature vectors of the target message text can more accurately and comprehensively characterize the features of the target message text, and a more accurate and better message text clustering result can be obtained based on the feature vectors of the target message text. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is one of the flow diagrams of the message text clustering method provided by the embodiment of the present application;

[0020] Figure 2 is a schematic diagram of the pyramid module provided by the embodiment of the present application;

[0021] Figure 3 is another flow diagram of the message text clustering method provided by the embodiment of the present application;

[0022] Figure 4 is a schematic structural diagram of the message text clustering device provided by the embodiment of the present application;

[0023] Figure 5 is a schematic structural diagram of the electronic device provided by the embodiment of the present application;

[0024] Figure 6 is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0026] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0027] Next, in conjunction with the accompanying drawings, a clustering method and device for message texts provided in the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0028] Figure 1 is one of the flow diagrams of the clustering method for message texts provided in the embodiments of the present application. Next, in conjunction with Figure 1 describe the clustering method for message texts provided in the embodiments of the present application. As Figure 1 shown, the method includes:

[0029] Step 101, obtain word vectors corresponding to N target message texts. Optionally, the execution subject of the clustering method for message texts provided in the embodiments of the present application is a clustering device for message texts.

[0030] The clustering device for message texts can be implemented in various forms. For example, the clustering device for message texts described in the embodiments of the present application can include mobile terminals such as mobile phones, smart phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), navigation devices, smart bracelets, smart watches, digital cameras, etc., and fixed terminals such as desktop computers. Below, it is assumed that the clustering device for message texts is a fixed terminal. However, those skilled in the art will understand that in scenarios specifically for mobile purposes, the configuration according to the embodiments of the present application can also be applied to mobile-type terminals.

[0031] Optionally, the target message text can be the message text received by an application (APP, Application) with message receiving functionality. Exemplarily, an application with message receiving functionality can be a text message application or any instant messaging application.

[0032] Optionally, based on any one of the word vector generation methods in natural language processing (NLP, Natural Language Processing), text word features can be extracted from each of the N target message texts respectively to obtain the word vector corresponding to the target message text.

[0033] Where N is a positive integer greater than 1.

[0034] A word vector (Word embedding) is a vectorized representation of a word, where a word or phrase from a vocabulary is mapped to a vector of real numbers.

[0035] The methods for generating word vectors mainly include two categories: statistic-based methods and language model (LanguageModel)-based methods.

[0036] Step 102: Extract features from the word vectors to obtain the feature vectors of the target message texts.

[0037] Optionally, for each target message text, a feature extraction model can be used to extract features from the word vector corresponding to the target message text to obtain the feature vector of the target message text.

[0038] Exemplarily, the word vectors corresponding to each target message text can be respectively input into a pre-trained feature extraction model, and the feature extraction model extracts features from the word vector corresponding to the target message text to obtain the feature vector of the target message text output by the feature extraction model.

[0039] The feature extraction model is a model obtained through pre-training. The feature extraction model is used to extract the context semantic information in the target message text based on the word vector corresponding to the target message text. The context semantic information in the target message text can describe the features of the target message text. Therefore, the vectorized representation of the context semantic information in the target message text can be used as the feature vector of the target message text.

[0040] Step 103: Perform clustering processing on the feature vectors of the N target message texts to obtain the clustering probabilities of the N target message texts.

[0041] Optionally, a clustering algorithm can be used to perform clustering processing on the feature vectors of the N target message texts to obtain the clustering probability of each target message text.

[0042] Exemplarily, the clustering algorithm may be a K-Means clustering algorithm (K-means clustering algorithm) or a SOM (Self-Organizing Maps) clustering algorithm, etc.

[0043] The clustering probability of the target message text is used for the similarity between the feature vector of the target message text and each clustering center.

[0044] Optionally, the clustering algorithm may be implemented through a clustering model obtained by pre-training. The feature vectors of N target message texts may be input into the clustering model, and the clustering model performs clustering processing on the input N feature vectors to obtain the clustering probabilities of the N target message texts output by the clustering model.

[0045] Optionally, a joint training method may be adopted to jointly train the feature extraction model and the clustering model based on the word vectors corresponding to the sample message texts.

[0046] Joint training is to jointly regard multiple models as a whole model through a certain method for training, which can better obtain the global optimal solution to complete a certain task.

[0047] It can be understood that the joint training may include multiple rounds of training.

[0048] Optionally, the idea of joint training is to regard the optimization problem in the feature extraction process as minimizing the distance between the probability distributions of the training data (i.e., the word vectors corresponding to the sample message texts) before and after training, and to make the clustering result guide the training of the feature extraction model parameters, so that the probability distribution of the predicted training data after training is closer to the probability distribution of the training data before training.

[0049] Optionally, the KL divergence may be used in the joint training to measure the distance between the probability distributions of the training data before and after each round of training. That is, the specific formula of the loss function Loss is as follows:

[0050]

[0051] where KL represents the KL divergence; P represents the probability distribution of the training data before training; Q represents the predicted probability distribution of the training data after training; q ij represents the similarity between the word vector z i corresponding to the i-th sample message text before training and the j-th clustering center μ j ; p ij represents the similarity between the word vector z i corresponding to the i-th sample message text predicted after training and the j-th clustering center μ j .

[0052] p ij can be calculated by the following formula:

[0053]

[0054] where j'≠j.

[0055] Optionally, the t-distribution can be used to measure the similarity between the word vector z corresponding to the i-th sample message text i and the j-th cluster center μ j The specific formula is as follows:

[0056]

[0057] where β is a preset constant; μ j' represents the j'-th cluster center. The embodiment of the present application does not specifically limit the value of β. Exemplarily, β can be set to 1.

[0058] After obtaining the distance between the probability distributions of the training data before and after each round of training, the gradient descent method can be used to optimize the word vector z corresponding to the i-th sample message text i and the j-th cluster center μ j The specific formula is as follows:

[0059]

[0060]

[0061] Step 104: Obtain the label corresponding to the target message text based on the clustering probability.

[0062] Optionally, for each target message text, the maximum value in the clustering probability of the target message text can be determined; the number of the cluster center corresponding to the maximum value is determined as the label corresponding to the target message text.

[0063] In the embodiment of the present application, by extracting the features of the word vector corresponding to the target message text, the feature vector of the target message text is obtained, the feature vectors of N target message texts are clustered, the clustering probabilities of N target message texts are obtained, and the label corresponding to the target message text is obtained based on the clustering probability, which can represent the context semantic information of the target message text more richly and accurately. The obtained feature vector of the target message text can more accurately and comprehensively characterize the features of the target message text, and a more accurate and better message text clustering result can be obtained based on the feature vector of the target message text.

[0064] Optionally, feature extraction is performed on the word vectors to obtain the feature vectors of the target message text, including: based on the self-attention mechanism, performing feature extraction on the word vectors to obtain the self-attention feature map corresponding to the target message text.

[0065] Optionally, the feature extraction model may include a self-attention layer and a multi-channel feature extraction layer. The self-attention layer is used for the first feature extraction; the multi-channel feature extraction layer is used for the second feature extraction.

[0066] For the word vectors corresponding to each target message text, they can be input into the self-attention layer; the self-attention layer learns the dependencies between the words within the sentence of the target message text based on the self-attention mechanism, obtains the internal structure and key semantic information of the sentence, and thus obtains the self-attention feature map (self-attention map) corresponding to the target message text.

[0067] Based on the multi-channel feature extraction layer in the feature extraction model, feature extraction is performed on the self-attention feature map corresponding to the target message text to obtain the feature vectors of the target message text.

[0068] Among them, the multi-channel feature extraction layer includes M feature extraction networks; M is a positive integer greater than 1; at least one feature extraction network is a pyramid convolutional neural network.

[0069] Optionally, the multi-channel feature extraction layer includes M parallel feature extraction networks. Each feature extraction network can serve as a channel for performing feature extraction on the self-attention feature map corresponding to the target message text to obtain the features of the self-attention feature map.

[0070] Input the self-attention feature map corresponding to the target message text into the multi-channel feature extraction layer, and the M feature extraction networks respectively output the features of the self-attention feature map; performing fusion processing on the features of the self-attention feature map obtained by the M feature extraction networks can obtain the feature vectors of the target message text.

[0071] Optionally, at least one of the above M feature extraction networks can be a pyramid convolutional neural network.

[0072] The pyramid convolutional neural network can be composed of multiple pyramid modules. The pyramid module can include a convolutional layer and a pooling layer, and the specific structure can be as Figure 2 shown.

[0073] Exemplarily, Figure 2 shows a pyramid module 200, which includes a convolutional layer 201 and a pooling layer 202. The convolutional layer 201 includes two equal-length convolutional sub-layers (equal-length convolutional sub-layer 2011 and equal-length convolutional sub-layer 2012) and a residual block 2013.

[0074] Exemplarily, the number of channels of both of the two same-length convolutional sub-layers can be set to 64. The convolutional manner of the same-length convolution can better extract the features of the word vectors, making their representations richer and more accurate.

[0075] The residual block 2013 is used to alleviate the problem of gradient dispersion caused by the increase in network depth.

[0076] Exemplarily, the size of the pooling layer 202 is 3 and the stride is 2. Each time the pooling layer 202 is passed through, the sequence length of the output of the convolutional layer can be halved.

[0077] In the pyramid convolutional neural network, by stacking multiple pyramid modules, the dimension of the final output feature vector can be made to be 1.

[0078] Exemplarily, Figure 3 is the second flow schematic diagram of the clustering method for message texts provided by the embodiments of the present application. As Figure 3 shown, the word vector 302 corresponding to the target message text 301 passes through the self-attention layer 3031 and the multi-channel feature extraction layer 3032 in the feature extraction model 303 in sequence, and a feature vector 304 can be obtained; the feature vector 304 is input into the clustering model 305, and a clustering label 306 can be obtained. The multi-channel feature extraction layer 3032 can be composed of 3 pyramid convolutional neural networks 30320. The sizes of the 3 pyramid convolutional neural networks 30320 are different and can be set to 3, 4, and 5 respectively.

[0079] The pyramid convolutional neural network can represent the context semantic information of the target message text more richly and accurately, and the obtained feature vector of the target message text can more accurately and comprehensively characterize the features of the target message text.

[0080] By fusing the self-attention mechanism and the multi-channel feature extraction layer, the embodiments of the present application can represent the context semantic information of the target message text more richly and accurately, and the obtained feature vector of the target message text can more accurately and comprehensively characterize the features of the target message text, so that a more accurate message text clustering result can be obtained based on the feature vector of the target message text.

[0081] Optionally, based on the multi-channel feature extraction layer in the feature extraction model, feature extraction is performed on the self-attention feature map corresponding to the target message text to obtain the feature vector of the target message text, including: respectively performing feature extraction on the self-attention feature map corresponding to the target message text based on each feature extraction network to obtain M first sub-vectors.

[0082] Optionally, the structures of the M feature extraction networks are different, so the M feature extraction networks can perform feature extraction on the self-attention feature map from different angles.

[0083] For each self-attention feature map corresponding to the target message text, the self-attention feature map can be input into each feature extraction network respectively to obtain the first sub-vectors output by each feature extraction network.

[0084] Concatenate the M first sub-vectors to obtain the feature vector of the target message text.

[0085] Optionally, after obtaining the M first sub-vectors, perform a concatenation process on the M first sub-vectors to obtain a vector, which is the feature vector of the target message text.

[0086] Optionally, the M first vectors can be directly concatenated to obtain the feature vector of the target message text.

[0087] Optionally, after performing max pooling on the M first vectors respectively, concatenate the M first vectors after max pooling to obtain the feature vector of the target message text.

[0088] The functions of pooling processing can include: (1) implementing downsampling; (2) reducing the number of parameters (reducing the dimension) and computational complexity while retaining the main features, preventing overfitting, removing redundant information, compressing the features, simplifying the network complexity and reducing memory consumption, etc.; (3) implementing non-linearity; (4) expanding the receptive field; (5) implementing invariance, where the invariance includes: translational invariance, rotational invariance, and scale invariance.

[0089] Max pooling can reduce the shift of the estimated mean caused by the parameter error of the convolutional layer, thereby reducing the error of feature extraction. Moreover, max pooling processing can eliminate non-maximum values, thereby reducing the computational complexity of the upper layer.

[0090] In the embodiments of the present application, multiple feature extraction networks are used to respectively perform feature extraction on the self-attention feature map corresponding to the target message text to obtain M first vectors, and the M first vectors are concatenated to obtain the feature vector of the target message text, which can more richly and accurately represent the context semantic information of the target message text, and the obtained feature vector of the target message text can more accurately and comprehensively characterize the features of the target message text, so that a more accurate message text clustering result can be obtained based on the feature vector of the target message text.

[0091] Optionally, based on the self-attention mechanism, perform feature extraction on the word vectors to obtain the self-attention feature map corresponding to the target message text, including: based on the self-attention mechanism, transform the word vectors to obtain the first vector, the second vector, and the third vector.

[0092] Optionally, for the word vector x corresponding to the i-th target message text i, the word vector x can be converted into Query, Key, and Value vectors Q i , K i i , V i .

[0093] The specific steps of the conversion include:

[0094] Let Q i = x i , K i = x i , V i = x i (6)

[0095] Among them, the vectors Q i , K i , V i are the first vector, the second vector, and the third vector respectively.

[0096] Perform linear transformations on the first vector, the second vector, and the third vector to obtain the fourth vector, the fifth vector, and the sixth vector.

[0097] Optionally, perform linear transformations on the vectors Q i , K i , V i to obtain the transformed Query, Key, and Value vectors Q' i , K' i , V' i . The vectors Q' i , K' i , V' i are the fourth vector, the fifth vector, and the sixth vector respectively.

[0098] The formula for the linear transformation is as follows:

[0099] Q' i = Q i * W Q (7)

[0100] K' i = K i * W K (8)

[0101] V' i = V i * W V (9)

[0102] Among them, W Q , W K , W V are the weight matrices corresponding to the Query, Key, and Value vectors respectively. ​

[0103] Calculate the similarity between the fourth vector and the fifth vector, and normalize the similarity to obtain a weight vector.

[0104] Optionally, the calculation formula of the attention weight vector α is as follows:

[0105]

[0106] where α i is the attention weight vector corresponding to the word vector x i ; d k represents the dimension of the word vector. The dimension d k of the word vector can be a preset value.

[0107] In formula (10), is the similarity between the fourth vector Q′ i and the fifth vector K′ i .

[0108] Multiply the weight vector by the sixth vector to obtain a self-attention feature map.

[0109] Optionally, the calculation formula of the self-attention feature map is as follows:

[0110] attention(query,key,value) = α i * V′ i (11)

[0111] In the embodiments of the present application, by extracting features from the word vectors corresponding to the target message text, learning the dependency relationships between the words within the sentence in the target message text, and obtaining the internal structure and key semantic information of the sentence, a self-attention feature map corresponding to the target message text is obtained. The obtained self-attention feature map can more richly and accurately represent the context semantic information of the target message text.

[0112] Optionally, obtaining the word vectors corresponding to N target message texts includes: performing word segmentation on each target message text to obtain the word segmentation result of the target message text.

[0113] Optionally, for each target message text, the target message text can be segmented based on a word segmentation tool or method to obtain the word segmentation result of the target message text.

[0114] The embodiments of the present application do not limit the specific word segmentation tool used. Exemplarily, the word segmentation tool can be an open-source word segmentation tool such as Jieba or LTP.

[0115] The embodiments of the present application do not limit the specific word segmentation method used. Optionally, when the text used in the target message text is Chinese, the Chinese word segmentation methods mainly include two categories: dictionary-based word segmentation algorithms and statistical machine learning algorithms. Any Chinese word segmentation method can be used to perform word segmentation processing on the target message text.

[0116] For the target message text after word segmentation processing, stop words can be removed according to the stop word dictionary to obtain the word segmentation result of the target message text.

[0117] Based on the word segmentation result of the target message text, obtain the word vector corresponding to the target message text.

[0118] Optionally, based on any word vector generation method, text word features can be extracted from the word segmentation result of the target message text to obtain the word vector corresponding to the target message text.

[0119] Exemplarily, the word segmentation result of the target message text can be input into the BERT pre-trained model. The BERT pre-trained model learns the text context semantic information of the word segmentation result of the target message text and represents the text context semantic information of the word segmentation result of the target message text in the form of a word vector to obtain the word vector corresponding to the target message text.

[0120] The embodiments of the present application perform word segmentation processing on the target message text to obtain the word segmentation result of the target message text. Based on the word segmentation result of the target message text, the word vector corresponding to the target message text is obtained, and a more reasonable word vector can be obtained, so that a more accurate message text clustering result can be obtained based on the word vector corresponding to the target message text.

[0121] It should be noted that for the message text clustering method provided by the embodiments of the present application, the execution subject can be a message text clustering device, or a control module in the message text clustering device for executing the message text clustering method. In the embodiments of the present application, the message text clustering method is executed by the message text clustering device as an example to illustrate the message text clustering device provided by the embodiments of the present application.

[0122] Figure 4 It is a schematic structural diagram of the message text clustering device provided by the embodiments of the present application. Optionally, as Figure 4 shown, the device includes a first acquisition module 401, an extraction module 402, a clustering module 403, and a second acquisition module 404, where:

[0123] The first acquisition module 401 is configured to acquire word vectors corresponding to N target message texts;

[0124] An extraction module 402, configured to perform feature extraction on the word vectors corresponding to each target message text based on a feature extraction model, to obtain the feature vectors of each target message text;

[0125] A clustering module 403, configured to perform clustering processing on the feature vectors of N target message texts by a clustering model, to obtain the clustering probabilities of the N target message texts;

[0126] A second acquisition module 404, configured to obtain the label corresponding to the target message text based on the clustering probability.

[0127] Optionally, the first acquisition module 401, the extraction module 402, the clustering module 403, and the second acquisition module 404 are electrically connected in sequence.

[0128] The first acquisition module 401 may perform text word feature extraction on each target message text respectively based on any one of the word vector generation methods, to obtain the word vector corresponding to the target message text.

[0129] For each target message text, the extraction module 402 may perform feature extraction on the word vector corresponding to the target message text through any one of the feature extraction methods, to obtain the feature vector of the target message text.

[0130] The clustering module 403 may perform clustering processing on N feature vectors through any one of the clustering algorithms, to obtain the clustering probability of each target message text.

[0131] For each target message text, the second acquisition module 404 may determine the maximum value in the clustering probability of the target message text; and determine the number of the clustering center corresponding to the maximum value as the label corresponding to the target message text.

[0132] Optionally, the extraction module 402 may include:

[0133] A first extraction unit, configured to perform feature extraction on the word vector based on the self-attention mechanism, to obtain the self-attention feature map corresponding to the target message text;

[0134] A second extraction unit, configured to perform feature extraction on the self-attention feature map corresponding to the target message text based on the multi-channel feature extraction layer in the feature extraction model, to obtain the feature vector of the target message text.

[0135] Wherein, the multi-channel feature extraction layer includes M feature extraction networks; M is a positive integer greater than 1; and at least one feature extraction network is a pyramid convolutional neural network.

[0136] Optionally, the second extraction unit may include:

[0137] An extraction subunit, configured to perform feature extraction on the self-attention feature map corresponding to the target message text respectively based on each feature extraction network, so as to obtain M first sub-vectors;

[0138] A splicing subunit, configured to splice and process the M first sub-vectors to obtain a feature vector of the target message text.

[0139] Optionally, the first extraction unit may include:

[0140] A vector conversion subunit, configured to convert word vectors based on the self-attention mechanism to obtain a first vector, a second vector, and a third vector;

[0141] A linear transformation subunit, configured to perform linear transformation on the first vector, the second vector, and the third vector to obtain a fourth vector, a fifth vector, and a sixth vector;

[0142] A first obtaining subunit, configured to calculate the similarity between the fourth vector and the fifth vector and perform normalization processing on the similarity to obtain a weight vector;

[0143] A second obtaining subunit, configured to multiply the weight vector by the sixth vector to obtain a self-attention feature map.

[0144] Optionally, the first obtaining module 401 may include:

[0145] A word segmentation unit, configured to perform word segmentation processing on each target message text to obtain a word segmentation result of the target message text;

[0146] An obtaining unit, configured to obtain word vectors corresponding to the target message text based on the word segmentation result of the target message text.

[0147] In the embodiment of the present application, feature extraction is performed on the word vectors corresponding to the target message text to obtain feature vectors of the target message text, clustering processing is performed on the feature vectors of N target message texts to obtain clustering probabilities of the N target message texts, and labels corresponding to the target message text are obtained based on the clustering probabilities, which can represent the context semantic information of the target message text more richly and accurately. The obtained feature vectors of the target message text can more accurately and comprehensively characterize the features of the target message text, and more accurate and better message text clustering results can be obtained based on the feature vectors of the target message text.

[0148] The clustering device for message texts in the embodiments of this application can be a device, or a component, integrated circuit, or chip in a terminal. This device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, tablet computer, laptop, palmtop computer, in-vehicle electronic device, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, Network Attached Storage (NAS), personal computer (PC), television (TV), teller machine, or self-service machine, etc. The embodiments of this application do not make specific limitations.

[0149] The clustering device for message texts in the embodiments of this application can be a device with an operating system. This operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of this application do not make specific limitations.

[0150] The clustering device for message texts provided in the embodiments of this application can implement Figures 1 to 3 each process implemented by the method embodiments. To avoid repetition, details are not described herein again.

[0151] Optionally, as Figure 5 shown, the embodiments of this application also provide an electronic device 500, including a processor 501, a memory 502, a program or instruction stored on the memory 502 and executable on the processor 501. When the program or instruction is executed by the processor 501, it implements each process of the above-mentioned clustering method embodiments for message texts and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0152] It should be noted that the electronic devices in the embodiments of this application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0153] Figure 6 is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application.

[0154] The electronic device 600 includes but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610, etc.

[0155] Those skilled in the art can understand that the electronic device 600 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 610 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 6 The structure of the electronic device shown in Figure 6 does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements, which will not be elaborated here.

[0156] Among them, the processor 610 is used to obtain word vectors corresponding to N target message texts;

[0157] The processor 610 is further used to perform feature extraction on the word vectors to obtain feature vectors of the target message texts;

[0158] The processor 610 is further used to perform clustering processing on the feature vectors of N target message texts to obtain clustering probabilities of the N target message texts;

[0159] The processor 610 is further used to obtain labels corresponding to the target message texts based on the clustering probabilities.

[0160] In the embodiment of the present application, feature extraction is performed on the word vectors corresponding to the target message texts to obtain feature vectors of the target message texts. Clustering processing is performed on the feature vectors of N target message texts to obtain clustering probabilities of the N target message texts. Labels corresponding to the target message texts are obtained based on the clustering probabilities, which can more richly and accurately represent the context semantic information of the target message texts. The obtained feature vectors of the target message texts can more accurately and comprehensively characterize the features of the target message texts, and more accurate and better message text clustering results can be obtained based on the feature vectors of the target message texts.

[0161] Optionally, the processor 610 is further used to perform feature extraction on the word vectors based on the self-attention mechanism to obtain a self-attention feature map corresponding to the target message text;

[0162] The processor 610 is further used to perform feature extraction on the self-attention feature map corresponding to the target message text based on a multi-channel feature extraction layer in the feature extraction model to obtain a feature vector of the target message text;

[0163] Among them, the multi-channel feature extraction layer includes M feature extraction networks; M is a positive integer greater than 1; at least one feature extraction network is a pyramid convolutional neural network.

[0164] Optionally, the processor 610 is further used to perform feature extraction on the self-attention feature map corresponding to the target message text based on each feature extraction network respectively to obtain M first sub-vectors;

[0165] The processor 610 is further configured to splice the M first sub-vectors to obtain a feature vector of the target message text.

[0166] Optionally, the processor 610 is further configured to transform the word vectors based on the self-attention mechanism to obtain a first vector, a second vector, and a third vector;

[0167] The processor 610 is further configured to perform a linear transformation on the first vector, the second vector, and the third vector to obtain a fourth vector, a fifth vector, and a sixth vector;

[0168] The processor 610 is further configured to calculate the similarity between the fourth vector and the fifth vector and normalize the similarity to obtain a weight vector;

[0169] The processor 610 is further configured to multiply the weight vector by the sixth vector to obtain a self-attention feature map.

[0170] Optionally, the processor 610 is further configured to perform word segmentation processing on each target message text to obtain a word segmentation result of the target message text;

[0171] The processor 610 is further configured to obtain word vectors corresponding to the target message text based on the word segmentation result of the target message text.

[0172] It should be understood that in the embodiments of the present application, the input unit 604 may include a Graphics Processing Unit (GPU) 6041 and a microphone 6042. The graphics processor 6041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 607 includes a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. The other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here. The memory 609 may be used to store software programs and various data, including but not limited to target application programs and operating systems. The processor 610 may integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, the user interface, and target application programs, and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 610.

[0173] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the embodiment of the above message text clustering method is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0174] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0175] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the embodiment of the above message text clustering method, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0176] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, a system chip, a chip system, or a system-on-chip.

[0177] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0179] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A clustering method for message texts, characterized in that, comprising: Obtaining word vectors corresponding to N target message texts; Performing feature extraction on the word vectors to obtain feature vectors of the target message texts; Performing clustering processing on the feature vectors of the N target message texts to obtain clustering probabilities of the N target message texts; Obtaining labels corresponding to the target message texts based on the clustering probabilities; Wherein, the performing feature extraction on the word vectors to obtain feature vectors of the target message texts includes: Performing feature extraction on the word vectors based on the self-attention mechanism to obtain a self-attention feature map corresponding to the target message texts; Performing feature extraction on the self-attention feature map corresponding to the target message texts based on a multi-channel feature extraction layer in the feature extraction model to obtain feature vectors of the target message texts; Wherein, the multi-channel feature extraction layer includes M feature extraction networks; M is a positive integer greater than 1; at least one of the feature extraction networks is a pyramid convolutional neural network, and the pyramid convolutional neural network is composed of multiple pyramid modules, and the pyramid module includes two convolutional sub-layers of equal length and a residual block.

2. The clustering method for message texts according to claim 1, characterized in that, The performing feature extraction on the self-attention feature map corresponding to the target message texts based on a multi-channel feature extraction layer in the feature extraction model to obtain feature vectors of the target message texts includes: Respectively performing feature extraction on the self-attention feature map corresponding to the target message texts based on each of the feature extraction networks to obtain M first sub-vectors; Performing splicing processing on the M first sub-vectors to obtain feature vectors of the target message texts.

3. The clustering method for message texts according to claim 1, characterized in that, The performing feature extraction on the word vectors based on the self-attention mechanism to obtain a self-attention feature map corresponding to the target message texts includes: Based on the self-attention mechanism, performing transformation on the word vectors to obtain a first vector, a second vector, and a third vector; Performing linear transformation on the first vector, the second vector, and the third vector to obtain a fourth vector, a fifth vector, and a sixth vector; Calculating the similarity between the fourth vector and the fifth vector, and performing normalization processing on the similarity to obtain a weight vector; Multiplying the weight vector by the sixth vector to obtain the self-attention feature map.

4. The clustering method for message texts according to any one of claims 1 to 3, characterized in that, The obtaining word vectors corresponding to N target message texts includes: Performing word segmentation processing on each of the target message texts to obtain word segmentation results of the target message texts; Based on the word segmentation results of the target message texts, obtaining word vectors corresponding to the target message texts.

5. A clustering device for message texts, characterized in that, comprising: A first obtaining module for obtaining word vectors corresponding to N target message texts; An extraction module for performing feature extraction on the word vectors to obtain feature vectors of the target message texts; A clustering module, configured to perform clustering processing on the feature vectors of the N target message texts to obtain the clustering probabilities of the N target message texts; A second obtaining module, configured to obtain the labels corresponding to the target message texts based on the clustering probabilities; Wherein, the extraction module includes: A first extraction unit, configured to perform feature extraction on the word vectors based on the self-attention mechanism to obtain a self-attention feature map corresponding to the target message text; A second extraction unit, configured to perform feature extraction on the self-attention feature map corresponding to the target message text based on a multi-channel feature extraction layer in the feature extraction model to obtain the feature vectors of the target message text; Wherein, the multi-channel feature extraction layer includes M feature extraction networks; M is a positive integer greater than 1; at least one of the feature extraction networks is a pyramid convolutional neural network, and the pyramid convolutional neural network is composed of multiple pyramid modules, and each pyramid module includes two convolutional sub-layers of equal length and a residual block.

6. The clustering device for message texts according to claim 5, wherein, the second extraction unit includes: An extraction subunit, configured to perform feature extraction on the self-attention feature map corresponding to the target message text based on each of the feature extraction networks to obtain M first sub-vectors; A splicing subunit, configured to perform splicing processing on the M first sub-vectors to obtain the feature vectors of the target message text.

7. The clustering device for message texts according to claim 5, wherein, the first extraction unit includes: A vector conversion subunit, configured to perform conversion on the word vectors based on the self-attention mechanism to obtain a first vector, a second vector, and a third vector; A linear transformation subunit, configured to perform linear transformation on the first vector, the second vector, and the third vector to obtain a fourth vector, a fifth vector, and a sixth vector; A first obtaining subunit, configured to calculate the similarity between the fourth vector and the fifth vector and perform normalization processing on the similarity to obtain a weight vector; A second obtaining subunit, configured to multiply the weight vector by the sixth vector to obtain the self-attention feature map.

8. The clustering device for message texts according to any one of claims 5 to 7, wherein, the first obtaining module includes: A word segmentation unit, configured to perform word segmentation processing on each of the target message texts to obtain the word segmentation results of the target message texts; An obtaining unit, configured to obtain the word vectors corresponding to the target message texts based on the word segmentation results of the target message texts.

9. An electronic device, wherein, it includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, the steps of the message text clustering method according to any one of claims 1-4 are implemented.

10. A readable storage medium, wherein, a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the message text clustering method according to any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Text classification algorithm based on self-attention mechanism and convolutional neural network

    CN113297380A

  • Text classification method based on self-attention mechanism and BiGRU

    CN113312483A