Community message classification method, device and equipment and storage medium

By acquiring messages to be corrected and generating similar category groups, updating historical datasets, and retraining the community message classification model, the problems of low efficiency and high cost in existing technologies are solved, achieving automated maintenance and improved accuracy.

CN115994273BActive Publication Date: 2026-05-05GUANGZHOU SENJI SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU SENJI SOFTWARE TECH CO LTD
Filing Date
2023-01-31
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing methods for classifying community messages are inefficient and costly. As the volume and types of messages increase, the classification ability of message classification models decreases significantly, requiring a large amount of manpower for maintenance and updates to ensure accuracy.

Method used

By acquiring messages to be corrected, generating similar category groups based on a comparison of their categories before and after correction, updating the historical message dataset, and retraining the community message classification model, the classification model can be automatically maintained and improved.

Benefits of technology

It reduces manual input costs, improves the accuracy of community message classification, and enables automated maintenance and continuous updating of the classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115994273B_ABST
    Figure CN115994273B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, device, and storage medium for classifying social media messages. The method includes: acquiring at least one message to be corrected; the message to be corrected includes a message category before correction and a message category after correction; the message category before correction is obtained based on the output of a social media message classification model; the social media message classification model is trained based on historical messages in a historical message dataset; selecting a target message to be corrected from the at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected; generating at least one candidate similarity category group; updating the historical message dataset based on the consistency between the target message in the candidate similarity category group and the historical messages in the historical message dataset; and training the social media message classification model using the updated historical message dataset to update the social media message classification model for message classification. This invention improves the accuracy of social media message classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to a method, apparatus, device, and storage medium for classifying social media messages. Background Technology

[0002] In the field of community management technology, it is necessary for community members to follow up with a group of users simultaneously. The density of users served within a given time frame is high, resulting in a large volume of messages. Therefore, it is essential to tag user messages and categorize them according to these tags to help community members quickly understand users and provide targeted responses or marketing.

[0003] Existing methods for classifying online communities include manual annotation, but this method is inefficient and costly. Another method involves automated classification using pre-trained message classification models. However, as the volume and types of messages increase daily, the classification ability of these models declines significantly, requiring higher human resources to maintain and update them to ensure accuracy. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and storage medium for classifying social media messages, thereby automating the maintenance of classification models used for classifying social media messages, reducing manual input costs, and improving the accuracy of classification models in classifying social media messages.

[0005] According to one aspect of the present invention, a method for classifying social media messages is provided, the method comprising:

[0006] Obtain at least one message to be corrected; the message to be corrected includes a message category before correction and a message category after correction; the message category before correction is obtained based on the output of a community message classification model; the community message classification model is trained based on historical messages in a historical message dataset;

[0007] Based on the comparison of the message categories before and after correction among the messages to be corrected, a target message to be corrected is selected from at least one message to be corrected;

[0008] Generate at least one candidate similarity category group; wherein the message similarity of the target correction messages in the same candidate similarity category group is adjacent.

[0009] The historical message dataset is updated based on the consistency between the target correction message in the candidate similarity category group and the historical messages in the historical message dataset.

[0010] The updated historical message dataset is used to train the community message classification model to update the community message classification model for message classification.

[0011] According to another aspect of the present invention, a social media message classification device is provided, characterized in that it comprises:

[0012] The message to be corrected acquisition module is used to acquire at least one message to be corrected; the message to be corrected includes a message category before correction and a message category after correction; the message category before correction is obtained based on the output of a community message classification model; the community message classification model is trained based on historical messages in a historical message dataset;

[0013] The target correction message selection module is used to select a target correction message from at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected.

[0014] A category group generation module is used to generate at least one candidate similar category group; wherein the message similarity of the target correction messages in the same candidate similar category group is adjacent.

[0015] The dataset update module is used to update the historical message dataset based on the consistency between the target correction messages in the candidate similar category group and the historical messages in the historical message dataset;

[0016] The message classification module is used to train the community message classification model with the updated historical message dataset to update the community message classification model for message classification.

[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the social message classification method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the social message classification method according to any embodiment of the present invention.

[0022] This invention provides a solution that obtains at least one message to be corrected, selects a target message to be corrected from the at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected, generates at least one candidate similar category group, updates the historical message dataset based on the consistency between the target message to be corrected in the candidate similar category group and the historical messages in the historical message dataset, and trains a community message classification model using the updated historical message dataset to update the community message classification model for message classification. This achieves automated maintenance of the classification model used for community message classification, reducing manual input costs. By updating the community message classification model based on the consistency between the target message to be corrected in the candidate similar category group and the historical messages in the historical message dataset, the community message classification model is continuously updated and improved, thereby improving the classification accuracy of the community message classification model and thus improving the classification accuracy of community messages.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a social media message classification method provided in Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart of a social media message classification method provided in Embodiment 2 of the present invention;

[0027] Figure 3 This is a flowchart of a social media message classification method provided in Embodiment 3 of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of a social media message classification device according to Embodiment 4 of the present invention;

[0029] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the social message classification method of this invention. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] Example 1

[0033] Figure 1 This is a flowchart of a social media message classification method provided in Embodiment 1 of the present invention. This embodiment is applicable to the case of automated social media message classification. The method can be executed by a social media message classification device, which can be implemented in hardware and / or software and can be configured in an electronic device.

[0034] S110. Obtain at least one message to be corrected; the message to be corrected includes the message category before correction and the message category after correction; the message category before correction is obtained based on the output of the community message classification model; the community message classification model is trained based on historical messages in the historical message dataset.

[0035] For example, at least one message to be corrected within a preset time period can be obtained. The preset time period can be pre-set by relevant technical personnel. For example, the preset time period can be 24 hours. The message to be corrected can be a social media message whose message category needs to be corrected. The message to be corrected includes a message category before correction and a message category after correction. The message category before correction is obtained based on the output of a social media message classification model; the message category after correction is obtained by relevant technical personnel manually correcting the message category of the message to be corrected.

[0036] It should be noted that the social media message classification model is used to classify at least one social media message. Specifically, it can be pre-trained by relevant technical personnel based on historical messages in a historical message dataset. After classifying the social media messages, the model obtains the corresponding message categories. For example, message categories may include inquiry messages, greeting messages, and news messages. It is understandable that the classification accuracy of the social media message classification model is related to the training set used to train the model; the larger the number of sample messages in the training set and the more accurate the message data, the higher the training accuracy of the social media message classification model.

[0037] It is understandable that when using message classification models to classify community messages, there may be instances where the output message categories are inaccurate. Therefore, relevant technical personnel can treat these inaccurately classified community messages as messages to be corrected. The original message category of the message to be corrected is the inaccurate message category output by the community message classification model. The corrected message category is the message category manually corrected by relevant technical personnel.

[0038] It should be noted that the purpose of manually determining the corrected message category of the messages to be corrected within the preset time period is to subsequently use the corrected message category of a small number of manually corrected messages to batch correct the historical message data in the historical message dataset used for the community message classification model, and to continuously introduce accurate community messages into the historical message dataset to increase the number of sample training sets, thereby improving the training accuracy of the community message classification model and thus improving the accuracy of community message classification.

[0039] S120. Based on the comparison of the message categories before and after correction among the messages to be corrected, select the target message to be corrected from at least one message to be corrected.

[0040] For example, at least one candidate message data group can be determined based on a comparison of the message categories before and after correction among the messages to be corrected; a target message data group can be determined based on the number of messages to be corrected in each candidate message data group, and the messages to be corrected in the target message data group can be identified as target correction messages.

[0041] In an optional embodiment, selecting a target message to be corrected from at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected includes: grouping messages to be corrected that have the same message categories before and after correction among the messages to be corrected to obtain at least one candidate message data group; determining the candidate message data group that meets the second preset message volume threshold condition as the target message data group; and determining the message to be corrected in the target message data group as the target message to be corrected.

[0042] For example, messages that share the same pre-correction message category and the same post-correction message category are grouped together and assigned to the same candidate message data group. For instance, if there are messages A, B, C, and D to be corrected, where messages A and B have a pre-correction message category of 'a' and a post-correction message category of 'b'; and messages C and D have a pre-correction message category of 'a' and a post-correction message category of 'c', then messages A and B are assigned to the same candidate message data group, and messages C and D are assigned to the same candidate message data group.

[0043] The second preset message volume threshold can be preset by relevant technical personnel. For example, the second preset message volume threshold can be that the number of messages to be corrected in the candidate message data group is not less than a preset second message volume threshold. The second message volume threshold can be preset by relevant technical personnel; for example, the second message volume threshold can be 100 messages.

[0044] For example, if the number of messages to be corrected in each candidate message data group is not less than a preset second message number threshold, then the candidate message data group is determined as the target message data group, and the messages to be corrected in the target message data group are determined as target correction messages.

[0045] S130. Generate at least one candidate similarity category group; wherein the message similarity of the target correction messages in the same candidate similarity category group is adjacent.

[0046] For example, at least one candidate similarity category group can be generated based on the message similarity between each target correction message. Specifically, the message similarity between each target correction message can be determined based on a pre-defined similarity determination algorithm, and based on the determination result, target correction messages with adjacent message similarities can be assigned to the same candidate similarity category group, thereby obtaining at least one candidate similarity category group. The similarity determination algorithm can be pre-defined by relevant technical personnel; for example, the similarity determination algorithm can be an Euclidean distance algorithm or a K-means algorithm, etc.

[0047] S140. Update the historical message dataset based on the consistency between the target correction message in the candidate similar category group and the historical message in the historical message dataset.

[0048] It should be noted that in the historical message dataset used to train the community message classification model, there may be historical messages that are identical to the target correction message. For example, if the target correction message is a greeting message like "Hello, Happy Holidays!", then there may be historical messages like "Hello, Happy Holidays!" in the historical message dataset. Optionally, there may not be any historical messages identical to the target correction message in the historical message dataset. Or, there may be target correction messages that are identical to historical messages in the historical message dataset, and there may also be target correction messages that are different from historical messages in the historical message dataset.

[0049] In an optional embodiment, updating the historical message dataset based on the consistency between the target correction message in the candidate similar category group and the historical message in the historical message dataset includes: if there is a target correction message in the candidate similar category group that is the same as a historical message in the historical message dataset, then the message category of the corresponding historical message in the historical message dataset is updated using the message category of the same target correction message; if there is a target correction message in the candidate similar category group that is different from a historical message in the historical message dataset, then the different target correction message and its corresponding message category are added to the historical message dataset as historical message data.

[0050] For example, if there is a target correction message in the candidate similarity group that is the same as a historical message in the historical message dataset, it can be determined whether the message category of the corresponding historical message is the same as the corrected message category of the corresponding target correction message. If so, there is no need to correct the corresponding historical message; if not, the message category of the corresponding historical message in the historical message dataset is updated using the message category of the same target correction message.

[0051] Specifically, if there is a candidate similarity where the target corrected message A is the same as the historical message C in the historical message dataset, and the corrected message category of the target corrected message A is message category a, then if the message category of the historical message C is message category a, then there is no need to correct the message category of the historical message C; if the message category of the historical message C is not message category a, then the message category of the historical message C is updated to message category a.

[0052] For example, if there is a target correction message in the candidate similarity group that is different from the historical message in the historical message dataset, it can be determined that the corresponding target correction message does not exist in the historical message dataset. Therefore, different target correction messages and their corresponding corrected message types can be added to the historical message dataset to improve the historical messages in the historical message dataset, increase the number of sample training sets, and thus facilitate the improvement of the training accuracy of the community message classification model.

[0053] This optional embodiment improves the number of historical messages in the historical sample dataset and enables batch correction of the message types of historical messages, thereby improving the accuracy of the community message classification model and facilitating more accurate message classification in the future.

[0054] S150. Train the community message classification model using the updated historical message dataset to update the community message classification model for message classification.

[0055] For example, by updating the message types of historical messages and adding new historical messages to the historical message dataset, the community message classification model is retrained, thereby continuously updating and improving the community message classification model. The updated and improved community message classification model is then used to classify the community messages acquired in real time.

[0056] This invention provides a solution that obtains at least one message to be corrected, selects a target message to be corrected from the at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected, generates at least one candidate similar category group, updates the historical message dataset based on the consistency between the target message to be corrected in the candidate similar category group and the historical messages in the historical message dataset, and trains a community message classification model using the updated historical message dataset to update the community message classification model for message classification. This achieves automated maintenance of the classification model used for community message classification, reducing manual input costs. By updating the community message classification model based on the consistency between the target message to be corrected in the candidate similar category group and the historical messages in the historical message dataset, the community message classification model is continuously updated and improved, thereby improving the classification accuracy of the community message classification model and thus improving the classification accuracy of community messages.

[0057] Example 2

[0058] Figure 2 This is a flowchart of a social media message classification method provided in Embodiment 2 of the present invention. This embodiment is an optimization and improvement based on the above technical solutions.

[0059] Furthermore, after the step "generating at least one candidate similar category group", the following step is added: "determine the candidate similar category group that meets the first preset message volume threshold condition as the target similar category group; determine at least one historical message to be corrected in the historical message dataset that has the same message category as the target correction message in the target similar category group; determine the target historical message based on the similarity between the target correction message in the target similar category group and each historical message to be corrected; update the message category of the target historical message using the message category of the target correction message to update the historical message dataset." This improves the method for updating the historical message dataset.

[0060] See Figure 2 The community message classification methods shown include:

[0061] S210. Obtain at least one message to be corrected; the message to be corrected includes the message category before correction and the message category after correction; the message category before correction is obtained based on the output of the community message classification model; the community message classification model is trained based on historical messages in the historical message dataset.

[0062] S220. Based on the comparison of the message categories before and after correction among the messages to be corrected, select the target message to be corrected from at least one message to be corrected.

[0063] S230. Generate at least one candidate similarity category group; wherein the message similarity of the target correction messages in the same candidate similarity category group is adjacent.

[0064] S240. The candidate similar category group that meets the first preset message volume threshold condition is determined as the target similar category group.

[0065] The first preset message volume threshold can be preset by relevant technical personnel. For example, the first preset message volume threshold can be that the number of target correction messages in the candidate message data group is not less than a preset first message volume threshold. The first message volume threshold can be preset by relevant technical personnel; for example, the first message volume threshold can be 100 messages.

[0066] For example, the candidate similar category group in each candidate similar category group whose number of target correction messages is not less than the first message number threshold is determined as the target similar category group.

[0067] S250. Identify at least one historical message in the historical message dataset that has the same message category as the target correction message in the target similarity category group before correction.

[0068] For example, if the historical message dataset contains historical messages of message category A, message category B, and message category C respectively, and if the original message category of the target message to be corrected is message category B, then at least one historical message corresponding to message category B in the historical message dataset will be identified as the historical message to be corrected.

[0069] S260. Determine the target historical messages based on the similarity between the target correction messages and each historical message to be corrected in the target similarity category group.

[0070] For example, the message similarity between the target correction message and each historical message to be corrected is determined, and the historical messages to be corrected that are similar to the target correction message are identified as target historical messages. It should be noted that, to further improve the accuracy of identifying target historical messages, the target historical messages can be determined by calculating message vector values.

[0071] In an optional embodiment, determining the target historical message based on the similarity between the target correction message in the target similarity category group and each historical message to be corrected includes: determining the message vector value corresponding to each target correction message in the target similarity category group; determining the target average vector value corresponding to the target similarity category group based on each message vector value; determining the historical vector value corresponding to each historical message to be corrected; and determining the target historical message based on the target average vector value and each historical vector value.

[0072] For example, the message vector values ​​corresponding to the target correction messages can be determined separately, and the average value of each message vector value can be determined as the target average vector value of the target similarity category group. The historical vector values ​​corresponding to each historical message to be corrected can be determined separately. Using the target average vector value and the historical vector values ​​of each historical message to be corrected, the similarity of the historical messages to be corrected can be calculated sequentially, and the historical messages to be corrected that meet the preset similarity threshold can be determined as the target historical messages.

[0073] S270. Using the corrected message category of the target corrected message, update the message category of the target historical message to update the historical message dataset.

[0074] For example, the message category of the target historical message is updated to the corrected message category of the target corrected message to update and improve the message category of historical messages in the historical message dataset.

[0075] S280. Update the historical message dataset based on the consistency between the target correction message in the candidate similar category group and the historical message in the historical message dataset.

[0076] S290. Train the community message classification model using the updated historical message dataset to update the community message classification model for message classification.

[0077] This embodiment's technical solution determines the candidate similarity category groups that meet the first preset message volume threshold condition from each candidate similarity category group as the target similarity category group; it determines at least one historical message in the historical message dataset that has the same message category as the target corrected message in the target similarity category group before correction; it determines the target historical message based on the similarity between the target corrected message in the target similarity category group and each historical message to be corrected; and it updates the message category of the target historical message using the corrected message category of the target corrected message to update the historical message dataset. This solution, by using the determined message category of the historical message to be corrected to correct the message category of the target historical message, achieves automated updating of the historical message dataset, improves the accuracy of historical messages in the historical message dataset, thereby improving the classification accuracy of the community message classification model, and ultimately improving the classification accuracy of community messages.

[0078] Example 3

[0079] Figure 3 This is a flowchart of a social media message classification method provided in Embodiment 3 of the present invention. This embodiment is an optimization and improvement based on the above technical solutions.

[0080] Furthermore, the step "generating at least one candidate similar category group" is refined to "vectorizing each target correction message to obtain the target vector value corresponding to each target correction message; determining the similarity between each target correction message based on the target vector value of each target correction message; and generating at least one candidate similar category group based on the similarity between each target correction message." This improves the method for generating candidate similar category groups.

[0081] See Figure 3 The community message classification methods shown include:

[0082] S310. Obtain at least one message to be corrected; the message to be corrected includes the message category before correction and the message category after correction; the message category before correction is obtained based on the output of the community message classification model; the community message classification model is trained based on historical messages in the historical message dataset.

[0083] S320. Based on the comparison of the message categories before and after correction among the messages to be corrected, select the target message to be corrected from at least one message to be corrected.

[0084] S330. Perform vectorization processing on each target correction message to obtain the target vector value corresponding to each target correction message.

[0085] It should be noted that this embodiment does not limit the vectorization processing method for the target correction message. It can be based on any vectorization processing method to obtain the target vector value corresponding to each target correction message.

[0086] S340. Based on the target vector values ​​of each target correction message, determine the similarity between each target correction message, and generate at least one candidate similarity category group based on the similarity between each target correction message; wherein the message similarity of target correction messages in the same candidate similarity category group is adjacent.

[0087] For example, the similarity between target correction messages can be determined based on their target vector values, and at least two target correction messages with adjacent similarities can be grouped into a candidate similarity category group. It should be noted that, to further improve the accuracy of determining candidate similarity category groups, the method of determining whether the similarity satisfies the similarity condition can also be used to determine the candidate similarity category group.

[0088] In an optional embodiment, the similarity between each target correction message is determined based on the target vector value of each target correction message, and at least one candidate similarity category group is generated based on the similarity between each target correction message, including: selecting any two target correction messages from each target correction message as a first message and a second message; determining a first target similarity between the first message and the second message based on the first vector value of the first message and the second vector value of the second message; if the first target similarity meets a preset similarity threshold condition, then the first message and the second message are assigned to the same candidate similarity category group, and a first average vector value of the first vector value and the second vector value is determined; other target correction messages are traversed sequentially, and a second target similarity between the first message, the second message, and the third message is determined based on the first average vector value and the third vector value of the traversed third message; wherein, other target messages include other target correction messages besides the first message and the second message; if the second similarity meets a preset similarity threshold condition, then the third message is added to the same candidate similarity category group as the first message or the second message, until the traversal is completed, resulting in at least one candidate similarity category group with similar message categories.

[0089] It should be noted that, to improve the accuracy of identifying candidate similarity category groups, a pairwise similarity determination method is used. For example, any two target correction messages are selected as the first message and the second message. Based on the first vector value corresponding to the first message and the second vector value corresponding to the second message, a first similarity is determined between the first message and the second message. If the first target similarity meets a preset similarity threshold, the first message and the second message are assigned to the same candidate similarity category group. The preset similarity threshold can be pre-set by relevant technical personnel. For example, the preset similarity threshold could be that the first target similarity is not less than a preset similarity threshold. The preset similarity threshold can also be pre-set by relevant technical personnel. For example, the preset similarity threshold could be 0.8. Specifically, if the first target similarity is not less than the preset similarity threshold, the first message and the second message are divided into the same candidate similarity category group, and the first average vector value of the first vector value and the second vector value is determined; if the first target similarity is less than the preset similarity threshold, it is indicated that the first message and the second message do not meet the preset similarity threshold condition, and the first message and the second message are further compared with other target correction messages.

[0090] If the first target similarity meets the preset similarity threshold, the first message and the second message are grouped into the same candidate similarity category group, and the first average vector value of the first vector value and the second vector value is determined. Other target correction messages besides the first and second messages are sequentially traversed, and the second target similarity between the first message, the second message, and the third message is determined based on the first average vector value and the third vector value of the traversed third message. If the second similarity meets the preset similarity threshold, the third message is added to the same candidate similarity category group as the first message or the second message, and the process continues to traverse other target correction messages besides the first message, the second message, and the third message, performing similarity calculations sequentially until the traversal is complete, resulting in at least one candidate similarity category group with similar message categories.

[0091] S350. Update the historical message dataset based on the consistency between the target correction message in the candidate similar category group and the historical message in the historical message dataset.

[0092] S360. The updated historical message dataset is used to train the community message classification model to update the community message classification model for message classification.

[0093] This embodiment of the scheme vectorizes each target correction message to obtain the target vector value corresponding to each target correction message; based on the target vector value of each target correction message, the similarity between each target correction message is determined, and based on the similarity between each target correction message, at least one candidate similar category group is generated. This achieves accurate determination of the candidate similar category group, improves the accuracy of subsequent updates to the historical message dataset, thereby improving the training accuracy of the subsequent community message classification model, and ultimately improving the classification accuracy of community messages.

[0094] Example 4

[0095] Figure 4 This is a schematic diagram of a social media message classification device provided in Embodiment 4 of the present invention. The social media message classification device provided in this embodiment of the present invention is applicable to the automated classification of social media messages. This social media message classification device can be implemented in hardware and / or software, such as... Figure 4 As shown, the device specifically includes: a message acquisition module 401 to be corrected, a target correction message selection module 402, a category group generation module 403, a dataset update module 404, and a message classification module 405. Among them,

[0096] The message to be corrected acquisition module 401 is used to acquire at least one message to be corrected; the message to be corrected includes a message category before correction and a message category after correction; the message category before correction is obtained based on the output of a community message classification model; the community message classification model is trained based on historical messages in a historical message dataset;

[0097] The target correction message selection module 402 is used to select a target correction message from at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected.

[0098] The category group generation module 403 is used to generate at least one candidate similar category group; wherein the message similarity of the target correction messages in the same candidate similar category group is adjacent.

[0099] The dataset update module 404 is used to update the historical message dataset based on the consistency between the target correction message in the candidate similar category group and the historical message in the historical message dataset;

[0100] The message classification module 405 is used to train the community message classification model with the updated historical message dataset to update the community message classification model for message classification.

[0101] This invention provides a solution that obtains at least one message to be corrected, selects a target message to be corrected from the at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected, generates at least one candidate similar category group, updates the historical message dataset based on the consistency between the target message to be corrected in the candidate similar category group and the historical messages in the historical message dataset, and trains a community message classification model using the updated historical message dataset to update the community message classification model for message classification. This achieves automated maintenance of the classification model used for community message classification, reducing manual input costs. By updating the community message classification model based on the consistency between the target message to be corrected in the candidate similar category group and the historical messages in the historical message dataset, the community message classification model is continuously updated and improved, thereby improving the classification accuracy of the community message classification model and thus improving the classification accuracy of community messages.

[0102] Optionally, the dataset update module 404 includes:

[0103] The message category update unit is used to update the message category of the corresponding historical message in the historical message dataset by adopting the message category of the same target correction message if there is a target correction message in the candidate similar category group that is the same as the historical message in the historical message dataset.

[0104] The message adding unit is used to add the different target correction messages and their corresponding message categories as historical message data to the historical message dataset if there are target correction messages in the candidate similarity category group that are different from the historical messages in the historical message dataset.

[0105] Optionally, the device further includes:

[0106] The similar category group determination module is used to determine the candidate similar category group that meets the first preset message volume threshold condition as the target similar category group after generating at least one candidate similar category group;

[0107] The module for determining historical messages to be corrected is used to determine at least one historical message to be corrected in the historical message dataset that has the same message category as the original message category of the target correction message in the target similarity category group.

[0108] The target historical message determination module is used to determine the target historical message based on the similarity between the target correction message in the target similarity category group and each of the historical messages to be corrected;

[0109] The historical message dataset update module is used to update the message category of the target historical message by adopting the corrected message category of the target corrected message, so as to update the historical message dataset.

[0110] Optionally, the target historical message determination module includes:

[0111] A message vector value determination unit is used to determine the message vector value corresponding to each target correction message in the target similarity category group;

[0112] The target average vector value determination unit is used to determine the target average vector value corresponding to the target similarity category group based on each of the message vector values;

[0113] A historical vector value determination unit is used to determine the historical vector value corresponding to each of the historical messages to be corrected.

[0114] The target historical message data determination unit is used to determine the target historical message data based on the target average vector value and each of the historical vector values.

[0115] Optionally, the target correction message selection module 402 includes:

[0116] The candidate message data group determination unit is used to combine messages to be corrected that have the same message category before correction and message category after correction among the messages to be corrected, to obtain at least one candidate message data group.

[0117] The target message data group determination unit is used to determine the candidate message data group that meets the second preset message volume threshold condition as the target message data group;

[0118] The target correction message determination unit is used to determine the message to be corrected in the target message data group as the target correction message.

[0119] Optionally, the category group generation module 403 includes:

[0120] The target vector value determination unit is used to perform vectorization processing on each of the target correction messages to obtain the target vector value corresponding to each of the target correction messages;

[0121] The category group determination unit is used to determine the similarity between each target correction message based on the target vector value of each target correction message, and to generate at least one candidate similar category group based on the similarity between each target correction message.

[0122] Optionally, the category group determination unit includes:

[0123] The message selection subunit is used to select any two target correction messages from each of the target correction messages as the first message and the second message.

[0124] The first target similarity determination subunit is used to determine the first target similarity between the first message and the second message based on the first vector value of the first message and the second vector value of the second message.

[0125] The first average vector value determination subunit is used to classify the first message and the second message into the same candidate similarity category group if the first target similarity meets the preset similarity threshold condition, and to determine the first average vector value of the first vector value and the second vector value.

[0126] The second target similarity determination subunit is used to sequentially traverse other target correction messages and determine the second target similarity between the first message, the second message, and the third message based on the first average vector value and the third vector value of the traversed third message data; wherein, the other target messages include other target correction messages besides the first message and the second message;

[0127] The candidate similarity category group determination subunit is used to add the third message to the same candidate similarity category group as the first message or the second message if the second similarity meets the preset similarity threshold condition, until the traversal is completed and at least one candidate similarity category group with similar message categories is obtained.

[0128] The social media message classification device provided in this embodiment of the invention can execute the social media message classification method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0129] Example 5

[0130] Figure 5 A schematic diagram of an electronic device 50 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0131] like Figure 5As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52 and a random access memory (RAM) 53, communicatively connected to the at least one processor 51. The memory stores computer programs executable by the at least one processor. The processor 51 can perform various appropriate actions and processes based on the computer program stored in the ROM 52 or loaded into the RAM 53 from storage unit 58. The RAM 53 can also store various programs and data required for the operation of the electronic device 50. The processor 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0132] Multiple components in electronic device 50 are connected to I / O interface 55, including: input unit 56, such as keyboard, mouse, etc.; output unit 57, such as various types of monitors, speakers, etc.; storage unit 58, such as disk, optical disk, etc.; and communication unit 59, such as network card, modem, wireless transceiver, etc. Communication unit 59 allows electronic device 50 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0133] Processor 51 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 51 performs the various methods and processes described above, such as social media message classification methods.

[0134] In some embodiments, the social media message classification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 50 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the social media message classification method described above may be performed. Alternatively, in other embodiments, processor 51 may be configured to perform the social media message classification method by any other suitable means (e.g., by means of firmware).

[0135] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0136] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0137] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0138] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0139] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0140] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0141] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0142] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for classifying social media messages, characterized in that, include: Obtain at least one message that needs to be corrected; The messages to be corrected include the message category before correction and the message category after correction; The original message category is obtained based on the output of the community message classification model; the community message classification model is trained based on historical messages in the historical message dataset. Based on the comparison of the message categories before and after correction among the messages to be corrected, a target message to be corrected is selected from at least one message to be corrected; Generate at least one candidate similarity category group; wherein the message similarity of the target correction messages in the same candidate similarity category group is adjacent. The historical message dataset is updated based on the consistency between the target correction message in the candidate similarity category group and the historical messages in the historical message dataset. The updated historical message dataset is used to train the community message classification model to update the community message classification model for message classification. The step of selecting a target correction message from at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected includes: The messages to be corrected that have the same pre-correction message category and post-correction message category are combined to obtain at least one candidate message data group. Candidate message data groups that meet the second preset message volume threshold condition are identified as target message data groups; The message to be corrected in the target message data group is identified as the target correction message; The step of updating the historical message dataset based on the consistency between the target correction messages in the candidate similarity category group and the historical messages in the historical message dataset includes: If there is a target correction message in the candidate similarity category group that is the same as a historical message in the historical message dataset, then the message category of the corresponding historical message in the historical message dataset is updated using the message category of the same target correction message. If there is a target correction message in the candidate similarity category group that is different from the historical messages in the historical message dataset, then the different target correction messages and their corresponding message categories are added to the historical message dataset as historical message data.

2. The method according to claim 1, characterized in that, After generating at least one candidate similar category group, the method further includes: The candidate similarity group that meets the first preset message volume threshold condition is determined as the target similarity group; Identify at least one historical message in the historical message dataset that has the same message category as the target correction message in the target similarity category group before correction; The target historical message is determined based on the similarity between the target correction message in the target similarity category group and each of the historical messages to be corrected; The message category of the target historical message is updated using the corrected message category of the target corrected message, thereby updating the historical message dataset.

3. The method according to claim 2, characterized in that, The step of determining the target historical message based on the similarity between the target correction message in the target similarity category group and each of the historical messages to be corrected includes: Determine the message vector value corresponding to each target correction message in the target similarity category group; Based on each of the message vector values, determine the target average vector value corresponding to the target similarity category group; Determine the historical vector value corresponding to each of the aforementioned historical messages to be corrected; The target historical messages are determined based on the target average vector value and each of the historical vector values.

4. The method according to claim 1, characterized in that, The generation of at least one candidate similar category group includes: The target correction messages are vectorized to obtain the target vector values ​​corresponding to each target correction message. Based on the target vector values ​​of each target correction message, the similarity between each target correction message is determined, and at least one candidate similarity category group is generated based on the similarity between each target correction message.

5. The method according to claim 4, characterized in that, The step of determining the similarity between each target correction message based on the target vector value of each target correction message, and generating at least one candidate similarity category group based on the similarity between each target correction message, includes: Select any two target correction messages from the aforementioned target correction messages as the first message and the second message; Based on the first vector value of the first message and the second vector value of the second message, a first target similarity between the first message and the second message is determined; If the first target similarity meets the preset similarity threshold condition, then the first message and the second message are divided into the same candidate similarity category group, and the first average vector value of the first vector value and the second vector value is determined. The other target correction messages are traversed sequentially, and the second target similarity between the first message, the second message, and the third message is determined based on the first average vector value and the third vector value of the traversed third message; wherein, the other target correction messages include other target correction messages besides the first message and the second message; If the second similarity meets the preset similarity threshold condition, the third message is added to the same candidate similarity category group as the first message or the second message, until the traversal is completed and at least one candidate similarity category group with similar message categories is obtained.

6. A social media message classification device, characterized in that, include: The message to be corrected acquisition module is used to acquire at least one message to be corrected; The messages to be corrected include the message category before correction and the message category after correction; The original message category is obtained based on the output of the community message classification model; the community message classification model is trained based on historical messages in the historical message dataset. The target correction message selection module is used to select a target correction message from at least one message to be corrected based on a comparison of the message categories before and after correction among the messages to be corrected. A category group generation module is used to generate at least one candidate similar category group; wherein the message similarity of the target correction messages in the same candidate similar category group is adjacent. The dataset update module is used to update the historical message dataset based on the consistency between the target correction messages in the candidate similar category group and the historical messages in the historical message dataset; The message classification module is used to train the community message classification model with the updated historical message dataset to update the community message classification model for message classification. The target correction message selection module includes: The candidate message data group determination unit is used to combine messages to be corrected that have the same message category before correction and message category after correction among the messages to be corrected, to obtain at least one candidate message data group. The target message data group determination unit is used to determine the candidate message data group that meets the second preset message volume threshold condition as the target message data group; A target correction message determination unit is used to determine the message to be corrected in the target message data group as a target correction message; The dataset update module includes: The message category update unit is used to update the message category of the corresponding historical message in the historical message dataset by adopting the message category of the same target correction message if there is a target correction message in the candidate similar category group that is the same as the historical message in the historical message dataset. The message adding unit is used to add the different target correction messages and their corresponding message categories as historical message data to the historical message dataset if there are target correction messages in the candidate similarity category group that are different from the historical messages in the historical message dataset.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the social message classification method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the social message classification method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Junk information identification method and equipment

    CN107515873A

  • Method and device for generating dialogue information, equipment and medium

    CN114625855A