Data Acquisition and Processing System for User Text Sentiment Tendency Analysis

By designing a data acquisition and processing system for sentiment analysis, the problem of data verification not being considered in the prior art is solved, and the reliability and accuracy of sentiment analysis data is improved.

CN119721007BActive Publication Date: 2025-06-27ZHONGHE YUNKE INFORMATION TECH GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510235398.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-27
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The prior art directly conducts unified sentiment analysis on the simplest text processed in sentiment analysis, without considering data verification, resulting in a decrease in data reliability and a decrease in the accuracy of sentiment tendency analysis.

Method used

Design a data acquisition and processing system, including a collection module, a cache module, a storage module, a trigger module, a sensitive information screening module, an interference information screening module and a verification module, and ensure the reliability of data processing by periodically obtaining session content, screening sensitive and interference information, and verifying data volume.

Benefits of technology

By screening and verifying the data, the reliability of sentiment analysis data is improved, thereby improving the accuracy of sentiment tendency analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721007B_ABST
    Figure CN119721007B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and particularly relates to a data acquisition and processing system for user text sentiment tendency analysis, including an acquisition module, a cache module, a storage module, a trigger module, a sensitive information screening module, an interference information screening module, and a verification module. It screens the data used for sentiment analysis, screens out interference information and sensitive information, and verifies the content to be verified after screening to determine whether the processing of the conversation content is qualified. When it is determined that the processing of the conversation content is abnormal in a timely manner, it determines the processing method for the information screening module, improving the reliability of the data used for sentiment analysis and the accuracy of sentiment tendency analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a data acquisition and processing system for analyzing the emotional tendency of user texts. Background Art

[0002] The rapid development of artificial intelligence (AI) technology has made natural language processing (NLP) one of the important fields of intelligent applications. As a part of NLP, emotion recognition aims to automatically detect and analyze the emotional state of users from texts, which is of great significance for improving the user experience and optimizing service processes.

[0003] In the prior art, emotion recognition usually relies on rule-based classification methods or traditional machine learning models, and existing emotion recognition solutions mainly focus on the output of emotion classification results.

[0004] Chinese Patent Publication No.: CN109840328B discloses a deep learning method for analyzing the emotional tendency of commodity review texts. The text is processed to obtain the simplest text, and the simplest text is divided into a training set and a test set. Word sequences are obtained according to the training set, and emotional features with emotional weight values are set, and the emotional weight values are adjusted to obtain corrected emotional weight values. Multiple phrase sequences are obtained according to different word combination methods of the word sequences, and the corrected emotional weight values in the phrase sequences are calculated to obtain the emotional labels of the text, and the emotional labels are compared with the artificial emotional evaluations to obtain the final emotional weight values and the final corrected emotional weight values. The above operations are performed on all texts in the training set, and the emotional model is obtained according to the final emotional weight values and the final corrected emotional weight values to perform emotional analysis on the test set for verification. It can be seen that the above technical solutions have the following problems: directly performing unified emotional analysis on the processed simplest text without considering the verification of the data used for emotional analysis, which affects the reliability of the data used for emotional analysis, and further leads to the problem of reduced accuracy of emotional tendency analysis. Summary of the Invention

[0005] Therefore, the present invention provides a data acquisition and processing system for analyzing the emotional tendency of user texts to overcome the problems in the prior art that directly perform unified emotional analysis on the processed simplest text without considering the verification of the data used for emotional analysis, which affects the reliability of the data used for emotional analysis, and further leads to the problem of reduced accuracy of emotional tendency analysis.

[0006] To achieve the above object, the present invention provides a data acquisition and processing system for analyzing the emotional tendency of user texts, including:

[0007] An acquisition module for collecting the conversation content of users;

[0008] A cache module, which is connected to the acquisition module and used to cache the session content of the user;

[0009] A storage module, which is connected to the cache module and used to obtain the session content cached by the cache module when the stored data volume of the cache module reaches the critical storage value, and send a content clearing instruction to the cache module when the caching of the session content is completed;

[0010] A trigger module, which is connected to the cache module. The trigger module periodically issues a trigger instruction at a preset trigger duration, and obtains all the session content in the cache module when the trigger instruction is issued;

[0011] A sensitive information screening module, which is connected to the trigger module and used to screen out sensitive information in the session content obtained by the trigger module;

[0012] An interference information screening module, which is connected to the sensitive information screening module and used to screen out interference information from the session information after the sensitive information is screened out, and record the session content after the interference information is screened out as the content to be verified;

[0013] A verification module, which is respectively connected to the interference information screening module, the sensitive information screening module and the trigger module, and used to determine whether the processing of the session content is qualified based on the data volume of the content to be verified and the data volume of the session content obtained by the trigger module, including:

[0014] Determine that the processing of the session content is qualified, output the content to be verified, and perform text sentiment tendency analysis on the content to be verified;

[0015] Determine that the processing of the session content is abnormal, and determine the processing method for the information screening module based on the data volume of the screened sensitive information and the data volume of the screened interference information.

[0016] Further, the sensitive information screening module is used to screen out sensitive information in the session content obtained by the trigger module, including:

[0017] Select a session window containing a preset number of sessions, perform natural sentence analysis on the sessions in the session window, and screen out the sensitive information marked by the analysis;

[0018] The interference information screening module is used to screen out interference information from the session information after the sensitive information is screened out, including:

[0019] Determine the session information after the sensitive information is screened out as desensitized session information;

[0020] Divide the desensitized session information into several keywords, compare each keyword with each screening word in the screening standard library, determine the keywords with a similarity greater than the preset similarity as interference information, and screen them out;

[0021] The similarity is the ratio of the number of characters in the longest consecutive identical part between the keyword and the screening word to the total number of characters in the screening word.

[0022] Further, the verification module is used to determine the screening ratio based on the data volume of the content to be verified and the data volume of the session content obtained by the trigger module, and determine whether the processing of the session content is qualified based on the screening ratio, including:

[0023] Used to determine the screening ratio, calculate the difference between the data volume of the session content obtained by the trigger module and the data volume of the content to be verified, and solve the ratio of this difference to the data volume of the session content obtained by the trigger module to obtain the screening ratio;

[0024] If the screening ratio is less than or equal to the first preset screening ratio, it is determined that the processing of the session content is qualified;

[0025] If the screening ratio is less than or equal to the second preset screening ratio and greater than the first preset screening ratio, it is determined whether the processing of the session content is qualified in combination with the data volume of the session content obtained by the trigger module;

[0026] If the screening ratio is greater than the second preset screening ratio, it is determined that the processing of the session content is abnormal, and the processing method for the information screening module is determined based on the data volume of the sensitive information screened out and the data volume of the interference information screened out.

[0027] Further, the verification module is used to determine whether the processing of the session content is qualified based on the data volume of the session content obtained by the trigger module, including:

[0028] Record the data volume of the session content obtained by the trigger module as the trigger data volume;

[0029] If the trigger data volume is less than or equal to the preset data volume, it is determined that the processing of the session content is qualified, and the preset trigger duration is adjusted to the corresponding value based on the trigger data volume and the historical data increment;

[0030] If the trigger data volume is greater than the preset data volume, it is determined that the processing of the session content is abnormal, and the processing method for the information screening module is determined based on the data volume of the sensitive information screened out and the data volume of the interference information screened out.

[0031] Further, the verification module is used to adjust the preset trigger duration to the corresponding value based on the trigger data volume, where:

[0032] The increase amplitude of the preset trigger duration determined based on the trigger data volume is inversely proportional to the trigger data volume.

[0033] Further, under the condition that the verification module completes the adjustment of the preset trigger duration based on the trigger data volume, the preset trigger duration is adjusted to the corresponding value based on the historical data increment, where:

[0034] The increase amplitude of the preset trigger duration determined based on the historical data increment is inversely proportional to the historical data increment;

[0035] The verification module is used to determine the historical data increment, draw a trigger data volume time domain curve based on each historical trigger data volume obtained by the trigger module, solve the slope of the trigger data volume time domain curve at the current time node, and determine the slope as the historical data increment.

[0036] Further, the verification module is used to determine the screening comparison parameter based on the data volume of the sensitive information screened out and the data volume of the interference information screened out, and determine the processing method for the information screening module based on the screening comparison parameter, including:

[0037] Record the ratio of the data volume of the interference information screened out to the data volume of the sensitive information screened out as the screening comparison parameter;

[0038] If the screening comparison parameter is less than or equal to the preset screening comparison parameter, the session window is adjusted to the corresponding value based on the screening comparison parameter;

[0039] If the screening comparison parameter is greater than the preset screening comparison parameter, calculate the interference coincidence parameter based on each screened interference information obtained, and determine the processing method for the interference information screening module based on the interference coincidence parameter.

[0040] Further, the verification module is used to determine the processing method for the interference information screening module based on the interference coincidence parameter, including:

[0041] Used to determine the interference coincidence parameter, obtain each screened interference information, divide each interference information into several data groups so that the same interference information belongs to the same data group, count the data volume of the interference information in each data group, and record the data volume of the data group with the largest data volume as the coincidence data volume; solve the ratio of the coincidence data volume to the total data volume of each screened interference information to obtain the interference coincidence parameter;

[0042] If the interference coincidence parameter is less than or equal to the preset interference coincidence parameter, the preset similarity is adjusted to the corresponding value based on the screening ratio;

[0043] If the interference coincidence parameter is greater than the preset interference coincidence parameter, the preset similarity is adjusted to the corresponding value based on the screening ratio, and the number of each screening word in the screening standard library is adjusted to the corresponding value based on the interference coincidence parameter.

[0044] Further, the verification module is used to adjust the preset similarity to the corresponding value based on the screening ratio, where:

[0045] The increase amplitude of the preset similarity determined based on the screening ratio is proportional to the screening ratio;

[0046] The verification module is used to adjust the number of each screening word in the screening standard library to the corresponding value based on the interference coincidence parameter, where:

[0047] The increase amplitude of the number of each screening word in the screening standard library determined based on the interference coincidence parameter is proportional to the interference coincidence parameter.

[0048] Further, the verification module is used to adjust the session window to the corresponding value based on the screening comparison parameter, where:

[0049] The increase amplitude of the session window determined based on the screening comparison parameter is inversely proportional to the screening comparison parameter.

[0050] Compared with the prior art, the beneficial effect of the present invention is that the data used for sentiment analysis is screened to screen out interference information and sensitive information, and the content to be verified after screening is verified to determine whether the processing of the session content is qualified. When it is determined that the processing of the session content is abnormal in time, the processing method for the information screening module is determined, which improves the reliability of the data used for sentiment analysis, and further improves the accuracy of sentiment tendency analysis.

[0051] Further, to determine whether the processing of the session content is qualified, when the screening ratio is greater than the second preset screening ratio, at this time, too much content is screened out, resulting in too low data volume for sentiment tendency analysis, and the accuracy of sentiment tendency analysis cannot be guaranteed. At this time, the processing method for the information screening module is determined to increase the data volume for sentiment tendency analysis; when the screening ratio is less than or equal to the second preset screening ratio and greater than the first preset screening ratio, at this time, it is determined whether the processing of the session content is qualified based on the data volume of the session content obtained by the comprehensive trigger module. When the trigger data volume is less than or equal to the preset data volume, at this time, due to too low original data, it cannot meet the data requirements for sentiment tendency analysis after a large amount of sensitive information and interference information are screened out. At this time, the preset trigger duration is increased to increase the data volume of the original data, further improving the acquisition efficiency of the data used for sentiment analysis.

[0052] Further, determine the processing method for the information screening module. When the screening comparison parameter is less than or equal to the preset screening comparison parameter, at this time, due to the excessive amount of sensitive information to be screened, the amount of data retained for sentiment analysis is relatively low. At this time, adjust the conversation window to improve the recognition accuracy of sensitive information; when the screening comparison parameter is greater than the preset screening comparison parameter, at this time, due to the low recognition accuracy of interference information, the amount of interference information to be screened is too large. At this time, obtain the interference coincidence parameter, which characterizes the coincidence of the recognized interference information. When the interference coincidence parameter is greater than the preset interference coincidence parameter, since the coincidence degree of the interference information to be screened is very high in this case, in order to ensure the accurate screening of interference information, while adjusting the preset similarity to improve the screening standard, increase the content of the screening standard library to ensure the screening accuracy. While ensuring the effective screening of interference information, streamline the content of the screening standard library to improve the screening efficiency of interference data, further improve the acquisition efficiency of data for sentiment analysis, thereby improving the efficiency of sentiment analysis and the accuracy of sentiment analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a block diagram of the data acquisition and processing system for user text sentiment analysis according to an embodiment of the present invention;

[0054] Figure 2 It is a logical decision diagram for the verification module according to an embodiment of the present invention to determine whether the processing of conversation content is qualified based on the screening ratio;

[0055] Figure 3 It is a logical decision diagram for the verification module according to an embodiment of the present invention to determine whether the processing of conversation content is qualified based on the data volume of the conversation content obtained by the trigger module;

[0056] Figure 4 It is a logical decision diagram for the verification module according to an embodiment of the present invention to determine the processing method for the information screening module based on the screening comparison parameter. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0059] It should be noted that in the description of the present invention, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the direction or positional relationship shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.

[0060] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0061] Please refer to Figure 1 、 Figure 2 、 Figure 3 and Figure 4 as shown, which are respectively the module block diagram of the data acquisition and processing system for user text sentiment tendency analysis in the embodiments of the present invention, the logical decision diagram of the verification module to determine whether the processing of the session content is qualified based on the screening ratio, the logical decision diagram of the verification module to determine whether the processing of the session content is qualified based on the data volume of the session content obtained by the trigger module, and the logical decision diagram of the verification module to determine the processing method of the information screening module based on the screening comparison parameter; A data acquisition and processing system for user text sentiment tendency analysis in an embodiment of the present invention includes:

[0062] An acquisition module for collecting the session content of users;

[0063] A cache module connected to the acquisition module for caching the session content of users;

[0064] A storage module connected to the cache module for obtaining the session content cached by the cache module when the stored data volume of the cache module reaches the critical storage value, and sending a content clearing instruction to the cache module when the caching of the session content is completed;

[0065] A trigger module connected to the cache module, which periodically issues a trigger instruction at a preset trigger duration, and obtains all the session content in the cache module when issuing the trigger instruction;

[0066] A sensitive information screening module connected to the trigger module for screening sensitive information in the session content obtained by the trigger module;

[0067] An interference information filtering module, which is connected to the sensitive information filtering module, is used to filter interference information from the session information after sensitive information filtering is completed, and record the session content after interference information filtering as the content to be verified;

[0068] A verification module, which is respectively connected to the interference information filtering module, the sensitive information filtering module and the triggering module, is used to determine whether the processing of the session content is qualified based on the data volume of the content to be verified and the data volume of the session content obtained by the triggering module, including:

[0069] Determine that the processing of the session content is qualified, and output the content to be verified for text sentiment tendency analysis;

[0070] Determine that the processing of the session content is abnormal, and determine the processing method for the information filtering module based on the data volume of the filtered sensitive information and the data volume of the filtered interference information.

[0071] Specifically, the verification module outputs the session content determined to be qualified for session content processing for text sentiment tendency analysis, which is the prior art and will not be elaborated here.

[0072] Specifically, the data for sentiment analysis is screened to filter out interference information and sensitive information, and the content to be verified after screening is verified to determine whether the processing of the session content is qualified. Even when it is determined that the processing of the session content is abnormal, determining the processing method for the information filtering module improves the reliability of the data for sentiment analysis, and thus improves the accuracy of sentiment tendency analysis.

[0073] Specifically, the sensitive information filtering module is used to filter sensitive information from the session content obtained by the triggering module, including:

[0074] Select a session window containing a preset number of sessions, perform natural language sentence analysis on the sessions within the session window, and filter out the sensitive information marked by the analysis;

[0075] The interference information filtering module is used to filter interference information from the session information after sensitive information filtering is completed, including:

[0076] Determine the session information after sensitive information filtering as desensitized session information;

[0077] Divide the desensitized session information into several keywords, compare each keyword with each screening word in the screening standard library, and determine the keywords with a similarity greater than the preset similarity as interference information and filter them out;

[0078] The similarity is the ratio of the number of characters in the longest consecutive identical part between the keyword and the screening word to the total number of characters of the screening word.

[0079] Specifically, sensitive information includes personal identity information, financial information, health information, trade secrets, passwords and access credentials, regulatory information for specific industries, personal privacy information, information related to national security, etc., which will not be elaborated here.

[0080] Specifically, there is no limitation on the specific method of performing natural language analysis on the conversations within the conversation window to obtain the sensitive information calibrated by the analysis. It can analyze the chat content within each conversation window in real time through the recognition model of AI NLP to calibrate the sensitive information, or import the conversations within the conversation window into the detection neural network, and the detection neural network outputs the content calibrated as sensitive information in the imported conversations, and screen out the content calibrated as sensitive information. This is the prior art and will not be elaborated here.

[0081] Specifically, the screening standard library is a pre-established set containing a series of screening words, which will not be elaborated here.

[0082] In order to make the objectives and advantages of the present invention more clearly understood, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0083] It should be noted that the data in this embodiment are comprehensively analyzed and evaluated from the historical detection data of the three months before this detection by the system described in the present invention and the corresponding historical detection results. The system described in the present invention determines the numerical values of various preset parameter standards for this detection based on whether the processing of the conversation content of 15,863 users detected cumulatively in the previous three months before this detection is qualified, the screening situation of sensitive information, the screening situation of interference information, and the specific analysis results of the processing of each text within each cycle. Those skilled in the art can understand that the determination method of the system described in the present invention for a single above-mentioned parameter can be to select the value with the highest proportion according to the data distribution as the preset standard parameter, use weighted summation to take the obtained value as the preset standard parameter, or other selection methods, as long as it satisfies that the system described in the present invention can clearly define different specific situations in the single determination process through the obtained values.

[0084] Specifically, the verification module is used to determine the screening ratio based on the data volume of the content to be verified and the data volume of the conversation content obtained by the trigger module, and determine whether the processing of the conversation content is qualified based on the screening ratio, including:

[0085] Used to determine the screening ratio, calculate the difference between the data volume of the conversation content obtained by the trigger module and the data volume of the content to be verified, and solve the ratio of this difference to the data volume of the conversation content obtained by the trigger module to obtain the screening ratio;

[0086] If the screening ratio is less than or equal to the first preset screening ratio, it is determined that the processing of the conversation content is qualified;

[0087] If the screening ratio is less than or equal to the second preset screening ratio and greater than the first preset screening ratio, it is determined whether the processing of the conversation content is qualified by combining the data volume of the conversation content obtained by the trigger module;

[0088] If the screening ratio is greater than the second preset screening ratio, it is determined that the processing of the conversation content is abnormal, and the processing method for the information screening module is determined based on the data volume of the sensitive information screened out and the data volume of the interference information screened out.

[0089] Specifically, the first preset screening ratio S1 is selected within the range [0.05, 0.15], and the second preset screening ratio S2 is selected within the range [0.2, 0.25];

[0090] Specifically, in this embodiment, it is preferably selected that 0.15 is the value of S1 and 0.25 is the value of S2. It is verified in 15863 experiments that when 0.15 is selected as the value of S1 and 0.25 is selected as the value of S2, the best division of the processing situation of the conversation content can be realized, and the division of the proportion of the screened content can be realized more clearly.

[0091] Specifically, the verification module is used to determine whether the processing of the conversation content is qualified based on the data volume of the conversation content obtained by the trigger module, including:

[0092] The data volume of the conversation content obtained by the trigger module is recorded as the trigger data volume;

[0093] If the trigger data volume is less than or equal to the preset data volume, it is determined that the processing of the conversation content is qualified, and the preset trigger duration is adjusted to the corresponding value based on the trigger data volume and the historical data increment;

[0094] If the trigger data volume is greater than the preset data volume, it is determined that the processing of the conversation content is abnormal, and the processing method for the information screening module is determined based on the data volume of the sensitive information screened out and the data volume of the interference information screened out.

[0095] Specifically, the preset data volume Y0 is selected within the range [200, 300], and the unit is MB;

[0096] Specifically, in this embodiment, it is preferably selected that 260 is the value of Y0. It is verified in 15863 experiments that when 260 is selected as the value of Y0, the division of the specific situation of the data volume of the conversation content obtained by the trigger module can be realized, and the situation of the obtained original data can be determined more clearly.

[0097] Specifically, to determine whether the processing of the conversation content is qualified, when the screening ratio is greater than the second preset screening ratio, too much content is screened out at this time, resulting in too low a data volume for sentiment analysis, and the accuracy of sentiment analysis cannot be guaranteed. At this time, determine the processing method for the information screening module to increase the data volume for sentiment analysis; when the screening ratio is less than or equal to the second preset screening ratio and greater than the first preset screening ratio, at this time, determine whether the processing of the conversation content is qualified based on the data volume of the conversation content obtained by the comprehensive trigger module. When the trigger data volume is less than or equal to the preset data volume, since the obtained original data is too low, it cannot meet the data requirements for sentiment analysis after a large amount of sensitive information and interference information are screened out. At this time, increase the preset trigger duration to increase the data volume of the original data, and further improve the acquisition efficiency of the data for sentiment analysis.

[0098] Specifically, the verification module is used to adjust the preset trigger duration to the corresponding value based on the trigger data volume, where:

[0099] The increase amplitude of the preset trigger duration determined based on the trigger data volume is inversely proportional to the trigger data volume.

[0100] In this embodiment, optionally,

[0101] Compare the trigger data volume with the first preset trigger comparison threshold and the second preset trigger comparison threshold;

[0102] If the trigger data volume is less than or equal to the first preset trigger comparison threshold, adjust the preset trigger duration to 1.21 times the current preset trigger duration;

[0103] If the trigger data volume is less than or equal to the second preset trigger comparison threshold and greater than the first preset trigger comparison threshold, adjust the preset trigger duration to 1.17 times the current preset trigger duration;

[0104] If the trigger data volume is greater than the second preset trigger comparison threshold, adjust the preset trigger duration to 1.11 times the current preset trigger duration;

[0105] The first preset trigger comparison threshold is taken as 0.5Y0, and the second preset trigger comparison threshold is taken as 0.7Y0.

[0106] Specifically, under the condition that the verification module completes the adjustment of the preset trigger duration based on the trigger data volume, the preset trigger duration is adjusted to the corresponding value based on the historical data increment, where:

[0107] The increase amplitude of the preset trigger duration determined based on the historical data increment is inversely proportional to the historical data increment;

[0108] The verification module is used to determine the historical data increment, draw the time-domain curve of the trigger data volume based on each historical trigger data volume obtained by the trigger module, solve the slope of the time-domain curve of the trigger data volume at the current time node, and determine the slope as the historical data increment.

[0109] In this embodiment, optionally,

[0110] Compare the historical data increment with the first preset increment and the second preset increment;

[0111] If the historical data increment is less than or equal to the first preset increment, adjust the preset trigger duration to 1.23 times the current preset trigger duration;

[0112] If the historical data increment is less than or equal to the second preset increment and greater than the first preset increment, adjust the preset trigger duration to 1.19 times the current preset trigger duration;

[0113] If the historical data increment is greater than the second preset increment, adjust the preset trigger duration to 1.12 times the current preset trigger duration;

[0114] The first preset increment is taken as -1.5, and the second preset increment is taken as -0.5.

[0115] Specifically, the verification module is used to determine the screening comparison parameter based on the data volume of the screened sensitive information and the data volume of the screened interference information, and determine the processing method for the information screening module based on the screening comparison parameter, including:

[0116] Record the ratio of the data volume of the screened interference information to the data volume of the screened sensitive information as the screening comparison parameter;

[0117] If the screening comparison parameter is less than or equal to the preset screening comparison parameter, adjust the session window to the corresponding value based on the screening comparison parameter;

[0118] If the screening comparison parameter is greater than the preset screening comparison parameter, calculate the interference coincidence parameter based on the obtained screened interference information, and determine the processing method for the interference information screening module based on the interference coincidence parameter.

[0119] Specifically, the preset screening comparison parameter L0 is selected within the range of [0.82, 0.93].

[0120] Specifically, the verification module is used to determine the processing method for the interference information screening module based on the interference coincidence parameter, including:

[0121] To determine the interference coincidence parameter, obtain each piece of interference information to be filtered out, divide each piece of interference information into several data groups so that the same interference information belongs to the same data group, count the data volume of the interference information in each data group, and record the data volume of the data group with the largest data volume as the coincidence data volume; solve the ratio of the coincidence data volume to the total data volume of each piece of interference information to be filtered out to obtain the interference coincidence parameter;

[0122] If the interference coincidence parameter is less than or equal to the preset interference coincidence parameter, then adjust the preset similarity to the corresponding value based on the filtering ratio;

[0123] If the interference coincidence parameter is greater than the preset interference coincidence parameter, then adjust the preset similarity to the corresponding value based on the filtering ratio, and adjust the number of each screening word in the screening standard library to the corresponding value based on the interference coincidence parameter.

[0124] Specifically, the preset interference coincidence parameter G0 is selected within the interval [0.29, 0.35].

[0125] Specifically, there is no limitation on the acquisition method of the newly added screening words in the screening standard library. Domain experts, business analysts or relevant staff can manually determine and add new screening words according to business changes and regulatory updates, or track new words and buzzwords on channels such as the Internet and social media in real time, and determine the text evaluated as interference information as screening words, which will not be elaborated here.

[0126] Specifically, determine the processing method for the information filtering module. When the filtering comparison parameter is less than or equal to the preset filtering comparison parameter, at this time, due to the excessive data volume of the sensitive information to be filtered out, the data volume for retention and sentiment analysis is relatively low. At this time, adjust the conversation window to improve the recognition accuracy of sensitive information; when the filtering comparison parameter is greater than the preset filtering comparison parameter, at this time, due to the low recognition accuracy of interference information, the data volume of the interference information to be filtered out is too large. At this time, obtain the interference coincidence parameter, which characterizes the coincidence of the recognized interference information. When the interference coincidence parameter is greater than the preset interference coincidence parameter, since the coincidence degree of the interference information to be filtered out is very high in this case, in order to ensure the accurate filtering of interference information, while adjusting the preset similarity to improve the screening standard, add content to the screening standard library to ensure the screening accuracy, while ensuring the effective filtering of interference information, streamline the content in the screening standard library to improve the filtering efficiency of interference data, and further improve the acquisition efficiency of the data for sentiment analysis, thereby improving the efficiency of sentiment analysis.

[0127] Specifically, the verification module is used to adjust the preset similarity to the corresponding value based on the filtering ratio, where:

[0128] The increase amplitude of the preset similarity determined based on the screening ratio is proportional to the screening ratio;

[0129] The verification module is used to adjust the number of each screening word in the screening standard library to the corresponding value based on the interference coincidence parameter, where:

[0130] The increase amplitude of the number of each screening word in the screening standard library determined based on the interference coincidence parameter is proportional to the interference coincidence parameter.

[0131] In this embodiment, optionally,

[0132] Compare the screening ratio with the first preset ratio comparison threshold and the second preset ratio comparison threshold;

[0133] If the screening ratio is less than or equal to the first preset ratio comparison threshold, adjust the preset similarity to 1.11 times the initial preset similarity;

[0134] If the screening ratio is less than or equal to the second preset ratio comparison threshold and greater than the first preset ratio comparison threshold, adjust the preset similarity to 1.18 times the initial preset similarity;

[0135] If the screening ratio is greater than the second preset ratio comparison threshold, adjust the preset similarity to 1.26 times the initial preset similarity;

[0136] The first preset ratio comparison threshold is taken as 1.7S2, and the second preset ratio comparison threshold is taken as 3S2.

[0137] In this embodiment, optionally,

[0138] Compare the interference coincidence parameter with the first preset interference comparison threshold and the second preset interference comparison threshold;

[0139] If the interference coincidence parameter is less than or equal to the first preset interference comparison threshold, adjust the number of each screening word in the screening standard library to 1.13 times the initial number;

[0140] If the interference coincidence parameter is less than or equal to the second preset interference comparison threshold and greater than the first preset interference comparison threshold, adjust the number of each screening word in the screening standard library to 1.23 times the initial number;

[0141] If the interference coincidence parameter is greater than the second preset interference comparison threshold, adjust the number of each screening word in the screening standard library to 1.33 times the initial number;

[0142] The first preset interference comparison threshold is taken as 0.8G0, and the second preset interference comparison threshold is taken as 2.1G0.

[0143] Specifically, the verification module is used to adjust the session window to the corresponding value based on the screening comparison parameter, where:

[0144] The increase amplitude of the session window determined based on the screening comparison parameter is inversely proportional to the screening comparison parameter.

[0145] In this embodiment, optionally,

[0146] Compare the screening comparison parameter with a first preset screening comparison threshold and a second preset screening comparison threshold;

[0147] If the screening comparison parameter is less than or equal to the first preset screening comparison threshold, adjust the number of preset sessions in the session window to 1.28 times the initial preset number;

[0148] If the screening comparison parameter is less than or equal to the second preset screening comparison threshold and greater than the first preset screening comparison threshold, adjust the number of preset sessions in the session window to 1.18 times the initial preset number;

[0149] If the screening comparison parameter is greater than the second preset screening comparison threshold, adjust the number of preset sessions in the session window to 1.11 times the initial preset number;

[0150] The first preset screening comparison threshold is taken as 0.63L0, and the second preset screening comparison threshold is taken as 0.84L0.

[0151] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or replacements to the relevant technical features, and the technical solutions after these changes or replacements will all fall within the protection scope of the present invention.

[0152] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data collection and processing system for user text sentiment analysis, which is used for analyzing and processing long texts in user text sentiment analysis, characterized in that: include: A collection module, which is used to collect user session content; A cache module, connected to the acquisition module, for caching the session content of the user; A storage module connected to the cache module, for obtaining the session content cached by the cache module when the amount of stored data in the cache module reaches a critical storage value, and for sending a content clearing instruction to the cache module when the caching of the session content is completed; A trigger module, which is connected to the cache module, periodically issues a trigger instruction with a preset trigger duration as a period, and obtains all session contents in the cache module when issuing the trigger instruction; A sensitive information screening module, which is connected to the trigger module and is used to screen out sensitive information in the conversation content obtained by the trigger module; An interference information screening module, which is connected to the sensitive information screening module, is used to screen out interference information from the session information after sensitive information screening is completed, and record the output session content after interference information screening is completed as content to be verified; A verification module, which is respectively connected to the interference information screening module, the sensitive information screening module and the trigger module, and is used to determine the screening ratio based on the data volume of the content to be verified and the data volume of the session content obtained by the trigger module, and determine whether the processing of the session content is qualified based on the screening ratio, including: To determine the screening ratio, the difference between the data volume of the session content obtained by the trigger module and the data volume of the content to be verified is calculated, and the ratio of the difference to the data volume of the session content obtained by the trigger module is solved to obtain the screening ratio; If the screening ratio is less than or equal to the first preset screening ratio, it is determined that the processing of the conversation content is qualified; If the screening ratio is less than or equal to the second preset screening ratio and greater than the first preset screening ratio, determining whether the processing of the session content is qualified in combination with the data volume of the session content obtained by the trigger module; If the screening ratio is greater than the second preset screening ratio, it is determined that the processing of the conversation content is abnormal, and the processing method for the information screening module is determined based on the data volume of the screened sensitive information and the data volume of the screened interference information; Determining whether the processing of the session content is qualified based on the data volume of the session content obtained by the trigger module, including: The data volume of the session content acquired by the trigger module is recorded as the trigger data volume; If the trigger data volume is less than or equal to the preset data volume, it is determined that the processing of the session content is qualified, and the preset trigger duration is adjusted to a corresponding value based on the trigger data volume and the historical data increment; If the trigger data volume is greater than the preset data volume, it is determined that the processing of the session content is abnormal, and the processing method for the information screening module is determined based on the data volume of the filtered sensitive information and the data volume of the filtered interference information; A time domain curve of the trigger data volume is drawn based on each historical trigger data volume obtained by the trigger module, and the slope of the time domain curve of the trigger data volume at the current time node is solved, and the slope is determined as the historical data increment.

2. The data collection and processing system for analyzing user text sentiment tendency according to claim 1 is characterized in that: The sensitive information screening module is used to screen out sensitive information in the session content obtained by the trigger module, including: Select a conversation window containing a preset conversation, perform natural sentence analysis on the conversation in the conversation window, and filter out sensitive information identified by the analysis; The interference information screening module is used to screen out interference information from the session information after sensitive information screening is completed, including: Determine the session information after sensitive information screening is completed as desensitized session information; Divide the desensitized conversation information into keywords, compare each keyword with each screening word in the screening standard library, determine the keywords with similarity greater than the preset similarity as interference information, and screen them out; The similarity is the ratio of the number of characters in the longest continuous identical part between the keyword and the filter word to the total number of characters in the filter word.

3. The data collection and processing system for analyzing user text sentiment tendency according to claim 2 is characterized in that: The verification module is used to adjust the preset trigger duration to a corresponding value based on the amount of trigger data, wherein: The increase range of the preset trigger duration determined based on the amount of trigger data is inversely proportional to the amount of trigger data.

4. The data collection and processing system for analyzing user text sentiment tendency according to claim 3 is characterized in that: The verification module adjusts the preset trigger duration to a corresponding value based on the historical data increment under the condition that the preset trigger duration is adjusted based on the trigger data amount, wherein: The increase in the preset trigger duration determined based on the historical data increment is inversely proportional to the historical data increment.

5. The data collection and processing system for analyzing user text sentiment tendency according to claim 4 is characterized in that: The verification module is used to determine the processing method for the information screening module based on the screening comparison parameter, including: The ratio of the amount of data of the filtered interference information to the amount of data of the filtered sensitive information is recorded as the filtering comparison parameter; If the screening comparison parameter is less than or equal to the preset screening comparison parameter, the session window is adjusted to a corresponding value based on the screening comparison parameter; If the screening comparison parameter is greater than the preset screening comparison parameter, the interference coincidence parameter is calculated based on the obtained interference information of each screening, and the processing method for the interference information screening module is determined based on the interference coincidence parameter.

6. The data collection and processing system for analyzing user text sentiment tendency according to claim 5 is characterized in that: The verification module is used to determine a processing method for the interference information screening module based on the interference coincidence parameter, including: To determine the interference overlap parameter, obtain the filtered interference information, divide the interference information into several data groups so that the same interference information is classified into the same data group, count the data volume of the interference information of each data group, and record the data volume of the data group with the largest data volume as the overlap data volume; solve the ratio of the overlap data volume to the total data volume of the filtered interference information to obtain the interference overlap parameter; If the interference coincidence parameter is less than or equal to the preset interference coincidence parameter, adjusting the preset similarity to a corresponding value based on the screening ratio; If the interference overlap parameter is greater than the preset interference overlap parameter, the preset similarity is adjusted to a corresponding value based on the screening ratio, and the number of each screening word in the screening standard library is adjusted to a corresponding value based on the interference overlap parameter.

7. The data collection and processing system for analyzing user text sentiment tendency according to claim 6 is characterized in that: The verification module is used to adjust the preset similarity to a corresponding value based on the screening ratio, wherein: The increase in the preset similarity determined based on the screening ratio is proportional to the screening ratio; The verification module is used to adjust the number of each screening word in the screening standard library to a corresponding value based on the interference overlap parameter, wherein: The increase in the number of each screening word in the screening standard library determined based on the interference overlap parameter is proportional to the interference overlap parameter.

8. The data collection and processing system for analyzing user text sentiment tendency according to claim 7 is characterized in that: The verification module is used to adjust the session window to a corresponding value based on the screening comparison parameter, wherein: The increase of the session window determined based on the screening comparison parameter is inversely proportional to the screening comparison parameter.

Citation Information

Patent Citations

  • Deep learning-based sentiment analysis methods for product review texts

    CN109840328B

  • Characteristic extraction method of structured interview recording transcription text based on intention slot

    CN116959754A

  • Data interaction system based on artificial intelligence

    CN119025539A