Processing Method for Monitoring Analysis System
By constructing a review data monitoring map, the problem that traditional analysis methods are difficult to understand the comment propagation path and identify key comment nodes is solved, and efficient analysis of comment data and accurate determination of target comment data are achieved.
Patent Information
- Application Number
- CN202510001009.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-02
AI Technical Summary
Traditional comment data analysis methods are difficult to capture the complex relationships between comment data, understand the spread paths of comments, and identify key comment nodes, making it difficult to efficiently extract useful information when facing a large amount of comment data.
By obtaining comment data on public network platforms, generating comment data sets, and abstracting user information into user nodes, generating user node sets and feature vector edge sets, and constructing comment data monitoring maps, thereby accurately determining target comment data.
It realizes efficient analysis of comment data, can accurately capture the characteristic relationship between comment data, identify the propagation path and key comment nodes, and quickly focus on key comment information published by disseminated users.
Smart Images

Figure CN119397066B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data processing technologies, and particularly to a processing method for a monitoring and analysis system. Background Art
[0002] With the rapid development of the Internet, public network platforms such as social media, forums, blogs, etc. have become important channels for people to express opinions and share views. The comment data on these platforms contains a large amount of user feedback and market dynamics information, which has extremely high analysis value for enterprises and individuals.
[0003] Traditional comment data analysis methods often rely on simple text mining technologies such as keyword search, sentiment analysis, etc. Although these methods can provide some basic information, they have obvious deficiencies in capturing the complex relationships between comment data, understanding the propagation paths of comments, and identifying key comment nodes.
[0004] Therefore, in the face of a large amount of comment data, how to efficiently extract useful information, especially those comments with transmissibility and influence, has become an urgent problem to be solved. Summary of the Invention
[0005] The present application provides a processing method for a monitoring and analysis system, which realizes a comment data monitoring graph and can accurately determine target comment data from a comment dataset, so as to quickly focus on the key comment information posted by users with transmissibility.
[0006] In a first aspect, the present application provides a processing method for a monitoring and analysis system, including:
[0007] Obtaining comment data on a public network platform within a preset time period range to generate a comment dataset, where the comment data includes comment time, comment information, and user information;
[0008] Generating a user node set according to the user information of each comment data in the comment dataset, where each user node in the user node set is used to represent the corresponding user information in the comment dataset;
[0009] Generating a set of feature vector edges between each user node in the user node set according to each comment data in the comment dataset, where each feature edge in the set of feature vector edges is used to represent the feature relationship between any two comment data in the comment dataset;
[0010] Generating a comment data monitoring graph corresponding to the comment data according to the user node set and the set of feature vector edges;
[0011] Determine target comment data from the comment dataset according to the comment data monitoring graph.
[0012] In the above solution, by obtaining comment data on the public network platform within a preset time period range, a large amount of user comment information can be systematically collected, covering a wide time span and diverse data sources, ensuring the comprehensiveness and timeliness of the data. Among them, the comment dataset not only includes the comment time and comment information, but also covers user information, providing multi-dimensional data support for subsequent data analysis. Then, a user node set is generated according to the user information in the comment dataset, and each user information is abstracted into a user node, realizing the structured representation of the comment data. Then, by generating a set of feature vector edges between each user node in the user node set, the feature relationship between any two comment data in the comment dataset can be accurately captured. Among them, the feature vector edge not only represents the association between comments, but also contains information such as the direction and strength of the relationship. Capturing this feature relationship is of great significance for understanding the propagation path of comment data and identifying key comment nodes, providing strong support for the subsequent determination of target comment data. Next, a comment data monitoring graph is generated according to the user node set and the set of feature vector edges. The graph presents the user nodes and their feature relationships in a topological structure, enabling the complex structure and internal laws of the data to be effectively expressed through the topological structure. Based on the comment data monitoring graph, the target comment data can be accurately determined from the comment dataset, so as to quickly focus on the key comment information posted by users with dissemination ability.
[0013] Optionally, the generating the set of feature vector edges between each user node in the user node set according to each comment data in the comment dataset includes:
[0014] Determine the comment time difference and user association information according to the first comment data and the second comment data in the comment dataset, where the comment time difference is determined according to the first comment time in the first comment data and the second comment time in the second comment data, and the user association information is determined according to the first user information in the first comment data and the second user information in the second comment data;
[0015] If the comment time difference meets the preset comment time difference condition and the user association information meets the preset user association relationship condition, an initial feature vector edge is generated between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data, where the vector direction of the initial feature vector edge is from the user node corresponding to the comment data with the earlier comment time among the first comment data and the second comment data to the other user node;
[0016] If the comment time difference does not meet the preset comment time difference condition and / or the user association information does not meet the preset user association relationship condition, no initial feature vector edge is generated between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data;
[0017] After generating the initial feature vector edge between the first user node and the second user node, determine the feature weight value of the initial feature vector edge according to the first comment information in the first comment data and the second comment information in the second comment data, so as to generate a corresponding feature vector edge according to the initial feature vector edge and the feature weight value.
[0018] In the above solution, first, the comment time difference and user association information are determined based on the first comment data and the second comment data in the comment dataset. This step provides a reference basis for the subsequent generation of feature vector edges by calculating the comment time difference and evaluating the association relationship between users. The calculation of the comment time difference helps to identify the time sequence between comments, so as to understand the propagation speed and trend of comments. At the same time, the determination of user association information can reveal the potential connections between commenters, such as friendship, social circles, etc., which is crucial for analyzing the influence and propagation path of comments. After determining the comment time difference and user association information, screening is carried out through preset comment time difference conditions and user association relationship conditions. Only when both conditions are met, an initial feature vector edge will be generated between the first user node and the second user node. This condition screening mechanism ensures that the generated feature vector edges have high relevance and accuracy, avoiding the generation of irrelevant or weakly related edges, thereby improving the quality and analysis efficiency of the graph. The generated initial feature vector edge has a clear vector direction, pointing from the user node with an earlier comment time to the user node with a later comment time. This directional representation not only reflects the propagation direction of comments but also helps to understand the flow trend of comments among users. Through the directional vector edges, the diffusion path of comments in the user network can be more intuitively displayed, providing strong support for subsequent public opinion monitoring and trend prediction. After generating the initial feature vector edge, the feature weight value is further determined according to the comment information. The feature weight value reflects the similarity or correlation between comments and is a key indicator for evaluating the importance of feature vector edges. By assigning weight values to feature vector edges, the association strength and influence between comments can be measured more accurately, which is of great significance for identifying key comments, mining potential topics, etc. Finally, the corresponding feature vector edge is generated according to the initial feature vector edge and its feature weight value. This step synthesizes multiple factors such as comment time difference, user association information, and comment information, ensuring that the generated feature vector edge meets the preset conditions and has high relevance and accuracy. Through this comprehensive consideration edge optimization mechanism, the generated comment data monitoring graph can more accurately reflect the comment relationship and propagation trend between users, providing strong support for subsequent data analysis and decision-making.
[0019] Optionally, the determining the feature weight value of the initial feature vector edge according to the first comment information in the first comment data and the second comment information in the second comment data includes:
[0020] Extract the keywords of the first comment information and the second comment information respectively to form corresponding first keyword set and second keyword set;
[0021] Generate a keyword similarity matrix according to the first keyword set and the second keyword set;
[0022] Determine the corresponding first keyword weight set according to the first keyword set, and determine the corresponding second keyword weight set according to the second keyword set;
[0023] Determine the feature weight value according to the first keyword weight set, the second keyword weight set, and the keyword similarity matrix.
[0024] In the above solution, first, keywords are extracted from the first comment information and the second comment information respectively to form the corresponding first keyword set and second keyword set. This step can effectively extract the core content in the comments, providing a basis for subsequent analysis. The construction of the keyword set not only simplifies the complexity of the comment information but also retains the core semantics of the comments, making the analysis more focused and accurate. Based on the first keyword set and the second keyword set, a keyword similarity matrix is generated. The similarity matrix provides a basis for evaluating the relevance between comments by quantifying the similarity between keywords. Through similarity calculation, the commonalities and differences in the comments can be accurately identified, providing a basis for determining the subsequent feature weight value. In addition, the corresponding first keyword weight set and second keyword weight set are further determined according to the first keyword set and the second keyword set respectively. The assignment of weights considers factors such as the importance and frequency of occurrence of keywords in the comments, making the weights of the keywords more reasonable and accurate. Through the reasonable assignment of weights, the influence of keywords in the comments can be evaluated more objectively, providing a basis for calculating the subsequent feature weight value. After determining the first keyword weight set, the second keyword weight set, and the keyword similarity matrix, the feature weight value is calculated by integrating these factors, so that the feature weight value can not only reflect the semantic similarity between comments but also consider the importance differences of keywords. By comprehensively considering multiple factors, a more accurate and comprehensive feature weight value can be generated, providing support for constructing the subsequent comment data monitoring graph. Through the above steps, the accurate extraction, similarity calculation, weight assignment, and comprehensive consideration of the feature weight value of keywords in the comment information are realized. These steps are progressive and jointly improve the accuracy and depth of data analysis. The accurate feature weight value helps to more accurately represent the feature relationship between user nodes in the comment data monitoring graph, providing strong data support for subsequent determination of target comment data, public opinion monitoring, etc.
[0025] Optionally, the determining the corresponding first keyword weight set according to the first keyword set and the corresponding second keyword weight set according to the second keyword set includes:
[0026] Determine the first keyword weight set according to the word frequency, inverse document frequency, and the number of occurrences in the comment dataset of each keyword in the first keyword set;
[0027] Determine the second keyword weight set according to the word frequency, inverse document frequency, and the number of occurrences in the comment dataset of each keyword in the second keyword set.
[0028] In the above solution, when determining the keyword weight, three dimensions of word frequency, inverse document frequency, and the number of occurrences of the keyword in the comment dataset are comprehensively considered. This multi-dimensional consideration method makes the distribution of keyword weights more comprehensive and objective. The word frequency reflects the importance of the keyword in a single comment, the inverse document frequency reflects the distinctiveness of the keyword in the entire comment dataset, and the number of occurrences further strengthens the universality and influence of the keyword. Through the comprehensive evaluation of these three indicators, the actual weight of the keyword can be measured more accurately. Through the comprehensive calculation of word frequency, inverse document frequency, and the number of occurrences, the deviation that may be brought by a single indicator can be avoided, thereby improving the accuracy of keyword weights. For example, relying solely on word frequency may ignore some low-frequency but highly relevant keywords, and the introduction of inverse document frequency can make up for this deficiency and highlight the uniqueness of these keywords. At the same time, the consideration of the number of occurrences also enhances the stability of weight distribution, making the weight value more in line with the actual distribution of keywords in the comment dataset. Precise keyword weight distribution helps to more precisely describe the semantic relationship between comments in subsequent data analysis. Among them, the keyword weight not only affects the generation and weight assignment of the feature vector edge, but also directly relates to the construction quality of the comment data monitoring graph. By assigning reasonable weights to keywords, the nodes and edges in the graph can more accurately reflect the actual structure and characteristics of the comment data, providing more powerful support for subsequent determination of target comment data, public opinion monitoring, etc.
[0029] Optionally, the determining the comment time difference and user association information according to the first comment data and the second comment data in the comment dataset includes:
[0030] Determine the comment time difference according to the time difference between the first comment time and the second comment time;
[0031] Determine the user association feature value according to the first user information and the second user information;
[0032] Correspondingly, when the comment time difference is greater than the preset comment time difference threshold and the user association feature value is greater than the preset user association feature threshold, it is determined that the comment time difference meets the preset comment time difference condition and the user association information meets the preset user association relationship condition.
[0033] In the above solution, first, calculate the time difference between the first comment time and the second comment time to determine the comment time difference, which helps to accurately identify the time interval between comments and provides basic data for subsequent analysis of the comment propagation speed and trend. Moreover, through the precise calculation of the time difference, the time correlation between comments can be evaluated more objectively, providing a basis for generating feature vector edges subsequently. Then, determine the user association eigenvalue according to the first user information and the second user information to quantify the degree of association between users. The calculation of the user association eigenvalue may involve multiple dimensions, such as the social relationship between users, the interaction frequency, and common interests. By evaluating the user association eigenvalue, the association strength between users can be judged more accurately, providing support for subsequent analysis of the diffusion path of comments in the user network. After determining the comment time difference and the user association eigenvalue, perform screening through preset comment time difference thresholds and user association feature thresholds. Only when the comment time difference is greater than the preset comment time difference threshold and the user association eigenvalue is greater than the preset user association feature threshold is it considered to meet the conditions. This conditional screening mechanism ensures that the generated feature vector edges have a high degree of relevance and importance, avoiding the generation of irrelevant or weakly relevant edges, thereby improving the quality and analysis efficiency of the comment data monitoring graph. It can be seen that by precisely calculating the comment time difference and evaluating the user association eigenvalue, and combining a strict conditional screening mechanism, more accurate and meaningful feature vector edges can be generated. These feature vector edges not only reflect the time sequence and user association between comments but also consider the importance of comment content and keywords, providing more accurate data support for subsequent public opinion monitoring.
[0034] Optionally, determining the target comment data from the comment data set according to the comment data monitoring graph includes:
[0035] Obtain multiple comment propagation paths from the comment data monitoring graph, where the comment propagation path includes a target user node sequence and target feature vector edges for connecting two adjacent target user nodes in the target user node sequence, and the target feature vector edges all point from the previous target user node in the target user node sequence to the next target user node;
[0036] Determine the comment data corresponding to the target user nodes on the comment propagation path as the target comment data.
[0037] In the above solution, multiple comment propagation paths are obtained from the comment data monitoring graph. These paths intuitively represent the diffusion process and direction of comments in the user network in a topological manner. Through the comment propagation paths, it can be analyzed how comments spread from the initial users to other users, as well as the key nodes and paths in the propagation process, providing a basis for subsequent analysis. After determining the comment propagation paths, the comment data corresponding to the target user nodes on the paths is further located as the target comment data. This positioning method ensures that the target comment data is closely related to the comment propagation paths and has important influence, providing data support for subsequent analysis and decision-making. By directly obtaining the comment propagation paths and target comment data from the comment data monitoring graph, the cumbersome data screening and filtering steps in traditional data analysis are avoided, improving the efficiency of data analysis. At the same time, since both the comment propagation paths and the target comment data are generated based on the topological comment data monitoring graph, they have higher accuracy and reliability. The accurately located target comment data provides decision-makers with clearer and more comprehensive information such as public opinion dynamics and user feedback. Based on this information, decision-makers can make responses and decisions more quickly, such as adjusting marketing strategies and optimizing product functions, thereby enhancing the competitiveness of the enterprise and the market response speed.
[0038] In a second aspect, the present application provides a monitoring and analysis system, including:
[0039] An acquisition module, configured to acquire comment data on a public network platform within a preset time period range to generate a comment data set, where the comment data includes comment time, comment information, and user information;
[0040] A processing module, configured to generate a user node set according to the user information of each comment data in the comment data set, where each user node in the user node set is used to represent the corresponding user information in the comment data set;
[0041] The processing module is further configured to generate a set of feature vector edges between each user node in the user node set according to each comment data in the comment data set, where each feature edge in the set of feature vector edges is used to represent the feature relationship between any two comment data in the comment data set;
[0042] The processing module is further configured to generate a comment data monitoring graph corresponding to the comment data according to the user node set and the set of feature vector edges;
[0043] The processing module is further configured to determine target comment data from the comment data set according to the comment data monitoring graph.
[0044] Optionally, the processing module is specifically configured to:
[0045] Determine the comment time difference and user association information based on the first comment data and the second comment data in the comment dataset, where the comment time difference is determined according to the first comment time in the first comment data and the second comment time in the second comment data, and the user association information is determined according to the first user information in the first comment data and the second user information in the second comment data;
[0046] If the comment time difference meets the preset comment time difference condition and the user association information meets the preset user association relationship condition, then generate an initial feature vector edge between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data, where the vector direction of the initial feature vector edge is from the user node corresponding to the comment data with the earlier comment time in the first comment data and the second comment data to the other user node;
[0047] If the comment time difference does not meet the preset comment time difference condition and / or the user association information does not meet the preset user association relationship condition, then no initial feature vector edge is generated between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data;
[0048] After generating the initial feature vector edge between the first user node and the second user node, determine the feature weight value of the initial feature vector edge according to the first comment information in the first comment data and the second comment information in the second comment data, so as to generate a corresponding feature vector edge according to the initial feature vector edge and the feature weight value.
[0049] Optionally, the processing module is specifically configured to:
[0050] Extract the keywords of the first comment information and the second comment information respectively to form corresponding first keyword sets and second keyword sets;
[0051] Generate a keyword similarity matrix according to the first keyword set and the second keyword set;
[0052] Determine a corresponding first keyword weight set according to the first keyword set, and determine a corresponding second keyword weight set according to the second keyword set;
[0053] Determine the feature weight value according to the first keyword weight set, the second keyword weight set and the keyword similarity matrix.
[0054] Optionally, the processing module is specifically configured to:
[0055] Determine the first keyword weight set according to the word frequency, inverse document frequency, and the number of occurrences in the comment dataset of each keyword in the first keyword set;
[0056] Determine the second keyword weight set according to the word frequency, inverse document frequency, and the number of occurrences in the comment dataset of each keyword in the second keyword set.
[0057] Optionally, the processing module is specifically configured to:
[0058] Determine the comment time difference according to the time difference between the first comment time and the second comment time;
[0059] Determine the user association feature value according to the first user information and the second user information;
[0060] Correspondingly, when the comment time difference is greater than the preset comment time difference threshold and the user association feature value is greater than the preset user association feature threshold, it is determined that the comment time difference meets the preset comment time difference condition and the user association information meets the preset user association relationship condition.
[0061] Optionally, the processing module is specifically configured to:
[0062] Obtain multiple comment propagation paths from the comment data monitoring graph, where the comment propagation path includes a target user node sequence and target feature vector edges for connecting two adjacent target user nodes in the target user node sequence, and the target feature vector edges point from the previous target user node in the target user node sequence to the next target user node;
[0063] Determine that the comment data corresponding to the target user node on the comment propagation path is the target comment data.
[0064] In a third aspect, the present application provides an electronic device, including:
[0065] A processor; and,
[0066] A memory for storing executable instructions of the processor;
[0067] Wherein, the processor is configured to execute any possible method described in the first aspect by executing the executable instructions.
[0068] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement any possible method described in the first aspect.
[0069] The processing method of the monitoring and analysis system provided by this application generates a comment dataset by obtaining comment data on a public network platform within a preset time range. Then, a user node set is generated based on the user information of each comment data in the comment dataset. Next, a feature vector edge set between each user node in the user node set is generated based on each comment data in the comment dataset. Thus, a comment data monitoring graph corresponding to the comment data is generated based on the user node set and the feature vector edge set, and target comment data is determined from the comment dataset based on the comment data monitoring graph, so as to accurately determine the target comment data from the comment dataset based on the comment data monitoring graph, and quickly focus on the key comment information posted by users with dissemination ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0071] Figure 1 is a schematic flowchart of the processing method of the monitoring and analysis system shown by this application according to an exemplary embodiment;
[0072] Figure 2 is a schematic flowchart of the processing method of the monitoring and analysis system shown by this application according to another exemplary embodiment;
[0073] Figure 3 is a schematic structural diagram of the monitoring and analysis system shown by this application according to an exemplary embodiment;
[0074] Figure 4 is a schematic structural diagram of an electronic device shown by this application according to an exemplary embodiment.
[0075] Through the above accompanying drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These accompanying drawings and written descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different accompanying drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0077] To solve the above problems, the embodiments provided by this application mainly include the following aspects in the main technical concept:
[0078] Multi-dimensional data collection and integration: First, by obtaining the comment data on public network platforms within a preset time range, a multi-dimensional comment data set including comment time, comment information, and user information was constructed. This data collection method ensures the comprehensiveness and timeliness of the data, providing a solid foundation for subsequent data analysis.
[0079] Structured representation of comment data: To better understand and analyze the comment data, the user information in the comment data set was abstracted into user nodes, generating a set of user nodes. Each user node not only contains the unique identification information of the user but also reflects the potential connections between users through the connection relationships between nodes. This structured representation method makes the comment data more intuitive, easy to process, and analyze.
[0080] Accurate capture of feature relationships: Further, the concept of feature vector edges was proposed to characterize the feature relationships between any two comment data in the comment data set. Feature vector edges not only contain the association information between comments but also reflect the direction and strength of the relationship through the vector direction and weight values. This way of accurately capturing feature relationships provides strong support for understanding the propagation path of comment data and identifying key comment nodes.
[0081] Construction of a comment data monitoring graph: Based on the set of user nodes and the set of feature vector edges, a comment data monitoring graph was generated. The graph presents the user nodes and their feature relationships in a topological structure, enabling the effective expression of the complex structure and internal laws of the data through the graph. This visual expression method not only helps users intuitively understand the comment data but also provides convenience for subsequent data analysis.
[0082] Quick determination of target comment data: By analyzing the comment data monitoring graph, the target comment data can be accurately determined from the comment data set. Target comment data usually has high dissemination and influence and can reflect the mainstream trend of public opinion. This way of quickly determining target comment data provides strong support for public opinion monitoring, market analysis, etc., and helps enterprises and individuals grasp market dynamics and user feedback in a timely manner.
[0083] Figure 1 It is a schematic flowchart of the processing method of the monitoring and analysis system shown in an exemplary embodiment of the present application. As Figure 1 shown, the processing method of the monitoring and analysis system provided in this embodiment includes:
[0084] S101. Obtain the comment data on public network platforms within a preset time range to generate a comment data set.
[0085] In this step, obtain the comment data on the public network platform within a preset time period to generate a comment data set. The comment data includes the comment time, comment information, and user information.
[0086] First, determine a preset time period, such as the past week or month. Then, collect the comment data within this time period from public network platforms (such as Weibo, news websites, etc.). The collected comment data should include the comment time, comment information (i.e., the content of the comment made by the user), and user information (such as the username, user ID, etc.). Organize this comment data into a data set for subsequent processing and analysis. It should be noted that the above-mentioned public network platform refers to a network platform where information can be legally collected. If user authorization is required, before obtaining the corresponding comment data, the user's authorization needs to be obtained.
[0087] S102. Generate a user node set according to the user information of each comment data in the comment data set.
[0088] In this step, generate a user node set according to the user information of each comment data in the comment data set. Each user node in the user node set is used to represent the corresponding user information in the comment data set. Traverse the comment data set and extract the user information in each comment data. For each unique user information, create a corresponding user node in memory. These user nodes form a user node set, and each user node is used to represent the corresponding user information in the comment data set.
[0089] S103. Generate a set of feature vector edges between each user node in the user node set according to each comment data in the comment data set.
[0090] In this step, generate a set of feature vector edges between each user node in the user node set according to each comment data in the comment data set. Each feature edge in the set of feature vector edges is used to represent the feature relationship between any two comment data in the comment data set.
[0091] Specifically, for any two comment data in the comment dataset, first calculate the comment time difference and user association information between them. The comment time difference is determined by comparing the comment times of the two comment data; the user association information can be determined by comparing the followed objects, followed-by objects, or historical comment records in the user information of the two comment data. If the comment time difference meets the preset conditions (such as being less than a certain time threshold or greater than a certain time threshold, which can be set accordingly according to the monitoring purpose) and the user association information meets the preset relationship conditions (such as belonging to the same social circle or having frequent interactions), then generate an initial feature vector edge between the user nodes corresponding to the two comment data. The vector direction of the initial feature vector edge points from the user node with an earlier comment time to the user node with a later comment time. Then, calculate the feature weight value of the initial feature vector edge according to the comment information in the two comment data. This can be determined by extracting the keywords of the comment information, calculating the keyword similarity matrix, and combining the keyword weights. Finally, generate the corresponding feature vector edge according to the initial feature vector edge and its feature weight value, and add it to the feature vector edge set.
[0092] S104. Generate a comment data monitoring graph corresponding to the comment data according to the user node set and the feature vector edge set.
[0093] Specifically, use the user node set as the vertex set of the graph and the feature vector edge set as the edge set of the graph to generate a comment data monitoring graph. The graph presents the feature relationships between user nodes in a topological structure, enabling the effective expression of the complex structure and internal laws of the data through the topological structure.
[0094] S105. Determine the target comment data from the comment dataset according to the comment data monitoring graph.
[0095] In the generated comment data monitoring graph, screen the comment data that meets specific conditions as the above-mentioned target comment data, or use search algorithms (such as depth-first search, breadth-first search, etc.) to search for comment propagation paths. The comment propagation path is a sequence of user nodes connected by feature vector edges, indicating the propagation path of the comment among users. After obtaining multiple such comment propagation paths from the graph, determine the comment data corresponding to the target user nodes on the path as the target comment data. These target comment data often have high propagation and influence, and are the key objects of concern in public opinion monitoring and market analysis.
[0096] In this embodiment, by obtaining the comment data on the public network platform within a preset time range to generate a comment data set, then generating a user node set according to the user information of each comment data in the comment data set, and then generating a set of feature vector edges between each user node in the user node set according to each comment data in the comment data set, so as to generate a comment data monitoring graph corresponding to the comment data according to the user node set and the set of feature vector edges, and determining target comment data from the comment data set according to the comment data monitoring graph, so as to realize that based on the comment data monitoring graph, the target comment data can be accurately determined from the comment data set, so as to quickly focus on the key comment information published by users with dissemination ability.
[0097] Figure 2 It is a schematic flowchart of the processing method of the monitoring and analysis system shown according to another exemplary embodiment of the present application. As Figure 2 shown, the processing method of the monitoring and analysis system provided in this embodiment includes:
[0098] S201. Obtain the comment data on the public network platform within a preset time range to generate a comment data set.
[0099] In this step, obtain the comment data on the public network platform within a preset time range to generate a comment data set, where the comment data includes the comment time, comment information, and user information.
[0100] First, determine a preset time period, such as the past week or month. Then, collect the comment data within this time period from public network platforms (such as Weibo, news websites, etc.). The collected comment data should include the comment time, comment information (i.e., the comment content published by the user), and user information (such as the username, user ID, etc.). Organize these comment data into a data set for subsequent processing and analysis. It should be noted that the above-mentioned public network platform refers to a network platform where information can be legally collected. If user authorization is required, the user's authorization needs to be obtained before obtaining the corresponding comment data.
[0101] S202. Generate a user node set according to the user information of each comment data in the comment data set.
[0102] In this step, generate a user node set according to the user information of each comment data in the comment data set, where each user node in the user node set is used to represent the corresponding user information in the comment data set. Traverse the comment data set and extract the user information in each comment data. For each unique user information, create a corresponding user node in memory. These user nodes form a user node set, and each user node is used to represent the corresponding user information in the comment data set.
[0103] S203. Determine the comment time difference and user association information based on the first comment data and the second comment data in the comment dataset.
[0104] In this step, determine the comment time difference and user association information based on the first comment data and the second comment data in the comment dataset. Among them, the comment time difference is determined according to the first comment time in the first comment data and the second comment time in the second comment data, and the user association information is determined according to the first user information in the first comment data and the second user information in the second comment data.
[0105] In a possible implementation, determine the comment time difference according to the time difference between the first comment time and the second comment time.
[0106] Use formula 3 and determine the user association eigenvalue according to the first user information and the second user information Formula 3 is:
[0107]
[0108] Among them, is the set of target objects concerned by the first user in the first user information, is the set of target objects concerned by the second user in the second user information, is the set of target objects concerned by the first user in the first user information, is the set of target objects concerned by the second user in the second user information, is the number of elements in the set.
[0109] In the above solution, the comment time difference is one of the basic data for time series analysis. By analyzing the comment time difference, the distribution and changes of comment data in the time dimension can be understood, providing an important basis for subsequent public opinion trend prediction, user behavior pattern recognition, etc. In real-time monitoring and analysis scenarios, the comment time difference can help the system quickly capture the real-time dynamics of comment data, such as the outbreak of hot events, the rapid change of user emotions, etc., so as to make timely responses and decisions. In the recommendation system, considering the comment time difference can help optimize the recommendation algorithm and improve the timeliness and relevance of recommendations. For example, for target objects that have generated a large number of comments recently, they can be preferentially recommended to users to meet the real-time needs of users.
[0110] The above formula 3 is used to calculate the user association eigenvalue, comprehensively considering the target objects jointly followed by users, the target objects jointly followed by others, as well as the number of target objects independently followed and followed by each user. This multi-dimensional consideration method can more comprehensively evaluate the association between users, representing the potential connections and interaction patterns between users. By calculating the user association eigenvalue, the user portrait can be further enriched and improved. For example, adding the association information between users to the user portrait helps to more accurately understand the characteristics of users' interest preferences, social circles, etc., thus providing strong support for applications such as personalized recommendation and community discovery. In community discovery and analysis, the user association eigenvalue can be used as an important indicator to measure the closeness between users. By calculating the association eigenvalues between different users, user groups with close connections can be identified, and then the structural characteristics and evolution laws of the community can be analyzed, providing a basis for community management and marketing strategy formulation. In the information dissemination network, the user association eigenvalue can help identify key dissemination nodes and paths. By preferentially pushing information to users with high association eigenvalues, the dissemination efficiency and coverage of information can be improved, thereby optimizing the information dissemination path and strategy.
[0111] In another possible implementation, formula 4 can also be used to determine the user association eigenvalue according to the first user information and the second user information. , formula 4 is:
[0112]
[0113] Wherein, is the set of target objects for which the first user in the first user information has at least historical one-way public comments, is the set of target objects for which the second user in the second user information has at least historical one-way public comments, is the number of elements in the set;
[0114] Correspondingly, when the comment time difference is greater than the preset comment time difference threshold and the user association eigenvalue is greater than the preset user association feature threshold, it is determined that the comment time difference meets the preset comment time difference condition and the user association information meets the preset user association relationship condition.
[0115] In the above solution, Formula 4 quantifies the relevance between users by calculating the ratio of the intersection to the union of the historical one-way public comments of two users on the target object. This quantification method helps to objectively evaluate the similarity and interaction intensity between users. By calculating the user association eigenvalue, the user profile can be further improved. Incorporating the relevance between users into the user profile can more comprehensively reflect features such as users' interest preferences and social circles, providing more accurate data support for applications such as personalized recommendation and community discovery. In the information dissemination network, the user association eigenvalue can be used as an important indicator to measure the information dissemination efficiency between users. By preferentially pushing information to users with high association eigenvalues, the information dissemination path can be optimized, and the information coverage rate and dissemination speed can be increased. In public opinion analysis, the user association eigenvalue helps to identify key opinion leaders and community structures. By analyzing the relevance between users, influential user groups and potential public opinion trends can be discovered, providing strong support for public opinion early warning and response.
[0116] When the comment time difference is greater than the preset comment time difference threshold and the user association eigenvalue is greater than the preset user association threshold, it can be determined that the comment time difference meets the preset conditions and the user association information meets the preset relationship conditions. This comprehensive judgment criterion helps to screen out comment data that is both timely and highly relevant, providing more accurate and valuable data support for subsequent public opinion monitoring, user behavior analysis, etc. It should be noted that the above preset comment time difference condition can also be set as the comment time difference being greater than the first preset comment time difference threshold and less than the second preset comment time difference threshold.
[0117] S204. Generate an initial feature vector edge between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data.
[0118] If the comment time difference meets the preset comment time difference condition and the user association information meets the preset user association relationship condition, an initial feature vector edge is generated between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data, where the vector direction of the initial feature vector edge is from the user node corresponding to the comment data with the earlier comment time in the first comment data and the second comment data to the other user node.
[0119] If the comment time difference does not meet the preset comment time difference condition and / or the user association information does not meet the preset user association relationship condition, no initial feature vector edge is generated between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data.
[0120] Specifically, for any two comment data in the comment dataset (denoted as the first comment data and the second comment data), the system first calculates the time difference between their comments and obtains their respective user information. Then, the system makes a judgment based on the preset comment time difference condition (such as the comment time difference threshold) and the user association relationship condition (such as the user association eigenvalue threshold). Only when the comment time difference meets the preset comment time difference condition and the user association information meets the preset user association relationship condition, will the system perform the next initial feature vector edge generation operation.
[0121] If both of the above conditions are satisfied, the system will generate an initial feature vector edge between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data. The vector direction of this initial feature vector edge is set to point from the user node corresponding to the comment data with an earlier comment time to the user node corresponding to the comment data with a later comment time. This directional setting helps to reflect the propagation path and temporal relationship of comments among users.
[0122] If the comment time difference does not meet the preset comment time difference condition and / or the user association information does not meet the preset user association relationship condition, the system will not generate an initial feature vector edge between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data. Such a processing logic helps to ensure that the generated feature vector edges have a high degree of relevance and accuracy, and avoid the generation of irrelevant or weakly related edges.
[0123] Through the above solution, the system fully considers the time sequence of comments and the association relationship among users when generating the initial feature vector edges, thus ensuring that the generated feature vector edges not only conform to the actual data characteristics but also can accurately reflect the propagation path and temporal relationship of comments in the user network. Through condition judgment and edge generation logic, the system can generate more accurate and meaningful feature vector edges, providing strong support for subsequent data analysis and public opinion monitoring. The directional setting of the initial feature vector edges makes the nodes and edges in the comment data monitoring graph more intuitively display the propagation path and temporal relationship of comments in the user network, helping users to more quickly understand the data and analysis results. By avoiding the generation of irrelevant or weakly related edges, the system can reduce unnecessary computing and analysis burdens, optimize resource utilization, and improve processing efficiency.
[0124] S205. Determine the feature weight value of the initial feature vector edge according to the first comment information in the first comment data and the second comment information in the second comment data, so as to generate the corresponding feature vector edge according to the initial feature vector edge and the feature weight value.
[0125] Specifically, after generating an initial feature vector edge between the first user node and the second user node, determine the feature weight value of the initial feature vector edge according to the first comment information in the first comment data and the second comment information in the second comment data, so as to generate a corresponding feature vector edge according to the initial feature vector edge and the feature weight value.
[0126] In a possible implementation, extract the keywords of the first comment information and the second comment information respectively to form a corresponding first keyword set and a second keyword set , where is the number of keywords in the first keyword set , is the number of keywords in the second keyword set ;
[0127] Generate a keyword similarity matrix according to the first keyword set and the second keyword set , where the keyword similarity matrix is: is:
[0128]
[0129] where the in the keyword similarity matrix is the similarity between the th keyword in the first keyword set and the th keyword in the second keyword set ;
[0130] Determine the corresponding first keyword weight set according to the first keyword set , and determine the corresponding second keyword weight set according to the second keyword set ; ; ;
[0131] Use formula 1 and determine the feature weight value according to the first keyword weight set , the second keyword weight set and the keyword similarity matrix , where formula 1 is:
[0132]
[0133] where is a preset constant.
[0134] In the above solution, by extracting the keywords of the first comment information and the second comment information and generating a keyword similarity matrix, the correlation between the two comments in terms of content can be quantified. Each element in the keyword similarity matrix represents the similarity between the corresponding keywords in the two comments. The calculation of this similarity is based on semantic-level matching, which can more accurately reflect the essential connection of the comment content. When calculating the feature weight value, not only the similarity between keywords is considered, but also a keyword weight set is introduced, which represents the importance of each keyword in the first comment information and the second comment information respectively. This processing method enables the importance degree of different keywords in the comment to be considered when calculating the feature weight value, avoiding the limitation of treating all keywords equally, and improving the accuracy and rationality of the weight value calculation. In addition, the above formula 1 effectively balances the influence of keyword similarity and keyword weight on the feature weight value. Through multiplication and exponentiation operations, formula 1 not only considers the similarity between keywords, but also weights the similarity through the weight value, enabling the feature weight value to more comprehensively reflect the comprehensive relationship between the two comments. At the same time, the preset constant in the formula is used to avoid calculation problems caused by a similarity of 0, ensuring the stability and reliability of the calculation. By quantifying the correlation between comments and considering the importance of keywords, the feature weight value between comment data can be determined more accurately, thereby improving the accuracy of data analysis. In addition, since a structured method is adopted to process comment data, the data analysis process is more efficient, capable of quickly processing large-scale data sets and meeting the requirements of real-time monitoring and analysis.
[0135] In other words, that is, the above formula 1 not only considers the similarity between keywords, but also introduces keyword weights, achieving a comprehensive quantification of the relationship between comment data. This comprehensive consideration method makes the calculated feature weight value more accurately reflect the actual correlation strength between comments, improving the accuracy of data analysis. By introducing a preset constant, formula 1 effectively avoids calculation problems caused when the keyword similarity is 0. This processing method ensures the stability and reliability of the calculation, and reasonable feature weight values can be obtained even in extreme cases. The exponential part in formula 1 realizes normalization processing, making the finally obtained feature weight value between 0 and 1 (or adjusted to other reasonable ranges according to specific situations), which is convenient for subsequent data processing and comparison. Since formula 1 adopts a structured calculation method, it can efficiently process large-scale data sets. When processing a large amount of comment data, it can quickly calculate the feature weight value between each pair of comments, providing strong support for subsequent monitoring and analysis work. By quantifying the feature weight value between comment data, formula 1 supports in-depth analysis of complex relationships. For example, in public opinion monitoring, key comment nodes and propagation paths can be identified based on the feature weight value, providing a basis for public opinion guidance and control.
[0136] Further, for the above-mentioned first keyword set to determine the corresponding first keyword weight set , and for the second keyword set to determine the corresponding second keyword weight set , specifically, it may include:
[0137] Using formula 2 and based on the first keyword set to determine the first keyword weight set , and based on the second keyword set to determine the corresponding second keyword weight set , where formula 2 is:
[0138]
[0139] where is the word frequency of the th keyword in the first keyword set , is the inverse document frequency of the th keyword in the first keyword set , is the number of occurrences of the th keyword in the first keyword set in the comment dataset, is the word frequency of the th keyword in the second keyword set , is the inverse document frequency of the th keyword in the second keyword set , is the number of occurrences of the th keyword in the second keyword set in the comment dataset.
[0140] In the above solution, formula 2 comprehensively considers three factors: word frequency, inverse document frequency, and the number of occurrences of keywords in the comment dataset. These factors together determine the weight of keywords. Word frequency reflects the importance of keywords in a single comment, inverse document frequency reflects the distinctiveness of keywords in the entire comment dataset, and the number of occurrences further emphasizes the universality and importance of keywords in the dataset. This multi-dimensional consideration method makes the calculation of keyword weights more accurate and can better reflect the actual role of keywords in the comment data.
[0141] Since the keyword weight is an important basis for subsequent calculation of feature weight values (as shown in Formula 1), improving the accuracy of keyword weight will directly enhance the precision of data analysis. In the monitoring and analysis system, this means being able to more accurately identify key information and trends in comment data, providing stronger support for decision-making.
[0142] Each factor in Formula 2 can be adjusted or optimized according to actual needs to adapt to different analysis scenarios and data characteristics. For example, in some cases, more attention can be paid to the role of word frequency or inverse document frequency, while in other cases, the number of occurrences of keywords in the dataset may be more emphasized. This flexibility enables the system to better adapt to different analysis requirements, improving processing efficiency and effectiveness.
[0143] In addition, the calculation process of Formula 2 is relatively simple and efficient, capable of supporting the processing of large-scale comment datasets. In the monitoring and analysis system, this means being able to quickly process a large amount of comment data, extract key information and conduct analysis, thus meeting the needs of real-time monitoring and rapid response. By calculating keyword weights and combining with the calculation of feature weight values (as shown in Formula 1), the system can better discover knowledge patterns and potential trends in comment data. This is of great significance for fields such as public opinion monitoring and market analysis, and can help decision-makers timely discover and respond to potential problems and opportunities.
[0144] S206. Obtain multiple comment propagation paths from the comment data monitoring graph.
[0145] In this step, multiple comment propagation paths are obtained from the comment data monitoring graph. Among them, the comment propagation path includes a target user node sequence and target feature vector edges used to connect two adjacent target user nodes in the target user node sequence. The target feature vector edges all point from the previous target user node in the target user node sequence to the next target user node.
[0146] Specifically, the system constructs a comment data monitoring graph based on the comment dataset. This graph consists of user nodes and feature vector edges. Among them, the user nodes represent the publishers of comment data, and the feature vector edges are generated according to the time difference and user association information between comment data, reflecting the propagation path and time sequence relationship of comments among users.
[0147] In the comment data monitoring graph, the system extracts multiple comment propagation paths. These paths consist of a series of target user node sequences. Each node in the sequence represents a user, and the feature vector edges used to connect two adjacent nodes indicate the direction of comment propagation from the previous user to the next user.
[0148] S207. Determine the comment data corresponding to the target user nodes on the comment propagation path as the target comment data.
[0149] The system traverses each extracted comment propagation path and analyzes each target user node on the path one by one. For each target user node on the path, the system maps it back to the original comment dataset to find the comment data corresponding to the user node. The system determines the comment data corresponding to all target user nodes on the comment propagation path as the target comment data. Since these comment data have been widely spread in the user network, they have relatively high attention and influence.
[0150] It should be noted that, based on the above embodiments, after obtaining multiple comment propagation paths from the comment data monitoring graph, the same target user nodes on different comment propagation paths can be merged according to the multiple comment propagation paths to generate a comment propagation network. Then, characteristic user nodes are determined according to the comment propagation network. The characteristic user nodes include the original comment user nodes, the propagation key user nodes, and the propagation active user nodes. Among them, the original comment user nodes include the starting user nodes of each comment propagation path in the comment propagation network. The propagation key user nodes include the user nodes whose number of characteristic vector edges with the vector direction pointing to other user nodes in the comment propagation network is greater than a preset number threshold. The propagation active user nodes include the user nodes whose vector direction points to other user nodes in the comment propagation network and at the same time there are other user nodes pointing to this node.
[0151] Among them, the system first traverses all the extracted comment propagation paths, identifies and merges the same target user nodes on different paths. This step is achieved by comparing the unique identifiers of user nodes (such as user IDs) to ensure that each user is represented only once in the network. After merging the same user nodes, the system constructs a comment propagation network based on the remaining nodes and characteristic vector edges. This network is a directed graph, where the nodes represent users and the edges represent the propagation paths and directions of comments between users. The system traverses each comment propagation path in the comment propagation network and marks the starting user node of the path as the original comment user node. These nodes are the starting points of comment propagation and are of great significance for understanding the origin of public opinion. The system counts the number of characteristic vector edges with the vector direction pointing to other user nodes for each user node in the comment propagation network. When the number of characteristic vector edges of a certain node exceeds the preset number threshold, this node is marked as a propagation key user node. These nodes play an important role as a transit in the comment propagation process and have a significant impact on the spread of public opinion. The system further analyzes the user nodes in the comment propagation network to find those nodes that both point to other user nodes and are at the same time pointed to by other user nodes. These nodes are marked as propagation active user nodes, and they show a high degree of interactivity in the comment propagation network and play an important role in the fermentation and spread of public opinion.
[0152] In the above solution, after obtaining multiple comment propagation paths, a comment propagation network is generated by merging the same target user nodes on different paths to effectively integrate the scattered comment propagation information and construct a global comment propagation view. The construction of the comment propagation network connects the originally isolated comment propagation paths with each other, forming a more complete and coherent information network, which helps to more comprehensively understand the propagation mode and trend of comments among user groups. Based on the comment propagation network, characteristic user nodes are further determined, including the original comment user node, the key propagation user node, and the active propagation user node. The original comment user node, as the starting point of comment propagation, its identification helps to trace the origin and initial impact of the comment. The key propagation user node and the active propagation user node respectively represent important transfer stations and active participants in the comment propagation process, and their identification helps to deeply understand the diffusion mechanism and influence distribution of comments in the user network. In addition, different types of characteristic user nodes are distinguished by setting different criteria (such as the number of characteristic vector edges whose vector direction points to other user nodes, whether there are other user nodes pointing to this node, etc.). This multi-dimensional user classification method makes the analysis more refined and accurate, and can better capture the roles and contributions of different users in the comment propagation process. After identifying the characteristic user nodes, the monitoring and analysis system can focus on and allocate resources to these nodes more targeted. For example, for the key propagation user node and the active propagation user node, the system can strengthen the monitoring of their comment content and propagation behavior to timely discover potential risks or opportunities, which helps to improve the monitoring efficiency and reduce unnecessary resource waste. Based on the analysis results of the comment propagation network and the characteristic user nodes, enterprises can formulate marketing strategies, crisis public relations plans, etc. more scientifically. For example, by understanding the characteristics and influence of the original comment user node, enterprises can better locate the target audience and formulate precise marketing plans; by analyzing the behavior patterns of the key propagation user node and the active propagation user node, enterprises can more effectively guide the trend of public opinion and optimize crisis response strategies. The analysis of the comment propagation network and the characteristic user nodes not only helps the current monitoring and analysis tasks, but also provides a basis for subsequent knowledge discovery and application. For example, by analyzing historical comment data, enterprises can discover information such as the changing trends of user interests and potential market demands, thereby promoting product and service innovation; at the same time, these analysis results can also provide valuable references for fields such as academic research and social public opinion analysis.
[0153] Figure 3 is a schematic structural diagram of a monitoring and analysis system shown according to an exemplary embodiment of the present application. As Figure 3 shown, the monitoring and analysis system 300 provided in this embodiment includes:
[0154] An acquisition module 310, configured to acquire comment data on a public network platform within a preset time period range to generate a comment data set, where the comment data includes a comment time, comment information, and user information;
[0155] A processing module 320, configured to generate a user node set according to the user information of each comment data in the comment data set, where each user node in the user node set is used to represent the corresponding user information in the comment data set;
[0156] The processing module 320 is further configured to generate a set of feature vector edges between each user node in the user node set according to each comment data in the comment data set, where each feature edge in the set of feature vector edges is used to represent the feature relationship between any two comment data in the comment data set;
[0157] The processing module 320 is further configured to generate a comment data monitoring graph corresponding to the comment data according to the user node set and the set of feature vector edges;
[0158] The processing module 320 is further configured to determine target comment data from the comment data set according to the comment data monitoring graph.
[0159] Optionally, the processing module 320 is specifically configured to:
[0160] Determine a comment time difference and user association information according to a first comment data and a second comment data in the comment data set, where the comment time difference is determined according to a first comment time in the first comment data and a second comment time in the second comment data, and the user association information is determined according to a first user information in the first comment data and a second user information in the second comment data;
[0161] If the comment time difference meets a preset comment time difference condition and the user association information meets a preset user association relationship condition, an initial feature vector edge is generated between a first user node corresponding to the first comment data and a second user node corresponding to the second comment data, where the vector direction of the initial feature vector edge is from the user node corresponding to the comment data with an earlier comment time among the first comment data and the second comment data to the other user node;
[0162] If the comment time difference does not meet the preset comment time difference condition and / or the user association information does not meet the preset user association relationship condition, no initial feature vector edge is generated between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data;
[0163] After generating the initial feature vector edge between the first user node and the second user node, determine the feature weight value of the initial feature vector edge according to the first comment information in the first comment data and the second comment information in the second comment data, so as to generate a corresponding feature vector edge according to the initial feature vector edge and the feature weight value.
[0164] Optionally, the processing module 320 is specifically configured to:
[0165] Extract the keywords of the first comment information and the second comment information respectively to form a corresponding first keyword set and a second keyword set;
[0166] Generate a keyword similarity matrix according to the first keyword set and the second keyword set;
[0167] Determine a corresponding first keyword weight set according to the first keyword set, and determine a corresponding second keyword weight set according to the second keyword set;
[0168] Determine the feature weight value according to the first keyword weight set, the second keyword weight set, and the keyword similarity matrix.
[0169] Optionally, the processing module 320 is specifically configured to:
[0170] Determine the first keyword weight set according to the word frequency, inverse document frequency, and the number of occurrences in the comment dataset of each keyword in the first keyword set;
[0171] Determine the second keyword weight set according to the word frequency, inverse document frequency, and the number of occurrences in the comment dataset of each keyword in the second keyword set.
[0172] Optionally, the processing module 320 is specifically configured to:
[0173] Determine the comment time difference according to the time difference between the first comment time and the second comment time;
[0174] Determine the user association feature value according to the first user information and the second user information;
[0175] Correspondingly, when the comment time difference is greater than the preset comment time difference threshold and the user association feature value is greater than the preset user association feature threshold, it is determined that the comment time difference satisfies the preset comment time difference condition and the user association information satisfies the preset user association relationship condition.
[0176] Optionally, the processing module 320 is specifically configured to:
[0177] Obtain a plurality of comment propagation paths from the comment data monitoring graph. Among them, the comment propagation path includes a target user node sequence and target feature vector edges for connecting two adjacent target user nodes in the target user node sequence. The target feature vector edges all point from the previous target user node in the target user node sequence to the next target user node;
[0178] Determine that the comment data corresponding to the target user nodes on the comment propagation path is the target comment data.
[0179] Figure 4 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present application. As Figure 4 shown, an electronic device 400 provided in this embodiment includes: a processor 401 and a memory 402; wherein:
[0180] The memory 402 is used to store a computer program, and this memory can also be a flash (flash memory).
[0181] The processor 401 is used to execute the execution instructions stored in the memory to implement each step in the above method. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0182] Optionally, the memory 402 can be either independent or integrated with the processor 401.
[0183] When the memory 402 is a device independent of the processor 401, the electronic device 400 may further include:
[0184] A bus 403 for connecting the memory 402 and the processor 401.
[0185] This embodiment also provides a readable storage medium. A computer program is stored in the readable storage medium. When at least one processor of the electronic device executes this computer program, the electronic device executes the methods provided by the above various embodiments.
[0186] This embodiment also provides a program product. The program product includes a computer program, and this computer program is stored in a readable storage medium. At least one processor of the electronic device can read this computer program from the readable storage medium, and at least one processor executes this computer program so that the electronic device implements the methods provided by the above various embodiments.
[0187] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.
[0188] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A processing method for a monitoring and analysis system, characterized in that: include: Acquire comment data on a public network platform within a preset time period to generate a comment data set, wherein the comment data includes comment time, comment information, and user information; Generate a user node set according to the user information of each comment data in the comment data set, wherein each user node in the user node set is used to represent the corresponding user information in the comment data set; Generating a feature vector edge set between each user node in the user node set according to each comment data in the comment data set, wherein each feature edge in the feature vector edge set is used to characterize a feature relationship between any two comment data in the comment data set; Generate a comment data monitoring graph corresponding to the comment data according to the user node set and the feature vector edge set; Acquire multiple comment propagation paths from the comment data monitoring graph, wherein the comment propagation paths are used to characterize the diffusion process and direction of the comment data in the user network of the public network platform, and the comment propagation paths include a target user node sequence and a target feature vector edge for connecting two adjacent target user nodes in the target user node sequence; Merge the same target user nodes on multiple comment propagation paths to determine, based on the generated comment propagation network: the original comment user node corresponding to the starting user node of each comment propagation path, the key user node for propagation whose vector direction points to other user nodes and whose number of characteristic vector edges is greater than a preset number threshold, and the active user node for propagation whose vector direction points to other user nodes and at the same time there are other user nodes pointing to the node; Target comment data is determined from the comment data set according to the comment data monitoring map.
2. The processing method for monitoring and analyzing system according to claim 1, characterized in that: The step of generating a feature vector edge set between each user node in the user node set according to each comment data in the comment data set includes: Determining a comment time difference and user association information according to the first comment data and the second comment data in the comment data set, wherein the comment time difference is determined according to the first comment time in the first comment data and the second comment time in the second comment data, and the user association information is determined according to the first user information in the first comment data and the second user information in the second comment data; If the comment time difference satisfies a preset comment time difference condition and the user association information satisfies a preset user association relationship condition, an initial feature vector edge is generated between a first user node corresponding to the first comment data and a second user node corresponding to the second comment data, wherein the vector direction of the initial feature vector edge is from the user node corresponding to the comment data with an earlier comment time in the first comment data and the second comment data to another user node; The feature weight value of the initial feature vector edge is determined according to the first comment information in the first comment data and the second comment information in the second comment data, so as to generate a corresponding feature vector edge according to the initial feature vector edge and the feature weight value.
3. The processing method for monitoring and analyzing system according to claim 2, characterized in that: If the comment time difference does not satisfy the preset comment time difference condition and / or the user association information does not satisfy the preset user association relationship condition, no initial feature vector edge is generated between the first user node corresponding to the first comment data and the second user node corresponding to the second comment data.
4. The processing method for monitoring and analyzing system according to claim 3, characterized in that: The determining the feature weight value of the initial feature vector edge according to the first comment information in the first comment data and the second comment information in the second comment data includes: Respectively extracting keywords from the first comment information and the second comment information to form a corresponding first keyword set and a second keyword set; generating a keyword similarity matrix according to the first keyword set and the second keyword set; Determine a corresponding first keyword weight set according to the first keyword set, and determine a corresponding second keyword weight set according to the second keyword set; The feature weight value is determined according to the first keyword weight set, the second keyword weight set, and the keyword similarity matrix.
5. The processing method for monitoring and analyzing system according to claim 4, characterized in that: The step of determining a corresponding first keyword weight set according to the first keyword set, and determining a corresponding second keyword weight set according to the second keyword set includes: Determine the first keyword weight set according to the word frequency, inverse document frequency and the number of occurrences of each keyword in the first keyword set in the comment data set; The second keyword weight set is determined according to the word frequency, the inverse document frequency and the number of occurrences of each keyword in the second keyword set in the comment data set.
6. The processing method for monitoring and analyzing system according to any one of claims 2 to 5, characterized in that: The determining of the comment time difference and user association information according to the first comment data and the second comment data in the comment data set includes: Determine the comment time difference according to the time difference between the first comment time and the second comment time; Determine a user association feature value according to the first user information and the second user information; Correspondingly, when the comment time difference is greater than the preset comment time difference threshold and the user association feature value is greater than the preset user association feature threshold, it is determined that the comment time difference satisfies the preset comment time difference condition and the user association information satisfies the preset user association relationship condition.
7. The processing method for monitoring and analyzing system according to any one of claims 2 to 5, characterized in that: The target feature vector edges all point from the previous target user node in the target user node sequence to the next target user node; The determining target comment data from the comment data set according to the comment data monitoring graph includes: Determine the comment data corresponding to the target user node on the comment propagation path as the target comment data.
8. A monitoring and analysis system, characterized in that: include: An acquisition module, used to acquire comment data on a public network platform within a preset time period to generate a comment data set, wherein the comment data includes comment time, comment information, and user information; A processing module, configured to generate a user node set according to user information of each comment data in the comment data set, wherein each user node in the user node set is used to represent corresponding user information in the comment data set; The processing module is further used to generate a feature vector edge set between each user node in the user node set according to each comment data in the comment data set, wherein each feature edge in the feature vector edge set is used to characterize a feature relationship between any two comment data in the comment data set; The processing module is further used to generate a comment data monitoring graph corresponding to the comment data according to the user node set and the feature vector edge set; The processing module is further used to obtain multiple comment propagation paths from the comment data monitoring map, wherein the comment propagation path is used to characterize the diffusion process and direction of the comment data in the user network of the public network platform, and the comment propagation path includes a target user node sequence and a target feature vector edge for connecting two adjacent target user nodes in the target user node sequence; merge the same target user nodes on multiple comment propagation paths to determine, based on the generated comment propagation network: the original comment user node corresponding to the starting user node of each comment propagation path, the key user node for propagation whose vector direction points to other user nodes and the number of feature vector edges is greater than a preset number threshold, and the active user node for propagation whose vector direction points to other user nodes and at the same time there are other user nodes pointing to the node; The processing module is further used to determine target comment data from the comment data set according to the comment data monitoring map.
9. An electronic device, characterized in that: include: processor; as well as, A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Emotion instability user detection method for video conventional comments
CN112214661A
Media content trueness analysis method based on artificial intelligence
CN113158082A