A social data collection method and device

By obtaining the relational data and characteristic word groups of social accounts in social networks, generating sub-communities and expanding the main community, the problem of monitoring group events is solved, and dynamic monitoring and information collection of group events in social networks are realized.

CN114880585BActive Publication Date: 2025-09-19XIAMEN MEIYA PICO INFORMATION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210632768.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-09-19
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

In social networks, it is difficult to effectively monitor and analyze the development of group events, especially because the anonymity, interactivity and dynamic changes of participants make it impossible to accurately track and analyze the role types and participation levels of participants.

Method used

By obtaining the relationship data of social accounts and a set of characteristic word groups in the main community, sub-communities are generated, and the relationship between social accounts is determined based on the characteristic word groups. The main community is dynamically expanded to monitor group events, and data processing and analysis are performed using social data collection terminals.

Benefits of technology

It realizes dynamic monitoring of mass incidents, can track relevant social accounts and collect event information, improves the accuracy and efficiency of monitoring, and supports real-time response to mass incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114880585B_ABST
    Figure CN114880585B_ABST
Patent Text Reader

Abstract

The present invention discloses a social data collection method and device. The method obtains relational data of all first social accounts in a main community to be observed and a set of characteristic word groups corresponding to the main community. Based on the relational data, the method obtains the second social account corresponding to each first social account to generate a corresponding sub-community. The method then obtains information about the second social account in the sub-community to generate a characteristic word group. The method determines the relationship between the second social account and the main community based on the relationship between the characteristic word group and the set of characteristic word groups. In this way, the method can track the second social account related to a group event in the main community in the current time period, add the second social account to the main community for monitoring, and thus effectively associate the participants of the group event and collect the corresponding event information. The addition of the second social account to the main community also realizes the dynamic expansion of the main community, thereby realizing the dynamic monitoring of the development of the group event.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data collection, and in particular to a method and device for collecting social data. Background Art

[0002] With the rapid development of mobile communication technology, social networking platforms, due to their openness, sharing, interactivity, and diverse, convenient, and practical applications, have become an increasingly important means of reflecting public sentiment. Consequently, hot topics on these platforms are constantly emerging. However, the combination of online and mass incidents has increased the instability, ramifications, and harmfulness of these hot topics. In particular, with the recent frequent occurrence of online mass incidents, it is necessary to leverage opportunities and avoid risks, keeping abreast of the evolving trends in these incidents so that timely countermeasures can be implemented and the social ecosystem can develop in an orderly and healthy direction.

[0003] The development of online mass incidents typically follows several stages: incubation and incubation, outbreak, climax, subsidence, and aftermath. Therefore, studying the mechanisms and patterns of these incidents can help prepare for effective responses to such incidents. However, the anonymity, interactivity, low-cost operation, and equal participation of the Internet have led to various technical obstacles that need to be overcome when observing group behavior in social networks.

[0004] For example, from a manager's perspective, monitoring the behavior of every participant in a distributed social network is difficult. This is especially true when the scope of participants in a mass incident is unclear and the number of participants is large. Furthermore, social network administrators lack accurate, multi-dimensional data to support their analysis, making it difficult to determine the role played by each participant in the incident or the extent of their involvement.

[0005] Alternatively, consider group events from the perspective of participants: Because humans are social creatures, every action is influenced by the actions of others. The behavior and trustworthiness of participants in each event are influenced by the behavior and trustworthiness of other participants in the social network—this is known as homophily within the social network. For example, if the behavior of all other participants surrounding a participant in a social network is trustworthy, then that participant's behavior will tend to be trustworthy. Therefore, the characteristics of group events, the roles played by participants, and the level of participation will all change dynamically over time. This makes it difficult to effectively track and analyze the participants in group events. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a social data collection method and device that can realize dynamic monitoring of the development of group events.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0008] A social data collection method comprises the steps of:

[0009] Obtain the relationship data of all primary social accounts in the main community to be observed;

[0010] Obtaining a set of characteristic word groups corresponding to the main community;

[0011] Acquire a second social account corresponding to each of the first social accounts according to the relationship data, and generate a subcommunity corresponding to the first social account according to the second social account;

[0012] Obtaining information of all second social accounts in the sub-community;

[0013] generating a characteristic word group corresponding to each of the second social account according to the information of the second social account;

[0014] It is determined whether the characteristic word group and the characteristic word group set have an intersection. If so, the second social account corresponding to the characteristic word group is added to the main community.

[0015] In order to solve the above technical problems, another technical solution adopted by the present invention is:

[0016] A social data collection terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step in the above-mentioned social data collection method is implemented.

[0017] The beneficial effects of the present invention are as follows: by obtaining relational data of all first social accounts in a main community to be observed and a set of characteristic word groups corresponding to the main community, and obtaining the second social accounts corresponding to each first social account based on the relational data to generate a corresponding sub-community, then obtaining information about the second social accounts in the sub-community to generate a characteristic word group, and judging the relationship between the second social accounts and the main community through the relationship between the characteristic word group and the set of characteristic word groups, the second social accounts related to group events in the main community in the current time period can be tracked, and the second social accounts can be added to the main community for monitoring, thereby effectively associating participants in the group event and collecting corresponding event information; and adding the second social accounts to the main community also realizes dynamic expansion of the main community, thereby realizing dynamic monitoring of the development of group events. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flowchart of the steps of a social data collection method according to an embodiment of the present invention;

[0019] Figure 2 This is a structural diagram of a social data collection device according to an embodiment of the present invention;

[0020] Figure 3 This is another step flow chart of a social data collection method according to an embodiment of the present invention;

[0021] Figure 4 This is a flowchart of another step of a social data collection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] To illustrate the technical content, achieved objectives and effects of the present invention in detail, the following description is given in conjunction with the embodiments and accompanying drawings.

[0023] Please refer to Figure 1 , a social data collection method, comprising the steps of:

[0024] Obtain the relationship data of all primary social accounts in the main community to be observed;

[0025] Obtaining a set of characteristic word groups corresponding to the main community;

[0026] Acquire a second social account corresponding to each of the first social accounts according to the relationship data, and generate a subcommunity corresponding to the first social account according to the second social account;

[0027] Obtaining information of all second social accounts in the sub-community;

[0028] generating a characteristic word group corresponding to each of the second social account according to the information of the second social account;

[0029] It is determined whether the characteristic word group and the characteristic word group set have an intersection. If so, the second social account corresponding to the characteristic word group is added to the main community.

[0030] As can be seen from the above description, the beneficial effects of the present invention are as follows: by obtaining the relational data of all first social accounts in the main community to be observed and the set of characteristic word groups corresponding to the main community, and obtaining the second social accounts corresponding to each first social account based on the relational data to generate a corresponding sub-community, then obtaining information about the second social accounts in the sub-community to generate a characteristic word group, and judging the relationship between the second social accounts and the main community through the relationship between the characteristic word group and the set of characteristic word groups, that is, it is possible to track the second social accounts related to the group events in the main community in the current time period, and add the second social accounts to the main community for monitoring, thereby effectively associating the participants of the group events and collecting the corresponding event information; and adding the second social accounts to the main community also realizes the dynamic expansion of the main community, thereby realizing dynamic monitoring of the development of the group events.

[0031] Furthermore, obtaining the characteristic word group set corresponding to the main community includes:

[0032] Obtaining text data of all first social accounts in the main community;

[0033] A characteristic word group corresponding to each of the first social accounts is generated according to the text data, and a characteristic word group set corresponding to the main community is generated.

[0034] From the above description, it can be seen that by obtaining the text data of all the first social accounts in the main community, and generating the characteristic word group corresponding to the first social account and the characteristic word group set corresponding to the main community based on the text data, the system can obtain the characteristic words of the group events that all the first social accounts in the main community pay attention to.

[0035] Furthermore, obtaining the characteristic word group set corresponding to the main community includes:

[0036] Set the text data for pre-monitoring;

[0037] A characteristic word group set corresponding to the main community is generated according to the pre-monitored text data.

[0038] From the above description, it can be seen that by setting the pre-monitored text data, the corresponding text data to be monitored can be actively set according to the current hot group events, and then the first social account most relevant to the current hot group events and the corresponding feature word group and feature word group set can be obtained to realize dynamic monitoring of the main community.

[0039] Furthermore, generating a characteristic word group set corresponding to the main community includes:

[0040] Cleaning the text data to obtain the main text segments of the text data;

[0041] Extracting keywords from the main text segment;

[0042] Marking the keywords to obtain a keyword set;

[0043] The popularity ranking of the keyword set is calculated by weighted counting to obtain the characteristic word group set.

[0044] From the above description, it can be seen that by cleaning, extracting and labeling text data, it is possible to accurately extract keywords from the text data, and rank the popularity of the keyword set through weighted counting calculation, and further select the hottest group events in the main community to achieve accurate monitoring of hot group events.

[0045] Furthermore, acquiring a second social account corresponding to each of the first social accounts according to the relationship data, and generating a subcommunity corresponding to the first social account according to the second social account includes:

[0046] Calculating, based on the relationship data, an intimacy ranking of all social accounts associated with the first social account;

[0047] The second social account is obtained according to the intimacy ranking, and the sub-community corresponding to the first social account is generated according to the second social account.

[0048] As can be seen from the above description, by calculating the intimacy ranking of social accounts associated with the first social account, accounts with lower intimacy can be excluded, which not only improves the accuracy of tracking related accounts, but also reduces the amount of calculation and improves processing speed.

[0049] Furthermore, obtaining the relationship data of all first social accounts in the main community to be observed includes:

[0050] Generating intimate relationships between all first social accounts according to the relationship data, and generating an intimate relationship set;

[0051] A set of strongly associated accounts is obtained by iteratively calculating the close relationship set.

[0052] From the above description, it can be seen that by establishing a set of close relationships between all first social accounts and obtaining a set of strongly associated accounts through iterative calculation, the first social accounts that are most relevant to hot events can be grouped together, so that a corresponding relationship can be established between hot events and their corresponding first social account groups, thereby improving the convenience of monitoring group events.

[0053] Furthermore, the step of obtaining the relationship data of all first social accounts in the main community to be observed includes:

[0054] Obtaining full data of all first social accounts;

[0055] The full amount of data includes the relational data;

[0056] The full amount of data is classified and the real-time files are stored.

[0057] As can be seen from the above description, by obtaining the full amount of data of the first social account, and classifying and storing the full amount of data in real time, the reliability of the data is improved.

[0058] Furthermore, classifying the full amount of data and storing the real-time files includes:

[0059] The entire amount of data is categorized and stored in a distributed search and analysis engine.

[0060] From the above description, it can be seen that by classifying the full amount of data and storing it in a distributed search and analysis engine, it is convenient for relevant personnel to search and analyze the data, thereby improving monitoring efficiency.

[0061] Furthermore, the classifying of the full amount of data and performing distributed real-time file storage includes:

[0062] Generate an index for the fields corresponding to the full data and correspond them to the real-time file.

[0063] From the above description, it can be seen that by generating indexes for the fields corresponding to the full data and corresponding to the real-time file storage shown, the corresponding data can be searched more quickly through the corresponding index fields, thereby improving the efficiency of data search.

[0064] Please refer to Figure 2 The present invention also provides a social data collection device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step in the above-mentioned social data collection method is implemented.

[0065] The above-mentioned social data collection method and device of the present invention are applicable to data collection of various online communities, such as social networking sites or social software such as Weibo and forums. The following is an explanation of the specific implementation method:

[0066] Example 1

[0067] Please refer to Figure 1 and Figure 3 , a social data collection method, comprising the steps of:

[0068] S1. Obtaining full data of all first social accounts, where the full data includes text data and relational data;

[0069] The full amount of community data to be observed is obtained as basic data through data mining technology. The full amount of data includes personal information, friend information, fan information, post information, comments, likes, reposts, and other community-related information. Personal information, post information, and comments are text data, while friend information, fan information, likes, and reposts are relational data. The full amount of data is classified and stored in a distributed open source search and analysis engine such as Elasticsearch, a distributed real-time file storage. Each field of the full amount of data is indexed to make it searchable. This can be expanded to hundreds of servers and process petabytes of structured or unstructured data.

[0070] S2. Obtain a set of characteristic word groups corresponding to the main community;

[0071] Generate a characteristic word group corresponding to each of the first social accounts based on the text data, and generate a characteristic word group set corresponding to the main community; through the combined use of Python and Spark, quickly clean and normalize the text data of the first social accounts in batches; specifically, the steps include:

[0072] S201, cleaning the text data to remove stop words such as punctuation and meaningless symbols; selecting a main text segment, extracting keywords using text segmentation technology, and eliminating character-level noise, including non-Chinese characters and common Chinese stop words;

[0073] S202, using a word segmentation tool to perform part-of-speech tagging on the cleaned keywords to obtain a set of keywords with different parts of speech;

[0074] S203: Obtain a popularity ranking of the keyword set using a weighted counting statistical method, extract several keywords with the highest popularity ranking as feature words, form a feature word group set corresponding to each of the first social accounts, and combine the feature word group sets to generate a feature word group set corresponding to the primary community;

[0075] S3. Acquire a second social account corresponding to each of the first social accounts based on the relationship data, and generate a subcommunity corresponding to the first social account based on the second social account;

[0076] Specifically, calculating the intimacy ranking of all social accounts associated with the first social account based on the relationship data; obtaining the second social account based on the intimacy ranking, and generating the subcommunity corresponding to the first social account based on the second social account;

[0077] Specifically, the social account that is pre-placed in the intimacy ranking is obtained and marked as the second social account;

[0078] S4. Obtain information of all second social accounts in the sub-community;

[0079] S5. Generate a characteristic word group corresponding to each second social account based on the information of the second social account;

[0080] S6. Determine whether the characteristic word group and the characteristic word group set have an intersection. If so, add the second social account corresponding to the characteristic word group to the main community.

[0081] Please refer to Figure 3 In a specific implementation scenario: define the network environment to be observed as a main community C, C = {V, U}; wherein V = {v1, v2, ..., vi, ..., vn} represents the set of all accounts in the social network C, vi represents the i-th account; n is the total number of accounts; U = {u1, u2, ..., ui, ..., un} represents the set of characteristic word groups of the social network C; ui represents the characteristic word group of the i-th account vi in ​​the group behavior set U; define a sub-community Ci, Ci = {W, X}; wherein W = {w1, w2, ..., wi, ..., wk} represents the set of all accounts in the social network Ci, wi represents the i-th account of the sub-community Ci; k is the total number of accounts; X = {x1, x2, ..., xi, ..., xk} represents the set of characteristic word groups of the sub-community Ci; xi represents the characteristic word group of the i-th account xi in the group behavior set U;

[0082] Step S1: Obtain text data and relational data of all social accounts vi in ​​the main community C, and generate an account set V;

[0083] Step S2: Obtain the characteristic word group ui corresponding to each social account vi based on the text data corresponding to each social account vi, and merge them to generate the characteristic word group set U corresponding to the main community C;

[0084] Step S3: Obtain the social account wi corresponding to each social account vi based on the relational data, and aggregate the social accounts wi into a subcommunity Ci. Taking account v1 as an example, calculate the intimacy of account v1 in the account set V to obtain the social accounts wk corresponding to the top k participants in the intimacy ranking corresponding to account v1. Generate a corresponding subcommunity C1 for the social account wk, and obtain the corresponding account set W.

[0085] Step S4: Obtain information of all social accounts wk in the account set W corresponding to the sub-community C1;

[0086] Step S5: Generate a feature word group xk corresponding to the social account wk based on the information of the social account wk;

[0087] Step S6: Determine whether the characteristic word group xk corresponding to the social account wk intersects with the characteristic word group set U corresponding to the main community C. If so, add the social account wk to the main community C.

[0088] Similarly, perform the same calculation for each account wk in the account set W corresponding to each account vi in ​​the account set V;

[0089] In an optional implementation, intimacy calculation is further performed on the set of accounts related to the main community in the sub-community, that is, a secondary iterative calculation. The number of iterative calculations can be adaptively adjusted according to specific needs. This enables intelligent expansion of the monitoring range and all-round and in-depth detection of group behavior. At the same time, by adding analysis of participant behavior characteristics such as likes, comments, reposts, comment support ratios, and repost support ratios on the basis of text cluster analysis, group behavior can be quantitatively analyzed from multiple dimensions while achieving dynamic monitoring that follows the evolution of group behavior.

[0090] Example 2

[0091] The difference between this embodiment and the first embodiment is that the characteristic word group set of the main community is actively adjusted to monitor the community;

[0092] The step S2 includes another optional method, specifically:

[0093] Set pre-monitored text data; generate a characteristic word group set corresponding to the main community based on the pre-monitored text data; if the monitoring personnel set keywords in a customized manner, or provide batch text data such as event super topics; that is, after obtaining the text data of the batch event super topics, generate the corresponding characteristic word group set Y from the text data according to the above steps S201-S203; if the keywords are set in a customized manner, perform synonym conversion and other processing based on the keywords to expand the number of accounts covered by the keywords, and generate the corresponding characteristic word group set Y;

[0094] In another optional embodiment, the characteristic word group set Y can be compared with the characteristic word group set U in the main community C to determine whether the characteristic word group set ui corresponding to each social account vi in ​​the main community C intersects with the characteristic word group set Y. If so, it indicates that there is an association between the social account vi and the pre-monitored group event. The corresponding account vi is then added to the set of strongly associated accounts for monitoring, and corresponding iterative calculation steps are performed.

[0095] Step S3 further includes: generating, based on the relationship data, close relationships between all first social accounts and generating a close relationship set; iteratively calculating based on the close relationship set to obtain a set of strongly associated accounts; and calculating the closeness based on a preset indicator system and indicator calculation method based on the account's friend relationships, follower relationships, and the number of likes, comments, reposts, comment support ratio, and repost support ratio of related posts.

[0096] Define the network environment to be observed as the main community C, C = {V, E, U}; E = {eij|i = 1, 2, ..., n; j = 1, 2, ..., n} represents the set of close relationships between any two accounts; eij represents the close relationship between the i-th account vi and the j-th account vj; if there is a close relationship between the i-th account vi and the j-th account vj, then eij = 1; otherwise, eij = 0;

[0097] Specifically, the corresponding relationship between each social account vi and another social account vj in the main community C is obtained based on relational data, and an intimate relationship set E is generated; the social accounts vi in ​​the main community C are sorted by intimacy based on the intimate relationship set E, and the social accounts after the preset order are excluded from the main community C; by dynamically adjusting the characteristic word group set U of the main community and the relevant accounts in the main community, the community environment can be monitored in a targeted manner, ensuring a strong correlation between participants and group events in the observation area, and providing effective information support for relevant departments to understand the dynamics of specific groups of people or group events in the social network.

[0098] Example 3

[0099] Please refer to Figure 2 A social data collection terminal includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, each step in a social data collection method as described in any one of Embodiment 1 or Embodiment 2 is implemented.

[0100] In summary, the present invention provides a social data collection method and apparatus. This method obtains relational data and text data from all first social accounts in a main community to be observed, processes the text data accordingly, or generates a characteristic word group set using pre-monitored text data. Furthermore, based on the relational data, a second social account corresponding to each first social account is obtained to generate a corresponding subcommunity. Information about the second social accounts in the subcommunity is then obtained to generate a characteristic word group. The relationship between the characteristic word group and the characteristic word group set is used to determine the relationship between the second social account and the main community. This method enables tracking of second social accounts associated with group events in the main community during the current time period, and adds the second social account to the main community for monitoring. Furthermore, by performing multiple iterative calculations on the second social account, participants in the group event can be effectively associated and corresponding event information collected. The second social account and its iteratively calculated accounts are added to the main community, achieving dynamic expansion of the main community. Furthermore, by dynamically adjusting the characteristic word group set U of the main community, the community environment is monitored in a targeted manner, thereby ensuring a strong correlation between participants and group events within the observed area, providing effective information support for relevant departments to understand the dynamics of specific groups of people or group events in social networks.

[0101] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A social data collection method, characterized in that: Including steps: Obtaining full data of all first social accounts in the main community to be observed, wherein the full data includes text data and relational data; Obtaining a characteristic word group set corresponding to the main community, specifically comprising: obtaining text data of all first social accounts in the main community; generating a characteristic word group corresponding to each first social account based on the text data, and generating a characteristic word group set corresponding to the main community, and using a combination of Python and Spark to achieve rapid batch cleaning and normalization of the text data of the first social accounts; Acquiring a second social account corresponding to each of the first social accounts based on the relationship data, and generating a sub-community corresponding to the first social account based on the second social account, specifically including: acquiring a corresponding relationship between each social account and another social account in the main community based on the relationship data, and generating a close relationship set, sorting the social accounts in the main community by closeness based on the close relationship set, and excluding social accounts below a preset order from the main community; Obtaining information of all second social accounts in the sub-community; generating a characteristic word group corresponding to each of the second social account according to the information of the second social account; It is determined whether the characteristic word group and the characteristic word group set have an intersection. If so, the second social account corresponding to the characteristic word group is added to the main community.

2. A social data collection method according to claim 1, characterized in that: The acquiring of the characteristic word group set corresponding to the main community includes: Set the text data for pre-monitoring; A characteristic word group set corresponding to the main community is generated according to the pre-monitored text data.

3. A social data collection method according to claim 1 or 2, characterized in that: Generating a characteristic word group set corresponding to the main community includes: Cleaning the text data to obtain the main text segments of the text data; Extracting keywords from the main text segment; Marking the keywords to obtain a keyword set; The popularity ranking of the keyword set is calculated by weighted counting to obtain the characteristic word group set.

4. A social data collection method according to claim 1, characterized in that: The step of obtaining the relationship data of all first social accounts in the main community to be observed includes: Generating intimate relationships between all first social accounts according to the relationship data, and generating an intimate relationship set; A set of strongly associated accounts is obtained by iteratively calculating the close relationship set.

5. A social data collection method according to claim 1, characterized in that: The step of obtaining the relationship data of all first social accounts in the main community to be observed includes: Obtaining full data of all first social accounts; The full amount of data includes the relational data; The full amount of data is classified and the real-time files are stored.

6. A social data collection method according to claim 5, characterized in that: Classifying the full amount of data and storing real-time files includes: The entire amount of data is categorized and stored in a distributed search and analysis engine.

7. A social data collection method according to claim 5, characterized in that: The classifying of the entire amount of data and performing distributed real-time file storage includes: Generate an index for the fields corresponding to the full data and correspond them to the real-time file.

8. A social data collection terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, each step of the social data collection method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Guide method for group behavior in social network

    CN104573038A