Data fusion method and system based on social network platform
By obtaining multi-dimensional data on the social network platform, cleaning and redundancy removal, extracting spatiotemporal attributes, building a social user and content node set, and performing relationship mining and data fusion, the problems of lack of correlation and redundancy in data fusion of social network platform are solved, and efficient and accurate user behavior analysis and personalized services are achieved.
Patent Information
- Application Number
- CN202510529471.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing multi-dimensional data of social network platforms, the prior art faces the problems of lack of correlation between data dimensions and data redundancy, resulting in inadequate data fusion quality and efficiency.
By obtaining multi-dimensional data from social network platforms, performing data cleaning and redundant removal, extracting geographical and spatiotemporal information for spatial and temporal attribute description, building a set of social users and content nodes, conducting social relationship mining and data fusion, obtaining social association edge weights, and realizing user social knowledge fusion.
It improves the neatness and analysis efficiency of the data set, accurately captures user behavior, builds accurate user portraits, reveals social interaction patterns, improves the quality and efficiency of data fusion, and supports personalized recommendations and social relationship analysis.
Smart Images

Figure CN120045796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a data fusion method and system based on a social network platform. Background Art
[0002] User-generated content (UGC) has become an important data source for social network platforms. Social network platforms not only provide users with a space for communication and interaction, but also gather a large amount of personal information, behavioral data, and social interaction data. Especially for social network platforms such as Twitter, Facebook, Instagram, and WeChat, they generate massive multi-dimensional data such as text, pictures, videos, and locations every day. How to extract valuable information from these huge and complex data has become an important topic in current research and applications. At the same time, the data on social network platforms is diverse, dynamic, and highly correlated. A single data source can no longer provide comprehensive insights. For example, text data can reflect users' emotions, opinions, and intentions, picture and video data contain more intuitive information, while geographical location information can reveal users' activity ranges and the topological structure of the social network. However, traditional data processing methods often face problems such as the lack of correlation between data dimensions and data redundancy in the process of processing the interaction and fusion of these multi-dimensional data, thus reducing the quality and efficiency of social data fusion. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a data fusion method and system based on a social network platform to solve at least one of the above technical problems.
[0004] To achieve the above object, a data fusion method based on a social network platform includes the following steps: Step S1: Obtain the corresponding multi-dimensional data published by each user from the social network platform, including text content, picture content, video content, and geographical spatio-temporal information; perform data cleaning and redundancy removal on the corresponding multi-dimensional data published by each user to obtain the corresponding multi-dimensional redundancy-removed data published by each user. Step S2: Extract spatio-temporal attribute descriptions from the geographical spatio-temporal information in the corresponding multi-dimensional redundancy-removed data published by each user to obtain spatio-temporal attribute description sub-elements corresponding to the data published by each user; classify the text content, picture content, and video content corresponding to each user's publication based on the spatio-temporal attribute description sub-elements corresponding to the data published by each user to obtain a sequence data set of the publication content corresponding to each user under the spatio-temporal attributes. Step S3: Obtain the corresponding social user entity node set and social post content entity node set through the dataset of post content sequences corresponding to each user's spatio-temporal attributes, and conduct social relationship mining and analysis based on the social user entity node set and social post content entity node set to obtain the social association relationships between each social user node and each social content node; Step S4: Obtain the social association edge weights between each social user node and each social content node through the social user entity node set and social post content entity node set; based on the social association edge weights between each social user node and each social content node and combined with the corresponding social association relationships, conduct user social data fusion on the social user entity node set and social post content entity node set to obtain the user social knowledge fusion result corresponding to the social network platform.
[0005] Further, Step S1 includes the following steps: Step S11: Obtain the corresponding text content published by each user from the social network platform; Step S12: Obtain the corresponding picture content published by each user from the social network platform; Step S13: Obtain the corresponding video content published by each user from the social network platform; Step S14: Obtain the corresponding geographical spatio-temporal information published by each user from the social network platform, including the publication time point and publication location corresponding to each user's published content; Step S15: Clean and standardize the multi-dimensional data corresponding to each user's publication to obtain the multi-dimensional standardized data corresponding to each user's publication; remove the publication dynamic redundancy from the multi-dimensional standardized data corresponding to each user's publication to obtain the multi-dimensional redundancy-removed data corresponding to each user's publication.
[0006] Further, the removal of publication dynamic redundancy from the multi-dimensional standardized data corresponding to each user's publication in Step S15 includes the following steps: Based on the geographical spatio-temporal information in the multi-dimensional standardized data corresponding to each user's publication, conduct spatio-temporal matching and division on the corresponding text content, picture content, and video content to obtain the text sub-content, picture sub-content, and video sub-content corresponding to each user's publication at each publication time and publication location; Divide the publication time points and publication locations that are close and the publication content patterns that are similar into the same publication spatio-temporal similarity window, and conduct similar window dynamic classification on the text sub-content, picture sub-content, and video sub-content corresponding to each user's publication at each publication time and publication location based on the same publication spatio-temporal similarity window to obtain the set of publication content groups corresponding to each user's publication under the same spatio-temporal similarity window; Perform text semantic parsing and analysis on the text sub - content within the corresponding set of published content groups for each user under the same spatio - temporal similarity window, to obtain the text content semantic parsing data for each user under the same spatio - temporal similarity window, including the emotional tendency of the text word content and the discourse field of the text word content; Use a pre - trained deep convolutional neural network model to perform visual semantic recognition and analysis on the picture sub - content and video sub - content within the corresponding set of published content groups for each user under the same spatio - temporal similarity window, to obtain the visual content semantic feature data for each user under the same spatio - temporal similarity window, including content scene, object, and human action content features; Based on the text content semantic parsing data and visual content semantic feature data for each user under the same spatio - temporal similarity window, perform dynamic redundancy removal on the corresponding set of published content groups for each user under the same spatio - temporal similarity window, to obtain multi - dimensional redundancy removal data corresponding to each user's publication.
[0007] Furthermore, the performing of dynamic redundancy removal on the corresponding set of published content groups for each user under the same spatio - temporal similarity window based on the text content semantic parsing data and visual content semantic feature data for each user under the same spatio - temporal similarity window includes the following steps: Calculate the content information entropy for the corresponding set of published content groups for each user under the same spatio - temporal similarity window, to obtain the published content information entropy for each user under the same spatio - temporal similarity window; Compare and judge the published content information entropy for each user under the same spatio - temporal similarity window according to a preset information entropy threshold. If the published content information entropy is greater than or equal to the preset information entropy threshold, it indicates that the content information redundancy degree of the corresponding set of published content groups is low, and then determine the corresponding set of published content groups under its spatio - temporal similarity window as a non - redundant published content group set; if the published content information entropy is less than the preset information entropy threshold, it indicates that the content information redundancy degree of the corresponding set of published content groups is high, and then determine the corresponding set of published content groups under its spatio - temporal similarity window as a redundant published content group set; Perform duplicate removal processing on the determined redundant published content group set to obtain a redundant published content duplicate removal group set; Obtain the corresponding spatio - temporal similarity window through the non - redundant published content group set to get a non - redundant spatio - temporal similarity window; based on the non - redundant spatio - temporal similarity window, perform synchronous screening processing on the corresponding text content semantic parsing data and visual content semantic feature data to obtain the text content semantic data and visual content semantic data corresponding to the non - redundant spatio - temporal similarity window; Based on the semantic data of the corresponding text content and the semantic data of the visual content in the non-redundant spatio-temporal similarity window, the corresponding non-redundant published content group set is processed for interaction preference cancellation to obtain a non-redundant published content interaction cancellation group set; the redundant published content repeated elimination group set and the non-redundant published content interaction cancellation group set are merged in the order of publication content to obtain multi-dimensional redundancy removal data corresponding to each user's publication.
[0008] Further, the interaction preference cancellation processing of the corresponding non-redundant published content group set based on the semantic data of the corresponding text content and the semantic data of the visual content in the non-redundant spatio-temporal similarity window includes the following steps: Obtain the corresponding number of likes and comments of the non-redundant published content through the non-redundant published content group set corresponding to the non-redundant spatio-temporal similarity window; Based on the number of likes and comments of the non-redundant published content, perform a social interaction heat analysis on the corresponding non-redundant published content group set to obtain the social interaction heat of the user's publication corresponding to the non-redundant spatio-temporal similarity window; Based on the social interaction heat of the user's publication corresponding to the non-redundant spatio-temporal similarity window, use the social interaction preference calculation formula of the published content to perform an interaction preference evaluation calculation on the semantic data of the corresponding text content and the semantic data of the visual content in the non-redundant spatio-temporal similarity window to obtain the interaction preference score of the user's publication corresponding to the non-redundant spatio-temporal similarity window; Based on the interaction preference score of the user's publication corresponding to the non-redundant spatio-temporal similarity window, perform interaction preference cancellation processing on the corresponding non-redundant published content group set to obtain a non-redundant published content interaction cancellation group set.
[0009] Further, the social interaction preference calculation formula of the published content is specifically: ; ; In the formula, P is the interaction preference score of the user's publication corresponding to the non-redundant spatio-temporal similarity window T, τ is the window time variable parameter, S t (τ) is the degree of attraction of the text emotional tendency to the user at time τ, ω 1 is the text emotional tendency contribution coefficient, D t (τ) is the degree of attraction of the discussion field to the user at time τ, ω 2 is the discussion field contribution coefficient, S v (τ) is the degree of attraction of the visual scene to the user at time τ, ω 3 is the visual scene contribution coefficient, O v (τ) is the degree of attraction of the visual object to the user at time τ, ω 4is the contribution coefficient of the visual object, A v (\tau) is the degree of attraction of the visual person's action to the user at time \(\tau\), \(\omega\) 5 is the contribution coefficient of the visual person's action. \(R(T)\) is the corresponding social interaction popularity of the user's post under the non-redundant spatio-temporal similarity window \(T\), \(a\) is the number of likes for the post content, \(b\) is the number of comments on the post content, \(\varepsilon\) is the influence coefficient of the post content, \(W(T)\) is the user's post time decay factor corresponding to the non-redundant spatio-temporal similarity window \(T\), and \(\eta\) is the correction coefficient of the user's post interaction preference score.
[0010] Further, step S2 includes the following steps: Step S21: Extract the time entity and location entity from the geographical spatio-temporal information in the multi-dimensional redundancy-removed data corresponding to each user's post, to obtain the post time entity and post location entity corresponding to each user's post; Step S22: Conduct entity attribute mining and analysis on the post time entity and post location entity corresponding to each user's post, to obtain the attribute description relationship between the corresponding time entity and location entity of each user's post; Step S23: Based on the attribute description relationship between the corresponding time entity and location entity of each user's post, conduct spatio-temporal attribute description extraction between the corresponding post time entity and post location entity, to obtain the spatio-temporal attribute description subset corresponding to each user's post; Step S24: Based on the spatio-temporal attribute description subset corresponding to each user's post, conduct spatio-temporal attribute classification on the text content, picture content, and video content corresponding to each user's post, to obtain the post content sequence data set corresponding to each user under the spatio-temporal attribute.
[0011] Further, step S3 includes the following steps: Step S31: Take the social user entity and social post content entity in the post content sequence data set corresponding to each user's spatio-temporal attribute as nodes, to obtain the corresponding social user entity node set and social post content entity node set; Step S32: Conduct post content association matching analysis between the corresponding social user nodes in the social user entity node set and the corresponding social content nodes in the social post content entity node set, to obtain the direct association relationship of post content matching between the social user nodes and social content nodes; Step S33: Conduct social interaction relationship mining and analysis between the corresponding social user nodes in the social user entity node set and the corresponding social content nodes in the social post content entity node set, to focus on the social interaction association behavior of the user on the social network platform for the post content corresponding to likes, comments, and forwards, to obtain the indirect association relationship of post content social between the social user nodes and social content nodes; Step S34: Merge the direct association relationship of the published content and the indirect social association relationship of the published content between the social user nodes and the social content nodes to obtain the social association relationship between each social user node and each social content node.
[0012] Further, step S4 includes the following steps: Step S41: Set corresponding social interaction time windows for each social user node in the social user entity node set and each social content node in the social published content entity node set, and count the number of interaction behaviors between each social user node and each social content node according to the social interaction time window, including the number of likes, comments, and forwards; construct an interaction frequency matrix according to the number of interaction behaviors between each social user node and each social content node to generate a two-dimensional matrix of social node interaction frequencies; Step S42: Based on the two-dimensional matrix of social node interaction frequencies, perform an in-depth analysis of the social association relationship between each social user node and each social content node to obtain the social interaction depth between each social user node and each social content node; Step S43: Obtain the social content flow rate and social stagnation time nodes between each social user node and each social content node, and based on the social content flow rate and social stagnation time nodes, perform an association time decay evaluation on the social association relationship between each social user node and each social content node to obtain the social association time decay degree between each social user node and each social content node; Step S44: Assign association weights according to the social interaction depth and social association time decay degree between each social user node and each social content node to obtain the social association edge weights between each social user node and each social content node; Step S45: Based on the social association edge weights between each social user node and each social content node and combined with the corresponding social association relationship, perform user social data fusion on the user published content data corresponding to the social user entity node set and the social published content entity node set to obtain the user social knowledge fusion result corresponding to the social network platform.
[0013] Further, the present invention also provides a data fusion system based on a social network platform for executing the data fusion method based on a social network platform as described above. The data fusion system based on a social network platform includes: A social network platform data processing module, which is used to obtain multi-dimensional data corresponding to each user's post from the social network platform, including text content, picture content, video content, and geographical spatio-temporal information; perform data cleaning and redundancy removal on the multi-dimensional data corresponding to each user's post to obtain multi-dimensional redundancy-removed data corresponding to each user's post; A user social data spatio-temporal classification module, which is used to extract spatio-temporal attribute descriptions from the geographical spatio-temporal information in the multi-dimensional redundancy-removed data corresponding to each user's post to obtain spatio-temporal attribute description sub-items corresponding to each user's post; classify the text content, picture content, and video content corresponding to each user's post based on the spatio-temporal attribute description sub-items corresponding to each user's post, so as to obtain a dataset of post content sequences corresponding to each user under spatio-temporal attributes; A user entity and content entity relationship mining module, which is used to obtain a corresponding set of social user entity nodes and a set of social post content entity nodes through the dataset of post content sequences corresponding to each user under spatio-temporal attributes, and perform social relationship mining analysis based on the set of social user entity nodes and the set of social post content entity nodes, so as to obtain the social association relationship between each social user node and each social content node; A user social data association and fusion module, which is used to obtain the social association edge weights between each social user node and each social content node through the set of social user entity nodes and the set of social post content entity nodes; perform user social data fusion on the set of social user entity nodes and the set of social post content entity nodes based on the social association edge weights between each social user node and each social content node and in combination with the corresponding social association relationship to obtain the user social knowledge fusion result corresponding to the social network platform.
[0014] Advantages of the present invention: 1. Compared with the prior art, the beneficial effect of the data fusion method based on the social network platform proposed by the present invention lies in obtaining multi-dimensional data published by each user from the social network platform, including text, picture, video content, and geographical spatio-temporal information, which is the basis of the entire social network analysis process. The key role of obtaining these multi-dimensional data is that it can comprehensively capture users' online behaviors, ensuring accurate modeling of user activities. The text content reflects users' emotional attitudes, interest tendencies, and interaction content, while the picture and video content provide rich visual information, reflecting users' lifestyles, activity scenarios, etc. The geographical spatio-temporal information can help accurately locate the time and place where users' activities occur, providing a spatial dimension for analyzing users' social behavior patterns. At the same time, by cleaning and removing redundancy from the multi-dimensional data, irrelevant information, duplicate data, outliers, etc. can be eliminated, effectively reducing noise and ensuring the cleanliness and efficiency of the data set. This process avoids misleading conclusions caused by data deviation or redundancy, and formats and standardizes the data, enabling data from different sources to participate in subsequent analysis consistently. Redundancy removal can also greatly reduce the computational complexity and improve the efficiency of subsequent analysis and mining, thereby ensuring that larger-scale data sets can be processed in a shorter time, thus solving the problem of data fusion redundancy on the social network platform and making the analysis results more timely and reliable. Secondly, by deeply extracting the geographical spatio-temporal information, a more explicit time and space background can be given to each user's behavior. For example, geographical information can not only indicate the location of users' activities but also reveal users' behavior habits and preferences at different time points. The extraction of spatio-temporal attributes helps analyze users' behaviors from the time dimension (such as users' active time periods, frequent activity time points) and the space dimension (such as locations and regions where users often act), and classifies the text content, picture content, and video content corresponding to each user's post according to the spatio-temporal attribute descriptors corresponding to each user's post, thereby forming a more accurate user behavior portrait. The classified data set can reveal users' interests and behavior patterns under specific spatio-temporal conditions, providing more guiding data support for further social relationship analysis. This classification based on spatio-temporal attributes can help analyze users' preferences and interaction methods in different situations, making the subsequent social relationship analysis more detailed, which can help better understand the user interaction patterns on the social network platform and discover group behavior characteristics under certain specific spatio-temporal conditions.Then, by constructing a set of social user entity nodes and a set of social post content entity nodes, and conducting social relationship mining and analysis, it marks the in-depth development of social network analysis. The core of this step is to construct a more complex social relationship network through the association between users and content, revealing the internal connections behind social behaviors. By establishing connections between each user and content node, interaction patterns, interest correlations, etc. between users and content can be discovered, and then the structure of the social network can be depicted. The process of social relationship mining can help the platform discover users' social circles, potential influential groups, and the dissemination paths of content. For example, the posts of certain users receive the attention of multiple social nodes, and this "influence" can be reflected through relationship mining. At the same time, by analyzing the social relationship network, it can also help identify the dissemination chain of information, reveal the interaction patterns between users and the dissemination dynamics of the social network. It can fundamentally grasp the interaction rules between user behaviors and posted content, thus better solving the problem of the lack of data dimension correlation between social network platforms. Finally, by obtaining the social association edge weights between each social user node and each social content node and performing data fusion, it is the key to constructing a deep social knowledge network. The edge weights of social relationships not only reflect the interaction intensity between users and content but also the degree of attention and feedback frequency of users to certain content. This process is actually to further construct a knowledge-driven social network through multi-level and multi-dimensional data fusion. By integrating these social data, more accurate user portraits and content feature models can be established, thereby providing services such as personalized content recommendations, precise advertising placement, and social relationship analysis for social platforms. The platform can extract valuable information from a large amount of user interaction data, thus improving the quality and efficiency of social data fusion.
[0015] 2. The data fusion system based on a social network platform proposed by the present invention is generally composed of a social network platform data processing module, a spatio-temporal classification module for user social data, a relationship mining module for user entity and content entity, and a correlation fusion module for user social data. It can implement any data fusion method based on a social network platform described in the present invention, and is used to realize the data fusion method based on a social network platform through the operation between computer programs running on each module. The internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower investment, and can quickly and effectively provide a more accurate and efficient data fusion process based on a social network platform, thus simplifying the operation process of the data fusion system based on a social network platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings: Figure 1Schematic diagram of the step process of the data fusion method based on the social network platform of the present invention; Figure 2 is Figure 1 detailed schematic diagram of step S1 in Specific embodiments
[0017] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art within the scope of the present invention without creative work belong to the scope of protection of the present invention.
[0018] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0019] It should be understood that although terms such as "first" and "second" may be used here to describe each unit, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed associated items.
[0020] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a data fusion method based on a social network platform, and the method includes the following steps: Step S1: Obtain the corresponding multi-dimensional data published by each user from the social network platform, including text content, picture content, video content, and geographical spatio-temporal information; perform data cleaning and redundancy removal on the corresponding multi-dimensional data published by each user to obtain the corresponding multi-dimensional redundancy-removed data published by each user; Step S2: Extract the spatio-temporal attribute descriptions from the geo-spatio-temporal information in the corresponding multi-dimensional redundancy-removed data published by each user to obtain the spatio-temporal attribute description sub-items corresponding to each user's publication; classify the text content, picture content, and video content corresponding to each user's publication based on the spatio-temporal attribute description sub-items corresponding to each user's publication to obtain the dataset of the published content sequences corresponding to each user under the spatio-temporal attributes; Step S3: Obtain the corresponding social user entity node set and social published content entity node set through the dataset of the published content sequences corresponding to each user under the spatio-temporal attributes, and perform social relationship mining and analysis based on the social user entity node set and the social published content entity node set to obtain the social association relationships between each social user node and each social content node; Step S4: Obtain the weights of the social association edges between each social user node and each social content node through the social user entity node set and the social published content entity node set; fuse the social user entity node set and the social published content entity node set based on the weights of the social association edges between each social user node and each social content node and in combination with the corresponding social association relationships to obtain the user social knowledge fusion result corresponding to the social network platform.
[0021] In the embodiment of the present invention, please refer to Figure 1 As shown, it is a schematic diagram of the step flow of the data fusion method based on the social network platform of the present invention. In this example, the data fusion method based on the social network platform includes the following steps: Step S1: Obtain the corresponding multi-dimensional data published by each user from the social network platform, including text content, picture content, video content, and geo-spatio-temporal information; perform data cleaning and redundancy removal on the corresponding multi-dimensional data published by each user to obtain the corresponding multi-dimensional redundancy-removed data published by each user; In the embodiments of the present invention, multi-dimensional data published by each user is obtained through API interfaces from major social network platforms (such as Weibo, WeChat, Facebook, etc.), including text content, picture content, video content, and geographical spatio-temporal information. The text data includes the published words, comments, and tags, etc. The picture data contains the pictures uploaded by the users, and the video data is the short videos or other forms of video content published by the users. At the same time, the geographical spatio-temporal information is obtained through the geographical positioning service provided by the social platform, recording the publishing time and geographical location of each piece of content. After obtaining the data, data cleaning techniques are used for processing according to different data sources. The text data is processed through natural language processing (NLP) tools to remove noise, correct spelling mistakes, remove invalid characters, and duplicate content; the picture and video data are processed through image processing algorithms (such as OpenCV) to identify and remove redundant pictures, and compression techniques are used to optimize the storage of images and videos. For geographical information, a deduplication algorithm is used to filter out data with similar or duplicate locations. After these processes, a high-quality multi-dimensional data set of each user can be obtained, and finally, the multi-dimensional redundancy-removed data corresponding to each user's publication is obtained.
[0022] Step S2: Extract spatio-temporal attribute descriptions from the geographical spatio-temporal information in the multi-dimensional redundancy-removed data corresponding to each user's publication to obtain spatio-temporal attribute description sub-items corresponding to each user's publication; classify the text content, picture content, and video content corresponding to each user's publication based on the spatio-temporal attribute description sub-items corresponding to each user's publication to obtain a sequence data set of the published content corresponding to each user under the spatio-temporal attribute. In the embodiments of the present invention, by extracting spatio-temporal attribute descriptions from the geographical spatio-temporal information in the obtained multi-dimensional redundancy-removed data, first, the geographical spatio-temporal information includes a timestamp and a geographical location. The time period of the user's activities (such as day, month, year) is extracted through the time information, and based on the geographical coordinates of the user's published content, combined with geographical information system (GIS) technology, the geographical features of the user's activities are extracted, such as cities, regions, specific locations (such as a scenic spot). Through a clustering algorithm (such as the K-means algorithm), the user's published content is classified according to time and space to form spatio-temporal attribute description sub-items of the user. Then, based on these spatio-temporal attribute description sub-items, the text, picture, and video content published by each user are classified according to spatio-temporal attributes. For example, if the time when the user publishes text content at a certain location coincides with the geographical coordinates, then the text is associated with the corresponding spatio-temporal attribute. In this way, a corresponding sequence data set of the published content can be formed according to the spatio-temporal attributes of each user, and finally, a sequence data set of the published content corresponding to each user under the spatio-temporal attribute is obtained through classification.
[0023] Step S3: Obtain the corresponding social user entity node set and social post content entity node set from the dataset of the post content sequences corresponding to each user's spatio-temporal attributes, and perform social relationship mining and analysis based on the social user entity node set and the social post content entity node set to obtain the social association relationships between each social user node and each social content node; In the embodiment of the present invention, by according to the spatio-temporal attributes of the user and the corresponding dataset of the post content sequences, a social user entity node set and a social post content entity node set are constructed. The social user entity node set includes each user as an independent node, and the social post content entity node set includes each piece of post content (text, picture, video) as an independent node. Then, a social relationship graph between the user and the post content is established through a social network analysis algorithm (such as PageRank or community detection algorithm). In this graph, the social user node and the social content node are connected by edges, and the weight of the edge represents the degree of association between the user and the content. For example, if a user frequently posts a certain type of content, or a certain content is frequently forwarded, commented on, or liked by the user, the edge weight between the user node and the content node will be higher. Through these social relationship analyses, the potential social network structure can be mined, and finally the social association relationships between each social user node and each social content node are obtained.
[0024] Step S4: Obtain the social association edge weights between each social user node and each social content node from the social user entity node set and the social post content entity node set; based on the social association edge weights between each social user node and each social content node and combined with the corresponding social association relationships, perform user social data fusion on the social user entity node set and the social post content entity node set to obtain the user social knowledge fusion result corresponding to the social network platform.
[0025] In an embodiment of the present invention, based on the social user entity node set and the social post content entity node set constructed in the previous step, social data fusion is performed on social user nodes and content nodes using edge weights. First, according to the association edge weights between social user nodes and social content nodes, combined with a user behavior analysis model (such as collaborative filtering algorithm or weighted graph convolutional neural network), the relationship between users and content is modeled, and by calculating the weight values of the edges, the interest degree and interaction intensity of users for a certain type of content can be measured, so as to identify users or content with greater influence in the social network. Then, based on the social association relationship, through graph data fusion technology (such as graph embedding algorithm), the social information of users and the social information of content are integrated to form a user social knowledge graph. This graph not only contains the direct relationship between users and content, but also can identify potential social associations through inference algorithms, providing more personalized recommendation services or content analysis for the platform. This fusion result will effectively improve the social platform's understanding of user behavior, optimize business processes such as recommendation engines and advertising placement, and finally obtain the user social knowledge fusion result corresponding to the social network platform.
[0026] Further, step S1 includes the following steps: Step S11: Obtain the corresponding text content published by each user from the social network platform; Step S12: Obtain the corresponding picture content published by each user from the social network platform; Step S13: Obtain the corresponding video content published by each user from the social network platform; Step S14: Obtain the corresponding geographical and temporal information published by each user from the social network platform, including the publication time point and the publication location corresponding to the content published by each user; Step S15: Perform data cleaning and standardization on the multi-dimensional data corresponding to each user's publication to obtain the multi-dimensional standardized data corresponding to each user's publication; perform removal of publication dynamic redundancy on the multi-dimensional standardized data corresponding to each user's publication to obtain the multi-dimensional redundancy-removed data corresponding to each user's publication.
[0027] As an embodiment of the present invention, referring to Figure 2 shown, for Figure 1 the detailed step flow diagram of step S1 in Step S11: Obtain the corresponding text content published by each user from the social network platform; In the embodiments of the present invention, by extracting the text content posted by users from the interfaces of social network platforms, and through API calls, the text data of each user is obtained, which usually includes posts, comments, private messages, etc. posted by users. Through precise screening conditions, the text information related to the specified time range, keywords, geographical location, etc. is extracted. During the data extraction process, different social network platforms are processed differently. For example, for the Twitter platform, the statuses / user_timeline interface of the Twitter API can be called to obtain the tweets posted by users; for Facebook, the public posts of the specified user can be obtained through the Graph API. All the text data needs to be encoded and standardized, adopting a unified character encoding method such as UTF-8 to ensure that information will not be lost due to character set incompatibility during the processing, and finally the text content corresponding to each user's post is obtained.
[0028] Step S12: Obtain the corresponding picture content posted by each user from the social network platform; In the embodiments of the present invention, the pictures uploaded by users are obtained by accessing the API interfaces of social network platforms. Taking Instagram as an example, the Instagram Graph API can be used to access all the public media information of users, including pictures, etc. through the media endpoint. When obtaining pictures, the URLs of the pictures need to be processed, downloaded to the local or directly process their links. The extracted picture data usually includes information such as the upload time of the picture, the user to whom it belongs, the picture description, the picture type, etc. The acquisition of picture data needs to ensure the permission requirements of the platform, such as user privacy settings and API access authorization. The downloaded pictures need to be format-converted into common picture formats such as JPEG or PNG for subsequent processing. To save storage space, the pictures can be compressed or scaled to remove unnecessary high-resolution data, and finally the picture content corresponding to each user's post is obtained.
[0029] Step S13: Obtain the corresponding video content posted by each user from the social network platform; In the embodiments of the present invention, when obtaining the video content published by a user, it is first necessary to call the API interface of the social network platform to obtain all the video resources published by the user. For Weibo, the Weibo Data API is used, and the videos.list method is called to obtain the video data published by a specified user, including the URL of the video, upload time, title, description, view count, etc. For WeChat, the WeChat API is used to obtain the video content through the / videos endpoint. In this process, it is necessary to extract the URL of the video content and ensure that the video format meets the target requirements, such as MP4 or WEBM format. After obtaining the video data, it is also necessary to deduplicate and standardize the video content to avoid the same video being recorded multiple times. The acquisition of video content can be filtered based on multiple dimensions such as publication time, video tags, and view count, and finally, the video content corresponding to each user's publication is obtained.
[0030] Step S14: Obtain the corresponding geographical and temporal information for each user's publication from the social network platform, including the publication time point and the publication location corresponding to each user's published content. In the embodiments of the present invention, by obtaining the geographical location and timestamp data of each user's published content, on many social platforms, when users publish content, they can attach location information and time information. For example, Twitter and Instagram allow users to attach geographical location information to each published post, and this information is usually stored in the form of longitude and latitude. When obtaining the geographical and temporal information, the geographical coordinate data in each piece of information is extracted through API calls and converted into a unified coordinate system format (such as WGS-84) for subsequent processing and analysis. At the same time, the obtained timestamp needs to be corrected according to the platform time zone and unified into UTC time. Through geographical information, the geographical distribution pattern of users' published content can be analyzed, and the timestamp helps to analyze the time pattern of users' activities. Finally, the corresponding geographical and temporal information for each user's publication is obtained, including the publication time point and the publication location corresponding to each user's published content.
[0031] Step S15: Perform data cleaning and standardization on the multi-dimensional data corresponding to each user's publication to obtain the multi-dimensional standardized data corresponding to each user's publication; remove the publication dynamic redundancy from the multi-dimensional standardized data corresponding to each user's publication to obtain the multi-dimensional redundancy-removed data corresponding to each user's publication.
[0032] In the embodiments of the present invention, by cleaning various types of previously extracted data, the goal of data cleaning is to remove useless or incorrect data items, such as missing values, format errors, duplicate data, etc. For text data, stop word removal, spelling correction, word segmentation, etc. are performed through natural language processing (NLP) techniques; for picture and video data, redundant files are cleaned, and duplicate picture and video resources are removed; for spatio-temporal information, time zones and formats are proofread to ensure that all time points conform to the standard time format (for example, the ISO 8601 standard). After cleaning, each data item is standardized. For text data, through vectorization methods such as TF-IDF, Word2Vec, etc., the text content is converted into a unified numerical format; for picture data, the image size is unified and the color is standardized; for video data, standardization can be performed by extracting key frames, unifying the video duration, etc. After this series of cleaning and standardization operations, the data becomes more standardized, and thus multi-dimensional standardized data corresponding to each user's post is obtained. At the same time, after data standardization, it enters the redundancy removal stage. The purpose of redundancy removal is to remove the duplicate or highly similar parts in the content published by the user, reduce the data volume and improve the data quality. For text data, by calculating the similarity between texts (such as using cosine similarity or Jaccard similarity), posts or comments with duplicate content are identified and removed; for pictures, image recognition technologies (such as Hash value algorithm or deep learning model) can be used to detect duplicate picture files and remove duplicate data; for video content, by comparing the extracted hash value or frame content of the video, similar videos are removed. Redundancy removal also includes removing overly frequent posts according to timestamps to avoid the same user from repeatedly publishing the same or similar content in a short period of time, and those content with low information quality and no practical value should also be excluded to ensure that each piece of data finally retained is valid and representative. Finally, multi-dimensional redundancy removal data corresponding to each user's post is obtained.
[0033] Further, the dynamic redundancy removal of the multi-dimensional standardized data corresponding to each user's post in step S15 includes the following steps: Based on the geographical spatio-temporal information in the multi-dimensional standardized data corresponding to each user's post, spatio-temporal matching division is performed on the corresponding text content, picture content, and video content, and text sub-content, picture sub-content, and video sub-content corresponding to each user's post at each post time and post location are obtained; In the embodiments of the present invention, key data related to geographical spatio-temporal information is extracted from the data published by users, including the publication time, publication location of each piece of content, and the location information of the user. These information are standardized through the integration of location services (such as GPS positioning or IP address positioning) and timestamp functions. By setting a fixed time interval and geographical area range, the published content of each user is divided spatio-temporally. Specifically, text content, picture content, and video content are marked according to their publication time and location, and a clustering algorithm (such as K-means) is used to perform spatio-temporal matching and classification on this content. For example, if a piece of text content, picture, or video is published at the same geographical location (such as a specific area of a city) and within a short period of time (such as within the same hour), they will be classified into the same spatio-temporal unit. Corresponding text sub-content, picture sub-content, and video sub-content will be formed under each spatio-temporal unit. This operation can be automatically completed by calling the Geographic Information System (GIS) module and time series analysis methods, and finally, the corresponding text sub-content, picture sub-content, and video sub-content published by each user at each publication time and location are obtained.
[0034] Preferably, publication time points, publication location points that are close, and publication content patterns that are similar are divided into the same publication spatio-temporal similarity window, and based on the same publication spatio-temporal similarity window, the corresponding text sub-content, picture sub-content, and video sub-content published by each user at each publication time and location are dynamically classified into similarity windows, and a set of publication content groups corresponding to each user published under the same spatio-temporal similarity window is obtained; In the embodiments of the present invention, by further dividing the content published by each user according to the proximity of each publication time point and geographical location, and aggregating similar spatio-temporal information together to form a "publication spatio-temporal similarity window". In this process, first, the similarity of the time and location of content publication is calculated to determine which published content belongs to the same spatio-temporal window. For time similarity, a time difference calculation method (such as Euclidean distance or cosine similarity) can be used to evaluate the similarity of publication time. For location similarity, the geographical distance between two locations (such as the Haversine formula) can be calculated to determine the similarity of locations. If the time difference between two published contents is within a predetermined threshold and the geographical locations are close (such as in the same city or adjacent areas), then these contents are classified into the same spatio-temporal similarity window. The text, picture, and video content within the same spatio-temporal similarity window will be dynamically classified together to form a set of publication content groups. This process is automatically divided using a clustering algorithm (such as DBSCAN) to ensure that content with similar characteristics can be correctly aggregated together, and finally, a set of publication content groups corresponding to each user published under the same spatio-temporal similarity window is obtained.
[0035] Preferably, text semantic parsing and analysis are performed on the text sub - contents within the corresponding set of published content groups of each user under the same spatio - temporal similarity window, and text content semantic parsing data corresponding to each user under the same spatio - temporal similarity window are obtained, including the emotional tendency of text word content and the field of discussion of text word content; In the embodiment of the present invention, through semantic parsing of the text sub - contents under each spatio - temporal similarity window, by means of natural language processing (NLP) technology, using a pre - trained language model (such as BERT or GPT) to perform sentiment analysis and topic analysis on the text content. The sentiment analysis module analyzes the keywords, sentence structures, and context relationships in the text to judge the emotional tendency of the text (such as positive, negative, or neutral). The topic analysis module, through topic modeling techniques (such as Latent Dirichlet Allocation, LDA) or deep learning methods, identifies the main field of discussion of the text (such as entertainment, sports, news, etc.). Each piece of text content will be labeled with the corresponding emotional tendency and field of discussion labels to form the semantic parsing data of the text content. If the text content involves specific emotional words or topics, weighted processing of emotional words will be carried out through a custom dictionary, and finally, text content semantic parsing data corresponding to each user under the same spatio - temporal similarity window are obtained, including the emotional tendency of text word content and the field of discussion of text word content.
[0036] Preferably, a pre - trained deep convolutional neural network model is used to perform visual semantic recognition and analysis on the picture sub - contents and video sub - contents within the corresponding set of published content groups of each user under the same spatio - temporal similarity window, and visual content semantic feature data corresponding to each user under the same spatio - temporal similarity window are obtained, including content scene, object, and human action content features; In the embodiments of the present invention, by using a deep learning model, especially a pre-trained convolutional neural network (CNN), visual semantic recognition is performed on each published picture and video. For picture content, the CNN model (such as ResNet, InceptionV3) will extract features of the scene, objects, and human actions in the picture, and identify the main content included in each picture (such as indoor and outdoor scenes, people, objects, etc.); for video content, a temporal convolutional neural network (TCN) or a model based on LSTM (long short-term memory network) is used to process the video frame sequence, and identify the dynamic scenes, human behaviors, or object movements therein. Through these deep learning technologies, semantic feature data of picture and video content can be extracted, including but not limited to feature information such as content scenes (such as beaches, city streets), objects (such as cars, mobile phones), and human actions (such as running, jumping), and finally, visual content semantic feature data corresponding to each user's publication in the same spatio-temporal similarity window is obtained, including content scene, object, and human action content features.
[0037] Preferably, based on the text content semantic parsing data and visual content semantic feature data corresponding to each user's publication in the same spatio-temporal similarity window, dynamic redundancy removal is performed on the publication content group set corresponding to each user's publication in the same spatio-temporal similarity window, so as to obtain multi-dimensional redundancy removal data corresponding to each user's publication.
[0038] In the embodiments of the present invention, by based on the semantic parsing data of text content and the semantic feature data of visual content, redundancy removal is performed on the publication content group set under each spatio-temporal similarity window. First, by comparing the semantic similarity between the text content, pictures, and video content within the same spatio-temporal window, duplicate or similar content is identified. Text redundancy removal uses similarity calculation (such as cosine similarity, Jaccard similarity) to compare the emotional tendency and discussion field of the text, and removes content with high similarity in emotional tendency or theme. For picture and video content, visual feature data is used for redundancy removal, and similarity calculation based on feature vectors (such as Euclidean distance) is used to compare the visual content of images and videos, and duplicate content with high visual content similarity is removed. During the redundancy removal process, the most representative content is retained to avoid repeated dissemination of information, thereby obtaining multi-dimensional redundancy removal data for each user's publication. This process can be comprehensively evaluated through a clustering algorithm combined with the features of text and visual content to ensure the efficiency and accuracy of redundant content removal, and finally obtain multi-dimensional redundancy removal data corresponding to each user's publication.
[0039] Further, the dynamic redundancy removal of the published content group sets corresponding to each user in the same spatio-temporal similarity window based on the semantic parsing data of the text content and the semantic feature data of the visual content corresponding to each user in the same spatio-temporal similarity window includes the following steps: Calculate the content information entropy of the published content group sets corresponding to each user in the same spatio-temporal similarity window, and obtain the published content information entropy of the published content group sets corresponding to each user in the same spatio-temporal similarity window; In the embodiment of the present invention, for the content published by each user in the same spatio-temporal similarity window, first, the information entropy of these published contents needs to be calculated. The information entropy is an index used to measure the uncertainty of information. The larger the entropy value, the higher the diversity and uncertainty of the content. In specific implementation, for the published content of each user, first, according to the limitations of the time window and geographical location, a "spatio-temporal similarity window" is determined. This window includes all the published content group sets in this spatio-temporal environment. Then, based on the content data in this window, the information entropy is calculated. For text content, methods such as TF-IDF and word frequency distribution can be used to extract content features, and then the entropy formula is used for calculation. The entropy formula is: , where p(x i ) is the probability of the occurrence of the i-th type of published content, and n is the total number of categories of the published content. Finally, the published content information entropy of the published content group sets corresponding to each user in the same spatio-temporal similarity window is obtained.
[0040] Preferably, compare and judge the published content information entropy of the published content group sets corresponding to each user in the same spatio-temporal similarity window according to a preset information entropy threshold. If the published content information entropy is greater than or equal to the preset information entropy threshold, it means that the content information redundancy degree of the corresponding published content group set is low, and then the published content group set corresponding to its spatio-temporal similarity window is determined as a non-redundant published content group set; if the published content information entropy is less than the preset information entropy threshold, it means that the content information redundancy degree of the corresponding published content group set is high, and then the published content group set corresponding to its spatio-temporal similarity window is determined as a redundant published content group set; In an embodiment of the present invention, by comparing the entropy value of the content published by each user according to a preset information entropy threshold, if the entropy value of the published content is greater than or equal to the preset entropy threshold, it indicates that the information redundancy of the content group set is relatively low; otherwise, the information redundancy is relatively high. First, it is necessary to determine this information entropy threshold. Usually, a reasonable threshold can be set through data analysis or statistical analysis methods of historical data. For example, by statistically analyzing the information entropy values of the content published on a large number of social platforms, calculating their distribution, and then selecting a quantile as the threshold. Then, for the content group set of the published content under each spatio-temporal similarity window, compare and judge the information entropy with the threshold. If the information entropy is greater than or equal to the threshold, it indicates that the content group set of the published content has high diversity and low redundancy, and confirm that this group set is a "non-redundant published content group set"; if the information entropy is less than the threshold, it indicates that there is high redundancy in this group set, and it should be determined as a "redundant published content group set".
[0041] Preferably, perform duplicate content removal processing on the determined redundant published content group set to obtain a redundant published content duplicate removal group set; In an embodiment of the present invention, for the content determined to be a redundant published content group set, implement "duplicate content removal processing". Redundant published content group sets usually have the situation of information duplication and high similarity. Therefore, it is necessary to use a certain deduplication algorithm to remove duplicate or similar content. Specifically, during implementation, first compare each piece of published content through a text similarity algorithm (such as Jaccard similarity, cosine similarity, etc.). If the similarity between two pieces of content exceeds the preset similarity threshold, it is considered that their content is repeated, and the repeated content is removed. For the deduplication processing of text content, sentence-level comparison can be used. After segmenting the text and representing it in vector form, calculate the similarity between the contents. If visual content (such as pictures, videos, etc.) is also included in the published content, image processing technology can be used, combined with a convolutional neural network (CNN) to extract the features of the pictures, compare the similarity between different pictures, and thus judge and remove duplicate visual content. After removing the duplicate content, finally obtain a redundant published content duplicate removal group set.
[0042] Preferably, obtain the corresponding spatio-temporal similarity window through the non-redundant published content group set to obtain a non-redundant spatio-temporal similarity window; based on the non-redundant spatio-temporal similarity window, perform synchronous screening processing on the corresponding text content semantic analysis data and visual content semantic feature data to obtain the corresponding text content semantic data and visual content semantic data under the non-redundant spatio-temporal similarity window; In an embodiment of the present invention, a corresponding spatio-temporal similarity window is obtained based on a non-redundant published content group set, so as to obtain a non-redundant spatio-temporal similarity window, and semantic parsing data of text and visual content is synchronously filtered through this window. The spatio-temporal similarity window is defined according to the time period and geographical location range corresponding to the non-redundant content published by the user. For each non-redundant published content group set, first, the corresponding spatio-temporal window information is extracted, and then, synchronous filtering processing is performed based on this window. In terms of text data, natural language processing (NLP) technology is used to perform semantic parsing on each piece of content, and feature information such as keywords, themes, and emotions is extracted. For visual data, image recognition technology (such as convolutional neural network in deep learning) is used to extract visual feature data of images or videos. Through synchronous filtering of the semantic features of text and visual content, content data irrelevant to the non-redundant published content is filtered out, and finally, text content semantic data and visual content semantic data corresponding to the non-redundant spatio-temporal similarity window are obtained.
[0043] Preferably, based on the text content semantic data and visual content semantic data corresponding to the non-redundant spatio-temporal similarity window, an interaction preference cancellation process is performed on the corresponding non-redundant published content group set to obtain a non-redundant published content interaction cancellation group set; the redundant published content repeated elimination group set and the non-redundant published content interaction cancellation group set are merged in terms of the time sequence of the published content to obtain multi-dimensional redundancy removal data corresponding to each user's publication.
[0044] In an embodiment of the present invention, based on the text and visual content semantic data extracted under the non-redundant spatio-temporal similarity window, an interaction preference cancellation process is performed on the non-redundant published content group set. The interaction preference cancellation process aims to reduce low-quality interaction content and improve the transmission efficiency and quality of information according to factors such as the user's interaction history and preference model. Specifically, the following methods can be used for processing. By analyzing the user's past interaction data such as likes, comments, and shares, an interaction preference model of the user is established to identify features such as the user's activity level and interest bias; according to the user's interaction records with specific content, the interaction score of each piece of content is calculated, and the content with a lower score is determined as "cancelled content"; through a filtering algorithm, the content with a lower score is filtered out from the non-redundant published content group set to obtain a non-redundant published content interaction cancellation group set. At the same time, the redundant published content repeated elimination group set and the non-redundant published content interaction cancellation group set will be merged in combination with the time sequence information to form multi-dimensional redundancy removal data finally published by the user. This process involves reorganizing multiple pieces of content in chronological order to ensure that the user's content presentation is time-sensitive and multi-dimensionally optimized, and finally, multi-dimensional redundancy removal data corresponding to each user's publication is obtained.
[0045] Further, the process of performing interaction preference cancellation processing on the corresponding non-redundant published content group set based on the semantic data of the corresponding text content and the semantic data of the visual content under the non-redundant spatio-temporal similarity window includes the following steps: Obtain the number of likes and the number of comments of the corresponding non-redundant published content from the non-redundant published content group set under the non-redundant spatio-temporal similarity window; In the embodiment of the present invention, by defining the non-redundant spatio-temporal similarity window, the non-redundant published content group set under specific time and space conditions is determined. This process is achieved through spatio-temporal analysis of the content published on the social network platform. The "spatio-temporal similarity window" is determined by comprehensive analysis of the user's geographical location and the content publishing time. Specifically, first, the timestamps and geographical location information of all user-published content are obtained, and then filtered according to the set time range and geographical range to filter out the content groups that meet the non-redundant conditions. Then, for these content groups that meet the conditions, the number of likes and comments of each content is further counted to ensure that these statistical data do not include redundant or duplicate content. Through this step, the specific social interaction data of each non-redundant published content, including the number of likes and comments, can be obtained, and finally the number of likes of the non-redundant published content and the number of comments of the non-redundant published content are obtained.
[0046] Preferably, perform a social interaction heat analysis on the corresponding non-redundant published content group set based on the number of likes of the non-redundant published content and the number of comments of the non-redundant published content to obtain the social interaction heat of the corresponding user's published content under the non-redundant spatio-temporal similarity window; In the embodiment of the present invention, the statistical calculation of the interaction heat of the corresponding user's social published content in the non-redundant published content group set is performed by combining the number of likes of the non-redundant published content and the number of comments of the non-redundant published content obtained from the previous analysis. In the specific operation, first, the number of likes and comments of each published content are added together to obtain the preliminary heat value of each content. Then, these heat values are further subdivided according to the time sequence and spatial correlation, so as to quantitatively calculate the preliminary heat value of each non-redundant published content group in a specific spatio-temporal window in combination with the corresponding influence and the decay situation of the publishing time, so as to calculate the interaction value of the user's published content, and finally obtain the social interaction heat of the corresponding user's published content under the non-redundant spatio-temporal similarity window.
[0047] Preferably, use the social interaction preference calculation formula of the published content to perform an interaction preference evaluation calculation on the semantic data of the corresponding text content and the semantic data of the visual content under the non-redundant spatio-temporal similarity window to obtain the user's published interaction preference score under the non-redundant spatio-temporal similarity window; In the embodiments of the present invention, by combining the window time variable parameter, text sentiment tendency, discussion field, visual scene, visual object, the degree of attraction of the visual human action to the user, the corresponding contribution coefficient, the social interaction heat of the user's published content, the number of likes of the published content, the number of comments of the published content, the influence coefficient of the published content, the user's publishing time decay factor, and related parameters, a suitable social interaction preference calculation formula for the published content is constructed for evaluation and calculation, so as to obtain a quantitative interaction preference score by integrating the user's interaction behaviors. This score can be calculated for different types of published content, thereby reflecting the user's interaction preferences under different content types, and finally obtaining the user's published interaction preference score corresponding to the non-redundant spatio-temporal similarity window.
[0048] Preferably, based on the user's published interaction preference score corresponding to the non-redundant spatio-temporal similarity window, the corresponding non-redundant published content group set is processed for low cancellation of interaction preference, and a non-redundant published content interaction low cancellation group set is obtained.
[0049] In the embodiments of the present invention, by comparing and judging the user's published interaction preference score obtained by previous quantification with a preset threshold, and by analyzing the user's interaction preference score, the non-redundant published content corresponding to the non-redundant spatio-temporal similarity window with a lower score is removed from the user's content stream, so as to prevent the user from seeing content that does not match their interests. Specifically, first, all published content is sorted according to the user's interaction preference score to obtain the degree of fit between each content and the user's preference, and then a threshold is set. All published content below this threshold will be marked as "low cancellation" content, that is, content with a lower interaction preference. These contents will be preferentially blocked in the user's content recommendation. To achieve this, the user's preference can be trained through a machine learning model (such as a decision tree, support vector machine, etc.) to further improve the accuracy and flexibility of the low cancellation process, and a filtered non-redundant published content group set is obtained, which contains content that highly matches the user's interaction preference, and finally a non-redundant published content interaction low cancellation group set is obtained.
[0050] Further, the social interaction preference calculation formula for the published content is specifically: ; ; In the formula, P is the user's published interaction preference score corresponding to the non-redundant spatio-temporal similarity window T, τ is the window time variable parameter, S t (τ) is the degree of attraction of the text sentiment tendency to the user at time τ, ω 1 is the text sentiment tendency contribution coefficient, D t (τ) is the degree of attraction of the discussion field to the user at time τ, ω 2 is the discussion field contribution coefficient, Sv (τ) is the degree of attraction of the visual scene to the user at time τ, ω 3 is the visual scene contribution coefficient, O v (τ) is the degree of attraction of the visual object to the user at time τ, ω 4 is the visual object contribution coefficient, A v (τ) is the degree of attraction of the visual human action to the user at time τ, ω 5 is the visual human action contribution coefficient, R(T) is the corresponding user - released social interaction heat under the non - redundant spatio - temporal similarity window T, a is the number of likes for the released content, b is the number of comments for the released content, ε is the released content influence coefficient, W(T) is the corresponding user - released time decay factor under the non - redundant spatio - temporal similarity window T, and η is the correction coefficient of the user - released interaction preference score.
[0051] The present invention obtains a social interaction preference calculation formula for released content through using a specific mathematical model and verification, which is used to evaluate and calculate the interaction preference for the semantic data of text content and visual content corresponding under the non - redundant spatio - temporal similarity window. This social interaction preference calculation formula for released content models the degree of attraction of different types of content (text, visual scene, visual object, visual human action, etc.) through multiple weight coefficients. This modeling method enables in - depth capture of the multi - dimensional preferences of users for released content. Specifically, the text sentiment tendency and discussion field can reflect the emotional resonance or topic interest of users when reading text, while visual elements focus on the attraction of visual presentation to users. Therefore, the formula can take into account the dual attraction of vision and text and provide a comprehensive user preference evaluation. The time variable and window time in the formula reflect the dynamic changes of the spatio - temporal similarity window, enabling the model to adapt to the temporal and spatial fluctuations of user behavior. For example, some content may be more attractive during a specific time period and lose attention at another time period. By modeling the weight decay in time and space, the formula can effectively evaluate the interaction preference of content within a certain time and space window. The advantage of this spatio - temporal analysis is that it can help the platform adjust the recommendation strategy according to the specific spatio - temporal background and improve the accuracy and timeliness of recommendations. The introduction of social interaction heat enables the formula to consider the direct influence of social interactions (such as the number of likes and comments). By analyzing the interaction heat of the released content, the formula can not only identify which content has a high user participation rate but also measure the dissemination potential and attraction of this content in the social network based on these interaction heats (number of likes, number of comments). The calculation formula of R(T): Among them, a and b represent the number of likes and comments respectively, ε is the content influence coefficient, and W(T) is the decay factor of the content. In this way, the formula combines the social interaction heat of the number of likes and comments with the decay factor of the content, ensuring a reasonable relationship between the heat of the content and the distance from the release time, and avoiding the excessive influence of outdated content on the user interaction preference. The user release interaction preference score is comprehensively calculated through the above formula. It not only considers the interaction heat and spatio-temporal characteristics of the content, but also adds a correction coefficient for the user preference. This correction coefficient can be adjusted according to the user's historical behavior data, personalized settings or other customized information, so that the calculation formula can reflect the unique preferences of individual users. This personalized evaluation helps the platform provide more accurate content recommendations according to the interest characteristics of each user, improving user participation and satisfaction. To sum up, the formula fully considers the user release interaction preference score P corresponding to the non-redundant spatio-temporal similarity window T, the window time variable parameter τ, the degree of attraction S t (τ) of the text sentiment tendency to the user at time τ, and the text sentiment tendency contribution coefficient ω 1 , the degree of attraction D t (τ) of the discussion field to the user at time τ, and the discussion field contribution coefficient ω 2 , the degree of attraction S v (τ) of the visual scene to the user at time τ, and the visual scene contribution coefficient ω 3 , the degree of attraction O v (τ) of the visual object to the user at time τ, and the visual object contribution coefficient ω 4 , the degree of attraction A v (τ) of the visual human action to the user at time τ, and the visual human action contribution coefficient ω 5 , the user release social interaction heat R(T) corresponding to the non-redundant spatio-temporal similarity window T, the number of likes a of the released content, the number of comments b of the released content, the content influence coefficient ε, the user release time decay factor W(T) corresponding to the non-redundant spatio-temporal similarity window T, and the correction coefficient η of the user release interaction preference score. Among them, by combining the number of likes a of the released content, the number of comments b of the released content, the content influence coefficient ε, and the user release time decay factor W(T) corresponding to the non-redundant spatio-temporal similarity window T, a functional relationship of the user release social interaction heat R(T) corresponding to the non-redundant spatio-temporal similarity window T is formed , and a functional relationship is formed according to the mutual correlation relationship between the user release interaction preference score P corresponding to the non-redundant spatio-temporal similarity window T and the above parameters , this formula can implement the interactive preference evaluation calculation process for the semantic data of text content and the semantic data of visual content corresponding to the non-redundant spatio-temporal similarity window. At the same time, the correction coefficient η of the interactive preference score published by the user can be adjusted according to the error situation that appears in the calculation process, so as to improve the accuracy and applicability of the calculation formula for the social interactive preference of the published content.
[0052] Further, step S2 includes the following steps: Step S21: Extract the time entity and location entity from the geographical spatio-temporal information in the multi-dimensional redundancy-removed data corresponding to each user's publication, and obtain the publication time entity and publication location entity corresponding to each user's publication; In the embodiment of the present invention, by extracting the time entity and location entity from the geographical spatio-temporal information in the multi-dimensional data corresponding to each user's publication, this process can be realized through the named entity recognition (NER) technology. Specifically, by training a named entity recognition model based on deep learning (such as BERT, BiLSTM-CRF, etc.), the model will identify time-related entities (such as "May 1, 2023", "5 pm", etc.) and location-related entities (such as "Beijing", "Pudong New Area, Shanghai", etc.) from the geographical spatio-temporal information published by the user. These entity information needs to be extracted from the user's published content and stored as structured data, such as the time entity (May 1, 2023) and the location entity (Beijing). The main task of this step is to accurately distinguish the time and location information in the text, extract and label them as independent entities, and finally obtain the publication time entity and publication location entity corresponding to each user's publication.
[0053] Step S22: Mine and analyze the entity attributes of the publication time entity and publication location entity corresponding to each user's publication, and obtain the attribute description relationship between the time entity and location entity corresponding to each user's publication; In the embodiments of the present invention, when mining and analyzing the entity attributes of the extracted time entities and location entities, a rule-based attribute extraction technique is combined with a machine learning algorithm. In terms of time entities, first, the time characteristics of the entities are analyzed. For example, "May 1, 2023" contains information such as year, month, and date. For location entities, the administrative division information of the location needs to be identified. For example, "Beijing" is a city and a municipality directly under the Central Government of China. By relying on the Geographic Information System (GIS) database and the time database, additional attributes can be added to each location and time entity. For example, "May 1, 2023" can be labeled as a working day, and "Beijing" can be labeled as the capital, a large city with a large population, etc. The core of this step is to extract relevant attributes for each time and location entity, and analyze the descriptive relationship between each time entity and the location entity, and finally obtain the attribute description relationship between the corresponding time entity and the location entity for each user's post.
[0054] Step S23: Extract the spatio-temporal attribute description between the corresponding published time entity and the published location entity based on the attribute description relationship between the corresponding published time entity and the published location entity for each user, so as to obtain the spatio-temporal attribute description sub for each user's post. In the embodiments of the present invention, by combining the previously extracted attributes of time and location, the spatio-temporal relationship between the two is further analyzed. For example, when analyzing the information "On May 1, 2023, the user posted in Beijing" mentioned in a published content, the specific time period attributes (such as holidays, working days, etc.) of the time entity "May 1, 2023" and the geographical attributes (such as city, economic center, etc.) of the location entity "Beijing" can be extracted. Then, based on the spatio-temporal attributes of the two, combined with the historical published data of the social network platform, a spatio-temporal relationship description is constructed. For example, it is judged that there are more posts in "Beijing" during "holidays", or the change in the number of posts at certain locations during "working days", etc., so as to provide data support for the spatio-temporal attribute description sub. Through spatio-temporal relationship analysis, the time-location attribute characteristics of each user's published content are constructed, that is, what social content the user posted at [time entity] in [location entity], and finally the spatio-temporal attribute description sub for each user's post is obtained.
[0055] Step S24: Classify the text content, picture content, and video content corresponding to each user's post based on the spatio-temporal attribute description sub for each user's post, and obtain the published content sequence data set corresponding to each user under the spatio-temporal attribute.
[0056] In the embodiments of the present invention, based on the spatio-temporal attribute descriptors released by each user, the text content, picture content, and video content released by the user are further classified according to spatio-temporal attributes. First, the spatio-temporal descriptions of each released content are matched with the spatio-temporal attributes of the user's historical releases. For example, if a user releases content related to "holidays" on "May 1, 2023", then the system will classify this content under the spatio-temporal attribute of "holidays". If the released content involves the location attribute of "Beijing", then this content will be further classified as relevant content in the "Beijing" area. For different forms of content such as text, pictures, and videos, spatio-temporal classification can be performed through multimodal learning techniques (such as image recognition models, natural language processing models, etc.). The text content can be classified by analyzing its theme, keywords, etc. The picture content extracts visual information related to spatio-temporal attributes through image recognition technology, while the video content is classified by matching time stamps and location information, generating a spatio-temporal attribute classification dataset for all the released content of each user, which contains multimodal content such as text, pictures, and videos corresponding to each spatio-temporal attribute, and finally obtaining the released content sequence dataset corresponding to each user under the spatio-temporal attributes.
[0057] Further, step S3 includes the following steps: Step S31: Taking the social user entities and social release content entities in the released content sequence dataset corresponding to each user's spatio-temporal attributes as nodes to obtain the corresponding social user entity node set and social release content entity node set; In the embodiments of the present invention, by extracting the user's spatio-temporal attributes and the released content sequence dataset, two main node sets in the social network are constructed, namely the social user entity node set and the social release content entity node set. The specific implementation method is to mark each released content through user behavior characteristics such as time stamps, geographical locations, and content categories, thereby constructing a sequence of content released by each user at different times and different locations. This process can be modeled using a graph database (such as Neo4j). Regarding each user as a social user entity node and each released content as a social release content entity node, these two node sets are connected through attributes to form a preliminary social network, and each node has clear spatio-temporal characteristics to further distinguish the release behaviors at different times and spaces, and finally obtaining the corresponding social user entity node set and social release content entity node set.
[0058] Step S32: Conducting release content association matching analysis between the corresponding social user nodes in the social user entity node set and the corresponding social content nodes in the social release content entity node set to obtain the direct release content matching association relationship between the social user nodes and the social content nodes; In the embodiments of the present invention, for the social user entity node set and the social post content entity node set, correlation matching analysis is performed through the text content, tags, and metadata of the post content. Specifically, the association relationship between social user nodes and social content nodes is determined by matching the similarity of user-posted content or the relevance of topics. For example, natural language processing (NLP) techniques are used to analyze the content posted by users, extract keywords, themes, or tags, and models such as TF-IDF (Term Frequency-Inverse Document Frequency) or BERT (Bidirectional Encoder Representations from Transformers) can be used to calculate the text similarity between each user-posted content and other content nodes, so as to construct a direct matching relationship between social user nodes and social content nodes, that is, the direct connection relationship between the corresponding social and its posted content. These matching relationships are further refined through graph algorithms. For example, the relationship strength between posted contents is calculated based on cosine similarity to obtain the direct association degree of the posted content, and finally, the direct association relationship of the posted content matching between social user nodes and social content nodes is obtained.
[0059] Step S33: Conduct mining and analysis of the social interaction relationship between the corresponding social user nodes in the social user entity node set and the corresponding social content nodes in the social post content entity node set, so as to focus on the social interaction association behavior of users on the social network platform for the posted content corresponding to the post, including likes, comments, and forwards, and obtain the indirect social association relationship of the posted content between social user nodes and social content nodes; In the embodiments of the present invention, by extracting the interaction data between users and their content according to the data interface of the platform, including the timestamps and user identifiers of behaviors such as likes, comments, and forwards, through social network analysis methods, an interaction relationship graph between users and content is constructed, and graph algorithms (such as PageRank, community discovery algorithms) are used to cluster these interaction behaviors to identify social user and content combinations with higher interaction frequencies. By calculating the weight values of the interaction behaviors between users and content, an indirect social association relationship between social user nodes and social content nodes is formed. For example, if a user forwards or comments on a piece of content multiple times, a strong association can be formed, and the content with a large number of comments will also become the indirect social focus of users. Finally, the indirect social association relationship of the posted content between social user nodes and social content nodes is obtained.
[0060] Step S34: Merge the direct association relationship of the posted content matching and the indirect social association relationship of the posted content between social user nodes and social content nodes to obtain the social association relationship between each social user node and each social content node.
[0061] In an embodiment of the present invention, by integrating the direct association relationship and the indirect association relationship between social user nodes and social content nodes obtained from previous analysis, the merger of social relationships is achieved to represent the all-round social association relationship between social user nodes and social content nodes, ensuring that different types of association relationships can reflect a more realistic social interaction network during the merger process, and finally obtaining the social association relationship between each social user node and each social content node.
[0062] Further, step S4 includes the following steps: Step S41: Set a corresponding social interaction time window for each social user node in the social user entity node set and each social content node in the social post content entity node set, and count the number of interaction behaviors between each social user node and each social content node according to the social interaction time window, including the number of likes, comments, and forwards; construct an interaction frequency matrix based on the number of interaction behaviors between each social user node and each social content node to generate a two-dimensional matrix of social node interaction frequencies; In an embodiment of the present invention, by setting a social interaction time window for each social user node and each social content node in the social post content entity node set, this time window can be dynamically adjusted according to user activity or fixedly set according to the actual usage pattern of the social platform (such as by day, by week, by month). By analyzing all the activities of the user within the time window, count the number of interaction behaviors between each social user node and the social content node, mainly including the number of likes, comments, and forwards, and all these interaction behaviors will be counted within the social interaction time window. Next, use the counted number of interaction behaviors to construct an interaction frequency matrix, where the rows of the matrix represent social user nodes, the columns represent social content nodes, and the value of each matrix element represents the interaction frequency of the social user node and the social content node (i.e., the combined number of likes, comments, and forwards). In this way, the interaction frequency between users and content can be accurately quantified, and finally a two-dimensional matrix of social node interaction frequencies is constructed and generated.
[0063] Step S42: Based on the two-dimensional matrix of social node interaction frequencies, perform a depth analysis of the social interaction of each social user node and each social content node to obtain the depth of social interaction between each social user node and each social content node; In an embodiment of the present invention, through in-depth analysis of the social interaction between social user nodes and social content nodes based on a previously constructed two-dimensional matrix of social node interaction frequencies. In this step, by performing weighted and normalization processing on the interaction frequencies, the interaction depth between each pair of social user nodes and social content nodes is calculated. The calculation formula for the interaction depth can be set as: Interaction depth = (Number of likes × Weight 1 + Number of comments × Weight 2 + Number of forwards × Weight 3) / Total number of interaction frequencies. Weights 1, 2, and 3 are set according to the actual application requirements of the social platform and are usually adjusted according to the user's participation in the content. By analyzing the interaction frequencies between social user nodes and social content nodes, the social interaction depth between each social user and the content can be evaluated, thereby understanding the user's attention and participation degree in specific content, and finally obtaining the social interaction depth between each social user node and each social content node.
[0064] Step S43: Obtain the social content flow rate and social stagnation time nodes between each social user node and each social content node, and perform an association time decay assessment on the social association relationship between each social user node and each social content node based on the social content flow rate and social stagnation time nodes, so as to obtain the social association time decay degree between each social user node and each social content node; In an embodiment of the present invention, by obtaining the flow rate and stagnation time nodes of social content, the social content flow rate can be calculated by monitoring the time interval for the social content to spread between different social user nodes. Specifically, the social content flow rate can be estimated by analyzing the time interval for the content to spread from one user to another. The faster the flow rate, the more extensive the spread of the content. On the other hand, the social stagnation time node is used to indicate the residence time of the content at a certain node. For example, when the content is published, if the interaction frequency of a certain social user with the content tends to zero within a certain period of time, it is considered that the content has "stagnated" at this node. At the same time, by performing a time decay assessment of the social association relationship based on the social content flow rate and the stagnation time node, the core of the time decay assessment is to evaluate the degree of decay of the relationship between the social user node and the social content node over time according to the spread speed and stagnation time of the social content, that is, the decay degree = spread speed / stagnation time. Specifically, if the social content spread speed is slow and the stagnation time is long, it is considered that the influence of the social content on the user gradually weakens, and then its association weight is decayed, and finally the social association time decay degree between each social user node and each social content node is obtained.
[0065] Step S44: Assign association weights according to the depth of social interaction and the degree of attenuation of social association time effect between each social user node and each social content node, so as to obtain the social association edge weights between each social user node and each social content node; In the embodiment of the present invention, by assigning social association edge weights according to the previously obtained depth of social interaction and the calculated degree of attenuation of social association time effect, in this step, the weight of each social association edge between a pair of social user nodes and social content nodes will be assigned according to the comprehensive result of their interaction depth and time effect attenuation degree. Specifically, the association edge weight between a social user node and a social content node can be calculated by the following method: weight = interaction depth × (1 - attenuation coefficient × stagnation time), where the attenuation coefficient is a constant set according to the flow rate of social content and the stagnation time node. The longer the stagnation time, the higher the attenuation coefficient, and the smaller the weight of the association edge. In this way, the intensity of the interaction relationship between social users and social content can be accurately quantified to reflect the influence of social content on users at different time periods, and finally the social association edge weights between each social user node and each social content node are obtained.
[0066] Step S45: Based on the social association edge weights between each social user node and each social content node and combined with the corresponding social association relationships, perform user social data fusion on the user publication content data corresponding to the social user entity node set and the social publication content entity node set, so as to obtain the user social knowledge fusion result corresponding to the social network platform.
[0067] In the embodiment of the present invention, by based on the previously obtained social association edge weights and combining the user publication content data in the social user entity node set and the social publication content entity node set, perform user social data fusion. The core of this step is to integrate the association information between social user nodes and social content nodes, and then construct a comprehensive user social knowledge graph. The specific operations include using a weighted fusion method to combine the interaction data of each social user node with the feature data of its related social content nodes, so as to construct an all-round user social behavior knowledge network spectrum. The user social knowledge fusion result can not only display the in-depth interaction relationship between users and content, but also provide accurate data support for applications such as personalized recommendation and social network analysis on the social platform. Through this fusion process, the social platform can better understand user interest preferences, content dissemination patterns, and social interaction trends, and then optimize the content recommendation and social interaction strategies of the platform, and finally obtain the user social knowledge fusion result corresponding to the social network platform.
[0068] Furthermore, the present invention also provides a data fusion system based on a social network platform for performing the data fusion method based on a social network platform as described above. The data fusion system based on a social network platform includes: A social network platform data processing module, configured to obtain multi-dimensional data corresponding to each user's post from the social network platform, including text content, picture content, video content, and geographical spatio-temporal information; perform data cleaning and redundancy removal on the multi-dimensional data corresponding to each user's post to obtain multi-dimensional redundancy-removed data corresponding to each user's post; A user social data spatio-temporal classification module, configured to extract spatio-temporal attribute descriptions from the geographical spatio-temporal information in the multi-dimensional redundancy-removed data corresponding to each user's post to obtain spatio-temporal attribute description sub-entities corresponding to each user's post; classify the text content, picture content, and video content corresponding to each user's post based on the spatio-temporal attribute description sub-entities corresponding to each user's post, so as to obtain a dataset of post content sequences corresponding to each user under spatio-temporal attributes; A user entity and content entity relationship mining module, configured to obtain a corresponding set of social user entity nodes and a set of social post content entity nodes through the dataset of post content sequences corresponding to each user under spatio-temporal attributes, and perform social relationship mining and analysis based on the set of social user entity nodes and the set of social post content entity nodes, so as to obtain social association relationships between each social user node and each social content node; A user social data association and fusion module, configured to obtain social association edge weights between each social user node and each social content node through the set of social user entity nodes and the set of social post content entity nodes; perform user social data fusion on the set of social user entity nodes and the set of social post content entity nodes based on the social association edge weights between each social user node and each social content node and in combination with the corresponding social association relationships, so as to obtain a user social knowledge fusion result corresponding to the social network platform.
[0069] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be encompassed within the present invention.
[0070] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data fusion method based on a social network platform, characterized in that: The following steps are involved: Step S1: Obtaining multi-dimensional data corresponding to each user's posting from the social networking platform, including text content, image content, video content, and geographic time and space information; Perform data cleaning and redundancy removal on the multi-dimensional data corresponding to each user's release to obtain the multi-dimensional redundancy-removed data corresponding to each user's release; Step S2: Extract the spatiotemporal attribute description of the geographic spatiotemporal information in the multi-dimensional redundant removal data corresponding to each user's posting, so as to obtain the spatiotemporal attribute descriptor corresponding to each user's posting; classify the spatiotemporal attributes of the corresponding text content, picture content and video content of each user's posting based on the spatiotemporal attribute descriptor corresponding to each user's posting, and obtain the corresponding posting content sequence data set under each user's spatiotemporal attribute; Step S3: Obtain the corresponding social user entity node set and social published content entity node set through the corresponding published content sequence data set under each user's spatiotemporal attributes, and perform social relationship mining and analysis based on the social user entity node set and the social published content entity node set to obtain the social association relationship between each social user node and each social content node; Step S4: obtaining the social association edge weights between each social user node and each social content node through the social user entity node set and the social publishing content entity node set; performing user social data fusion on the social user entity node set and the social publishing content entity node set based on the social association edge weights between each social user node and each social content node and in combination with the corresponding social association relationships, so as to obtain the user social knowledge fusion result corresponding to the social network platform.
2. The data fusion method based on social network platform according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: obtaining the corresponding text content posted by each user from the social networking platform; Step S12: Obtaining the corresponding picture content posted by each user from the social networking platform; Step S13: Obtaining the corresponding video content posted by each user from the social networking platform; Step S14: obtaining the corresponding geographic and temporal information of each user's posting from the social networking platform, including the posting time point and posting location corresponding to each user's posting content; Step S15: Perform data cleaning and standardization on the multi-dimensional data corresponding to each user's release to obtain the multi-dimensional standardized data corresponding to each user's release; perform dynamic redundancy removal on the multi-dimensional standardized data corresponding to each user's release to obtain the multi-dimensional redundancy-removed data corresponding to each user's release.
3. The data fusion method based on social network platform according to claim 2, characterized in that: The step S15 of performing dynamic redundancy removal on the multi-dimensional standardized data corresponding to each user includes the following steps: Based on the geographic spatiotemporal information in the multi-dimensional standardized data corresponding to each user's posting, the corresponding text content, picture content and video content are divided into time and space matching to obtain the text sub-content, picture sub-content and video sub-content corresponding to each user's posting time and place; The publishing time points and publishing locations that are close and the publishing content modes that are similar are divided into the same publishing time and space similarity window, and based on the same publishing time and space similarity window, the text sub-content, picture sub-content and video sub-content corresponding to each user's publishing time and publishing location are dynamically classified into similar windows to obtain the corresponding publishing content group set of each user's publishing in the same time and space similarity window; Performing text semantic analysis on the text sub-contents in the corresponding content group set published by each user in the same spatiotemporal similar window to obtain the text content semantic analysis data corresponding to each user published in the same spatiotemporal similar window, including the sentiment tendency of the text word content and the discussion field of the text word content; Use the pre-trained deep convolutional neural network model to perform visual semantic recognition analysis on the image sub-content and video sub-content in the corresponding content group set published by each user in the same spatiotemporal similar window, and obtain the visual content semantic feature data corresponding to each user's posting in the same spatiotemporal similar window, including content scene, object and character action content features; Based on the text content semantic analysis data and visual content semantic feature data corresponding to each user's posting in the same spatiotemporal similarity window, dynamic redundancy removal is performed on the content group set corresponding to each user's posting in the same spatiotemporal similarity window to obtain multi-dimensional redundant removal data corresponding to each user's posting.
4. The data fusion method based on social network platform according to claim 3 is characterized in that: The method of performing dynamic redundancy removal on a content group set corresponding to each user's content published in the same spatiotemporal similar window based on the text content semantic analysis data and visual content semantic feature data corresponding to each user's content published in the same spatiotemporal similar window comprises the following steps: The content information entropy is calculated for the corresponding content group set published by each user in the same spatiotemporal similarity window to obtain the corresponding content information entropy published by each user in the same spatiotemporal similarity window; According to the preset information entropy threshold, the information entropy of the corresponding published content published by each user in the same spatiotemporal similarity window is compared and judged. If the information entropy of the published content is greater than or equal to the preset information entropy threshold, it means that the content information redundancy of the corresponding published content group set is low, and the corresponding published content group set under its spatiotemporal similarity window is determined as a non-redundant published content group set; if the information entropy of the published content is less than the preset information entropy threshold, it means that the content information redundancy of the corresponding published content group set is high, and the corresponding published content group set under its spatiotemporal similarity window is determined as a redundant published content group set; Performing a publishing content duplication elimination process on the group set of redundant publishing contents determined as redundant publishing contents to obtain a redundant publishing content duplication elimination group set; Obtaining a corresponding spatiotemporal similarity window through a non-redundant published content group set to obtain a non-redundant spatiotemporal similarity window; performing synchronous screening processing on the corresponding text content semantic analysis data and visual content semantic feature data based on the non-redundant spatiotemporal similarity window to obtain the corresponding text content semantic data and visual content semantic data under the non-redundant spatiotemporal similarity window; Based on the corresponding text content semantic data and visual content semantic data in the non-redundant spatiotemporal similarity window, the corresponding non-redundant published content group set is processed with interactive preference low consumption to obtain a non-redundant published content interactive low consumption group set; the redundant published content duplicate elimination group set and the non-redundant published content interactive low consumption group set are merged in time sequence to obtain multi-dimensional redundant removal data corresponding to each user's release.
5. The data fusion method based on social network platform according to claim 4 is characterized in that: The interactive preference low-consumption processing of the corresponding non-redundant published content group set based on the text content semantic data and visual content semantic data corresponding to the non-redundant spatiotemporal similarity window includes the following steps: Obtain the corresponding number of likes and comments on the non-redundant published content by using the corresponding non-redundant published content group set under the non-redundant spatiotemporal similarity window; Based on the number of likes and comments on non-redundant published content, a social interaction heat analysis is performed on the corresponding non-redundant published content group set to obtain the corresponding user published social interaction heat under the non-redundant spatiotemporal similarity window; Based on the corresponding user posting social interaction heat in the non-redundant spatiotemporal similarity window, the social interaction preference calculation formula of the published content is used to evaluate the interaction preference of the text content semantic data and the visual content semantic data in the non-redundant spatiotemporal similarity window, and the corresponding user posting interaction preference score in the non-redundant spatiotemporal similarity window is obtained; Based on the corresponding user posting interaction preference scores in the non-redundant spatiotemporal similarity window, the corresponding non-redundant posting content group set is subjected to interaction preference low-consumption processing to obtain a non-redundant posting content interaction low-consumption group set.
6. The data fusion method based on social network platform according to claim 5, characterized in that: The calculation formula for the social interaction preference of published content is specifically: ; ; Where P is the user’s interaction preference score corresponding to the non-redundant spatiotemporal similarity window T, τ is the window time variable parameter, S t (τ) is the attractiveness of the text sentiment tendency to the user at time τ, ω1 is the contribution coefficient of the text sentiment tendency, D t (τ) is the attractiveness of the discussion field to users at time τ, ω2 is the contribution coefficient of the discussion field, S v (τ) is the attractiveness of the visual scene to the user at time τ, ω3 is the contribution coefficient of the visual scene, O v (τ) is the degree of attraction of the visual object to the user at time τ, ω4 is the contribution coefficient of the visual object, A v (τ) is the degree of attraction of the visual character action to the user at time τ, ω5 is the contribution coefficient of the visual character action, R(T) is the corresponding social interaction heat of the user under the non-redundant spatiotemporal similarity window T, a is the number of likes for the published content, b is the number of comments on the published content, ε is the influence coefficient of the published content, W(T) is the corresponding user publishing time attenuation factor under the non-redundant spatiotemporal similarity window T, and η is the correction coefficient of the user publishing interaction preference score.
7. The data fusion method based on social network platform according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: extracting the time entity and location entity from the geographical spatiotemporal information in the multi-dimensional redundant removal data corresponding to each user's posting, and obtaining the posting time entity and posting location entity corresponding to each user's posting; Step S22: performing entity attribute mining analysis on the publishing time entity and publishing location entity corresponding to each user's publishing, and obtaining the attribute description relationship between the publishing time entity and location entity corresponding to each user's publishing; Step S23: extracting the spatiotemporal attribute description between the corresponding publishing time entity and publishing location entity based on the attribute description relationship between the corresponding time entity and location entity published by each user, so as to obtain the spatiotemporal attribute descriptor corresponding to each user publication; Step S24: Based on the corresponding spatiotemporal attribute descriptor of each user's posting, the corresponding text content, image content and video content of each user's posting are classified according to the spatiotemporal attributes to obtain a corresponding posting content sequence data set under each user's spatiotemporal attributes.
8. The data fusion method based on social network platform according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: taking the social user entity and the social published content entity in the published content sequence data set corresponding to each user's spatiotemporal attribute as nodes to obtain the corresponding social user entity node set and social published content entity node set; Step S32: performing a published content association matching analysis between the corresponding social user nodes in the social user entity node set and the corresponding social content nodes in the social published content entity node set to obtain a published content matching direct association relationship between the social user nodes and the social content nodes; Step S33: performing social interaction relationship mining and analysis on the corresponding social user nodes in the social user entity node set and the corresponding social content nodes in the social published content entity node set, so as to focus on the social interaction related behaviors of users corresponding to the published content on the social network platform, including likes, comments and forwarding, and obtain the social indirect association relationship of the published content between the social user nodes and the social content nodes; Step S34: merge the published content matching direct association relationship between the social user node and the social content node and the published content social indirect association relationship to obtain the social association relationship between each social user node and each social content node.
9. The data fusion method based on social network platform according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: setting a corresponding social interaction time window for each social user node in the social user entity node set and each social content node in the social content publishing entity node set, and counting the number of interaction behaviors between each social user node and each social content node according to the social interaction time window, including the number of likes, comments and reposts; constructing an interaction frequency matrix according to the number of interaction behaviors between each social user node and each social content node, so as to generate a two-dimensional matrix of social node interaction frequencies; Step S42: performing a user social interaction depth analysis on the social association relationship between each social user node and each social content node based on the two-dimensional matrix of social node interaction frequency, and obtaining the social interaction depth between each social user node and each social content node; Step S43: obtaining the social content flow speed and social stagnation time node between each social user node and each social content node, and performing association time decay evaluation on the social association relationship between each social user node and each social content node based on the social content flow speed and the social stagnation time node, to obtain the social association time decay degree between each social user node and each social content node; Step S44: assigning association weights according to the social interaction depth between each social user node and each social content node and the degree of social association time decay, so as to obtain the social association edge weights between each social user node and each social content node; Step S45: Based on the social association edge weights between each social user node and each social content node and in combination with the corresponding social association relationships, user social data fusion is performed on the social user entity node set and the user published content data corresponding to the social published content entity node set to obtain the user social knowledge fusion result corresponding to the social network platform.
10. A data fusion system based on a social network platform, characterized in that: Used to execute the data fusion method based on a social network platform as claimed in claim 1, the data fusion system based on a social network platform comprises: The social network platform data processing module is used to obtain the multi-dimensional data corresponding to each user's posting from the social network platform, including text content, picture content, video content and geographic time and space information; perform data cleaning and redundancy removal on the multi-dimensional data corresponding to each user's posting to obtain the multi-dimensional redundancy-removed data corresponding to each user's posting; The user social data spatiotemporal classification module is used to extract the spatiotemporal attribute description of the geographic spatiotemporal information in the multi-dimensional redundant removal data corresponding to each user's posting, so as to obtain the spatiotemporal attribute descriptor corresponding to each user's posting; based on the spatiotemporal attribute descriptor corresponding to each user's posting, the spatiotemporal attribute classification of the corresponding text content, image content and video content of each user's posting is performed, so as to obtain the corresponding posting content sequence data set under each user's spatiotemporal attribute; The user entity and content entity relationship mining module is used to obtain the corresponding social user entity node set and social publishing content entity node set through the corresponding publishing content sequence data set under each user's spatiotemporal attributes, and perform social relationship mining and analysis based on the social user entity node set and the social publishing content entity node set, so as to obtain the social association relationship between each social user node and each social content node; The user social data association fusion module is used to obtain the social association edge weights between each social user node and each social content node through the social user entity node set and the social publishing content entity node set; based on the social association edge weights between each social user node and each social content node and in combination with the corresponding social association relationships, the social user entity node set and the social publishing content entity node set are used to fuse user social data to obtain the user social knowledge fusion result corresponding to the social network platform.