Identity generation system and method fitting social network group characteristics
By employing multimodal data processing and dynamic feature fitting, the challenge of multimodal data integration in social network identity generation has been solved, achieving dynamic adaptability and accuracy of user identities, improving the reliability of social robots integrating into groups, and meeting the needs of network security and intelligent monitoring.
Patent Information
- Application Number
- CN202511061302.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing methods for generating social network identities struggle to effectively integrate multimodal data and fail to fully capture user behavior patterns across different social scenarios. This results in significant discrepancies between the generated user identities and the real social environment, failing to meet the accuracy and dynamic adaptability requirements of network security and intelligent monitoring.
By employing multimodal data acquisition and preprocessing, natural language processing techniques, gray-level co-occurrence matrix feature extraction, random forest model analysis, and dynamic feature fitting, an identity generation system that fits the characteristics of social network groups is constructed. Through the fusion of text feature vectors and video feature vectors, the user identity generation is dynamically adjusted to achieve dynamic adaptability and accuracy of user identity.
It achieves dynamic adaptability and accuracy of user identity, improves the reliability of social robots integrating into groups, significantly enhances the comprehensiveness and accuracy of identity generation, and provides technical support for network security, public opinion monitoring, and intelligent interaction.
Smart Images

Figure CN120995435A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an identity generation system that fits the characteristics of social network groups, and also to a corresponding identity generation method, belonging to the field of data processing technology. Background Technology
[0002] In the digital age, social networks have become an indispensable part of human society. Their user behavior analysis, group characteristic modeling, and identity generation technologies have wide applications in cybersecurity, public opinion monitoring, and intelligent interaction. With the rapid development of internet technology, the behavioral data of social network users has experienced explosive growth. The data format has evolved from simple text information to a complex collection of multimodal data including text, images, video, and audio. This places higher demands on identity generation technologies in social networks. Traditional identity generation methods often focus on users' textual information, constructing user profiles and determining corresponding user identities by analyzing textual features such as historical speech content, vocabulary usage habits, and grammatical structure. This method was effective in early social network environments because its data format was simple and its processing logic was relatively straightforward, allowing for rapid modeling of users' textual interaction behavior.
[0003] However, with the rise of short videos, live streaming, and other multimedia formats, user data has become increasingly diversified and dynamic. Video data, containing visual features, action sequences, and scene information, can more richly reflect users' interests, preferences, social habits, and behavioral patterns. Relying solely on textual features for identity generation makes it difficult to comprehensively capture users' complete behavioral trajectories in different social scenarios, leading to significant discrepancies between the generated user identities and the real social environment. For example, in video-based social scenarios, users' actions such as liking, commenting, and sharing video content, as well as the visual elements of the video content itself, contain a wealth of crucial information beyond the text. Traditional methods, unable to effectively process this data, result in one-sided identity generation results.
[0004] Meanwhile, with the increasing demands for network security and intelligent monitoring, there is an urgent need to improve the accuracy, dynamic adaptability, and group fitting ability of social network identity generation technology. How to effectively integrate multimodal data, capture dynamic group characteristics, and improve the real-time performance and accuracy of identity generation has become a current research challenge and hot topic in the field. Summary of the Invention
[0005] The primary technical problem to be solved by this invention is to provide an identity generation system that fits the characteristics of social network groups.
[0006] Another technical problem to be solved by the present invention is to provide an identity generation method that fits the characteristics of social network groups.
[0007] To achieve the above-mentioned technical objectives, the present invention adopts the following technical solution:
[0008] According to a first aspect of the present invention, an identity generation system for fitting social network group characteristics is provided, including a collection module, a classification module, a processing module, and a generation module;
[0009] Specifically, the acquisition module sends the preprocessed group dataset and group data with the same network address to the classification module; the classification module sends the group feature vector, group feature sequence, group response sequence and non-response sequence to the processing module; the processing module sends the target fitting feature set to the generation module; and the generation module determines the comprehensive fitting value and determines the user identity based on the comprehensive fitting value.
[0010] Preferably, the acquisition module acquires all user data of the social network platform according to a preset acquisition time, preprocesses all user data to determine a group dataset, and traverses the group dataset to extract all group data of the same network address.
[0011] The classification module determines the group feature vector based on the group data, constructs the group feature sequence based on all the group feature vectors, determines the vector parameter value of each group feature vector in the group feature sequence based on the vector model, and divides the group feature sequence according to the vector parameter value to construct the group response sequence and the group non-response sequence.
[0012] The processing module extracts the same type of group feature vectors from each group response sequence and establishes a fitting feature sequence. Based on the fitting feature sequence, it determines the fitting feature value and determines the corresponding fitting feature value according to the remaining same type of group feature vectors in the group response sequence. It establishes a fitting feature set for all the fitting feature values, compares the fitting feature set with historical data, and determines whether to adjust the fitting feature set based on the comparison result, and determines the target fitting feature set.
[0013] The generation module performs feature fitting based on the target fitting feature set, determines a comprehensive fitting value based on the feature fitting result, and determines the user identity based on the comprehensive fitting value.
[0014] Preferably, the preprocessing includes data noise reduction, removal of duplicate data, and processing of missing values;
[0015] The missing value processing uses a preset interpolation algorithm to interpolate the values of adjacent data.
[0016] Preferably, the group feature vector includes text feature vector and video feature vector;
[0017] The classification module uses NLP technology to analyze all group data, performs word segmentation, stop word removal and stemming on the text data in all group data, determines the target text data, analyzes the relationship between words and extracts noun phrases;
[0018] Noun phrases that conform to syntactic structure are used as keywords and the number of keywords is counted. The data size, data creation time and the number of keywords of each target text data are determined as text features. Each parameter in the text features is normalized to determine the text feature vector.
[0019] The video data in all group data is converted into corresponding grayscale images, and a grayscale co-occurrence matrix is determined based on the grayscale relationship between each pixel and other pixels in the grayscale image. The video feature vector is then determined based on the grayscale co-occurrence matrix.
[0020] Preferably, the classification module acquires a vector dataset and divides the vector dataset into a training set and a test set.
[0021] A random forest model is pre-selected, and the random forest model is iteratively trained based on the training set. The iteratively trained random forest model is then tested based on the test set.
[0022] If the test value of the random forest model after the current iteration is greater than or equal to the test value of the random forest model after the previous iteration, then the iterative training stops and the random forest model after the current iteration is used as the vector model. Otherwise, a training set is determined based on the test set and the training set, and iterative training continues according to the training set until the preset number of iterations is reached.
[0023] The vector model outputs the vector parameter value of each group feature vector. When the vector parameter value is greater than or equal to the preset vector parameter value, the group feature vector corresponding to the vector parameter value is assigned to the group response sequence; otherwise, it is assigned to the group non-response sequence.
[0024] Preferably, the classification module extracts samples from the test set and each training set respectively, extracts features from all samples according to the pre-trained recognition model, estimates the probability density based on the extracted features, and determines the feature distribution of the test set and each training set based on the result of the probability density estimation.
[0025] A loss function is constructed based on the feature distribution, and the extraction probability of each training set is determined according to the loss function. Samples are extracted from the training set according to the extraction probability and formed into the training set.
[0026] Preferably, the processing module determines a preset group feature vector range corresponding to the fitted feature sequence, the preset group feature vector range including a first preset group feature vector and a second preset group feature vector, wherein the first preset group feature vector is greater than the second preset group feature vector;
[0027] The processing module divides the same type of group feature vectors that are greater than the first preset group feature vector into the first feature sequence;
[0028] The processing module divides the same type of group feature vectors in the fitted feature sequence that are less than or equal to the first preset group feature vector and greater than or equal to the second preset group feature vector into the second feature sequence;
[0029] The processing module divides the same type of group feature vectors in the fitted feature sequence that are smaller than the second preset group feature vector into a third feature sequence;
[0030] The fitted feature values are determined based on the first feature sequence, the second feature sequence, and the third feature sequence.
[0031] Preferably, the historical data includes multiple historical fitted feature values;
[0032] When all fitting feature values in the fitting feature set exist in the historical data, the processing module determines not to adjust the fitting feature set and determines the target fitting feature set based on the fitting feature set; otherwise, the fitting feature set is adjusted and the target fitting feature set is determined based on the adjustment result.
[0033] Each data point in the target fitting feature set is denoted as the value to be fitted.
[0034] Preferably, the processing module records the fitted feature values in the fitted feature set that do not exist in the historical data as replacement feature values;
[0035] Take all historical data containing the replacement feature value as the dataset to be aggregated, determine the expected number of clusters k as 2, initialize the parameters of the Gaussian distribution, calculate the probability that each data in the dataset to be aggregated belongs to each Gaussian distribution, and obtain the responsibility value.
[0036] Based on the responsibility value, determine the cluster to which the replacement feature value belongs, and take all historical fitting feature values contained in the cluster as a similar feature set, and obtain the mean value of historical fitting features in the similar set;
[0037] The historical fitting feature mean is replaced with the replacement feature value in the fitting feature set, and the target fitting feature set is determined based on the replacement result.
[0038] According to a second aspect of the present invention, an identity generation method for fitting social network group characteristics is provided, comprising the following steps:
[0039] S1: Obtain all user data from the social network platform according to the preset collection time, preprocess all user data to determine the group dataset, traverse the group dataset, and extract all group data from the same network address;
[0040] S2: Determine the population feature vector based on the population data, construct the population feature sequence based on all population feature vectors, determine the vector parameter value of each population feature vector in the population feature sequence based on the vector model, divide the population feature sequence according to the vector parameter value, and construct the population response sequence and the population non-response sequence;
[0041] S3: Extract the feature vectors of the same type of group in each group response sequence and establish a fitting feature sequence. Determine the fitting feature value based on the fitting feature sequence, and determine the corresponding fitting feature value according to the remaining feature vectors of the same type of group in the group response sequence. Establish a fitting feature set for all fitting feature values, compare the fitting feature set with historical data, and determine whether to adjust the fitting feature set based on the comparison results. Determine the target fitting feature set.
[0042] S4: Perform feature fitting based on the target feature set, determine the comprehensive fitting value based on the feature fitting results, and determine the user identity based on the comprehensive fitting value.
[0043] Compared with existing technologies, this invention innovatively achieves dynamic generation of user identities that fit the social environment through multimodal data acquisition and preprocessing, natural language processing technology and gray-level co-occurrence matrix feature extraction, random forest model analysis, dynamic feature fitting and statistical matching. It effectively integrates the common characteristics of the group and adjusts in real time according to changes in the social environment, making the generated user identities dynamically adaptable and avoiding the risk of being out of touch with reality. Simultaneously, by determining the comprehensive fitting value through feature fitting, the generated user identities more accurately match the user's real characteristics, significantly improving the reliability of social robots integrating into groups and greatly enhancing the comprehensiveness and accuracy of identity generation. This provides strong technical support for fields such as network security, public opinion monitoring, and intelligent interaction. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the structure of an identity generation system that fits the characteristics of a social network group in the first embodiment of the present invention;
[0045] Figure 2 This is a flowchart of an identity generation method for fitting social network group characteristics, as shown in the second embodiment of the present invention. Detailed Implementation
[0046] The technical content of the present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0047] First Embodiment
[0048] like Figure 1 As shown, the first embodiment of the present invention provides an identity generation system that fits the characteristics of social network groups, including a collection module, a classification module, a processing module and a generation module.
[0049] Specifically, the acquisition module sends the preprocessed group dataset and group data with the same network address (including but not limited to IP addresses) to the classification module; the classification module sends the group feature vector, group feature sequence, group response sequence and non-response sequence to the processing module; the processing module sends the target fitting feature set to the generation module; the generation module determines the comprehensive fitting value and determines the user identity based on the comprehensive fitting value.
[0050] In one embodiment of the present invention, the acquisition module acquires all user data of the social network platform according to a preset acquisition time, preprocesses all user data to determine the group dataset, traverses the group dataset, and extracts all group data of the same network address.
[0051] The process of preprocessing all user data to determine the group dataset includes: preprocessing includes data noise reduction, removal of duplicate data, and handling of missing values. The handling of missing values uses a preset interpolation algorithm to interpolate the values of adjacent data.
[0052] In practical applications, user data is often affected by noise interference, such as errors in network transmission and abnormal data records caused by internet failures. Group data, on the other hand, is preprocessed user data, while a group dataset is the collection of all group data. Data denoising can eliminate the influence of this noise, allowing the group data in the group dataset to accurately reflect the true behavioral characteristics of users, thereby improving the overall quality of the group dataset. The presence of duplicate data can introduce biases during data analysis. For example, if a user likes a video five times, the video like data will appear five times. During preprocessing, only the video like data from the single like is retained. Furthermore, duplicate data may lead to the overemphasis of certain features or patterns, affecting the accuracy of the analysis results. Removing duplicate data ensures that each data sample contributes independently to the analysis results, avoiding bias caused by data duplication, effectively reducing data redundancy, and improving data processing efficiency. Preset interpolation algorithms include linear interpolation, polynomial interpolation, and spline interpolation. The specific algorithm chosen depends on the actual amount of user data and is not limited here. In user data, missing values may appear due to various reasons (such as equipment failure, data collection errors, etc.). Directly deleting data records containing missing values may result in the loss of useful information. By interpolating the values of adjacent data based on a preset interpolation algorithm, missing values can be reasonably filled without deleting data records, thereby preserving data integrity and improving the overall data quality of the population dataset.
[0053] In one embodiment of the present invention, the classification module determines the group feature vector based on the group data, constructs the group feature sequence based on all the group feature vectors, determines the vector parameter value of each group feature vector in the group feature sequence based on the vector model, divides the group feature sequence according to the vector parameter value, and constructs the group response sequence and the group non-response sequence.
[0054] When determining the group feature vector based on group data, the group feature vector includes: text feature vector and video feature vector.
[0055] The classification module analyzes all group data according to natural language processing (NLP) technology, performs word segmentation, stop word removal, and stemming on the text data in all group data to determine the target text data, analyzes the relationships between words and extracts noun phrases, uses the noun phrases that conform to the syntactic structure as keywords and counts the number of keywords, determines the data size, data creation time, and keyword number of each target text data as text features, normalizes each parameter in the text features to determine the text feature vector, converts the video data in all group data into corresponding grayscale images, determines the gray-level co-occurrence matrix according to the gray-level relationship between each pixel and other pixels in the grayscale image, and determines the video feature vector based on the gray-level co-occurrence matrix.
[0056] Natural language processing technology reduces the amount of data and highlights key information by removing stop words (such as words with no actual semantic meaning or very weak semantic meaning like "de", "le", "zai", etc.). Stemming can unify different forms of words (such as different tenses of verbs, singular and plural of nouns, etc.) into the stem form (such as "beautiful" for "美丽的"). According to the dependency syntax analysis model of NLP technology, the grammatical and semantic relationships between words are analyzed, and noun phrases are extracted from them. Dependency syntax analysis can accurately identify noun phrases (such as "local users", "event guests"), which reflect the theme of the target text data and are the key basis for keyword extraction. The target text data processed by natural language processing technology can effectively perform feature extraction. The data size, data creation time, and keyword number of each target text data are determined as text features, which quantitatively describe the target text data from multiple dimensions. The data size can reflect the length of the text, the data creation time represents the current time after being processed by natural language processing technology, and the keyword number reflects the information content of the text and the importance of the theme. The text features comprehensively describe the characteristics of the target text data. Then, each parameter in the text features is normalized, and finally, the text feature vector is determined. The process of determining the text feature vector is long and mature, and will not be elaborated here.
[0057] Converting video data into grayscale images reduces the dimension and complexity of the data while retaining important visual information in the video. Grayscale images only contain brightness information. Compared with color images, grayscale images focus more on the structure and texture information of the image. For the analysis of the image structure (calculation of the gray-level co-occurrence matrix), grayscale images can remove interference factors such as color, facilitating the extraction of the video feature vector of the video data. The process of calculating the gray-level co-occurrence matrix and determining the corresponding video feature vector is long and mature, and will not be elaborated here. By extracting the text feature vector and the video feature vector, the feature dimension of user data becomes more extensive, improving the reliability of the social robot's integration into the group and ensuring the stability of user identity generation.
[0058] In determining the vector parameter values of each group feature vector in the group feature sequence based on the vector model, and dividing the group feature sequence according to the vector parameter values to construct the group response sequence and the group non-response sequence, the classification module obtains a vector dataset containing relevant attributes of the group feature vectors (such as sentence length and number of paragraphs corresponding to text feature vectors, color distribution and optical flow features corresponding to video feature vectors, etc.), and divides the vector dataset into training set and test set. A random forest model is pre-selected, and the random forest model is iteratively trained based on the training set, and the iteratively trained random forest model is tested based on the test set. If the test value of the random forest model after the current iteration is greater than or equal to the test value of the random forest model after the previous iteration, the iterative training stops and the random forest model after the current iteration is used as the vector model. Otherwise, a training set is determined based on the test set and the training set, and iterative training continues based on the training set until the preset number of iterations is reached. The vector parameter value of each population feature vector is output according to the vector model. When the vector parameter value is greater than or equal to the preset vector parameter value, the population feature vector corresponding to the vector parameter value is assigned to the population response sequence. Otherwise, it is assigned to the population non-response sequence.
[0059] Since the group feature vectors include both text and video feature vectors, constructing a group feature sequence from all group feature vectors allows for targeted feature analysis. The vector dataset contains data such as sentence length, number of paragraphs, video color, and video optical flow features. The vector dataset is divided into training and test sets, typically with 10-20% of the data used as a training set (divided into 3-4 sets), and the remainder used as a test set to improve the model's generalization performance. The training set is used to train the random forest model, while the test set is used to evaluate the model's performance, with metrics including accuracy, loss function value, and recall. In each iteration, the model attempts to learn patterns and relationships in the data to improve its prediction or classification capabilities. If the test value of the random forest model after the current iteration is greater than or equal to the test value of the random forest model after the previous iteration, it indicates that the model performance has improved or remained stable, and the model is considered to have reached a satisfactory performance level. In this case, the iterative training stops, and the random forest model after the current iteration is used as the vector model. Otherwise, a training set is determined based on the test set and the training set, and iterative training continues. During the continued iterative training, if the test value is still less than the test value of the random forest model after the previous iteration, a training set is determined again based on the test set and the training set. This ensures that the model has enough training time to learn complex patterns and features in the data.
[0060] Meanwhile, based on the vector model, the vector parameter values of each group feature vector are output. When the vector parameter value is greater than or equal to the preset vector parameter value, the group feature vector corresponding to that vector parameter value is assigned to the group response sequence; otherwise, it is assigned to the group non-response sequence. For example, if the group feature vector of the color feature of a very small area is not representative and is unrelated to elements such as theme or plot, its corresponding vector parameter value is low, so it is assigned to the group non-response sequence. More important group feature vectors are assigned to the group response sequence. Targeted processing of the group response sequence improves the efficiency of data processing, thereby closely resembling the user's real identity and ensuring that the generated user identity has dynamic adaptability.
[0061] When determining the training set based on the test set and the training set, the classification module extracts samples from the test set and each training set respectively, extracts features from all samples according to the pre-trained recognition model, estimates the probability density based on the extracted features, determines the feature distribution of the test set and each training set based on the probability density estimation results, constructs a loss function based on the feature distribution, determines the extraction probability of each training set based on the loss function, and extracts samples from the training set according to the extraction probability to form the training set.
[0062] Specifically, by drawing samples from the test and training sets, extracting features using the recognition model, and estimating probability density, the feature distribution of the test and training sets can be accurately grasped, avoiding data bias during model training. A loss function is constructed based on the feature distribution, and the sampling probability of the training set is determined accordingly. This allows for flexible adjustment of the sampling ratio of different training set samples, enabling the model to focus on learning key feature data, improving its ability to capture complex data features, and thus optimizing training performance. The resulting training set comprehensively considers the data features of both the test and training sets, effectively reducing model overfitting. It also exposes the model to representative and diverse data during training, thereby improving the accuracy and reliability of testing.
[0063] In one embodiment of the present invention, the processing module extracts the same type of group feature vectors from each group response sequence and establishes a fitting feature sequence. Based on the fitting feature sequence, it determines the fitting feature value and determines the corresponding fitting feature value according to the remaining same type of group feature vectors in the group response sequence. It establishes a fitting feature set for all the fitting feature values, compares the fitting feature set with historical data, determines whether to adjust the fitting feature set based on the comparison result, and determines the target fitting feature set.
[0064] Specifically, when extracting the same type of group feature vectors from each group response sequence and establishing a fitted feature sequence, and determining the fitted feature value based on the fitted feature sequence, the processing module determines the preset group feature vector range corresponding to the fitted feature sequence. The preset group feature vector range includes a first preset group feature vector and a second preset group feature vector. The first preset group feature vector is greater than the second preset group feature vector. The processing module divides the same type of group feature vectors in the fitted feature sequence that are greater than the first preset group feature vector into the first feature sequence. The processing module divides the same type of group feature vectors in the fitted feature sequence that are less than or equal to the first preset group feature vector and greater than or equal to the second preset group feature vector into the second feature sequence. The processing module divides the same type of group feature vectors in the fitted feature sequence that are less than the second preset group feature vector into the third feature sequence. The fitted feature value is determined based on the first feature sequence, the second feature sequence, and the third feature sequence.
[0065] Specifically, the same-type group feature vectors represent type parameters with consistent video feature vectors and type parameters with consistent text feature vectors, such as punctuation marks, characters, scene frames, action frames, etc. For example, text feature vectors that all contain characters or video feature vectors that all contain action frames are used to establish a fitted feature sequence. The preset group feature vector range corresponding to the fitted feature sequence is represented by the range of the number of characters or the number of action frames. The accuracy of determining the fitted feature values is improved by dividing the feature sequence into a first feature sequence, a second feature sequence, and a third feature sequence. The first quantity of the first feature sequence, the second quantity of the second feature sequence, and the third quantity of the third feature sequence are counted. The fitted feature values are then determined according to the following formula:
[0066]
[0067] Wherein, P represents the fitted feature value; n is the first quantity; Hmax is the maximum value of the vector magnitude of the feature vectors of the same type in the first feature sequence; Hi is the vector magnitude of the i-th feature vector of the same type in the first feature sequence; m is the second quantity; Gmax is the maximum value of the vector magnitude of the feature vectors of the same type in the second feature sequence; Gj is the vector magnitude of the j-th feature vector of the same type in the second feature sequence; z is the third quantity; Lmax is the maximum value of the vector magnitude of the feature vectors of the same type in the third feature sequence; and Lu is the vector magnitude of the u-th feature vector of the same type in the third feature sequence. It should be noted that the vector magnitude in this invention refers to the length of the feature vector of the same type in its direction.
[0068] When comparing the fitted feature set with historical data and determining whether to adjust the fitted feature set based on the comparison results, the historical data includes several historical fitted feature values. When all fitted feature values in the fitted feature set exist in the historical data, the processing module determines not to adjust the fitted feature set and determines the target fitted feature set based on the fitted feature set. Otherwise, the fitted feature set is adjusted, and the target fitted feature set is determined based on the adjustment results. Each data point in the target fitted feature set is recorded as a value to be fitted.
[0069] When all fitted feature values in the fitted feature set exist in the historical data, it is determined that the fitted feature set will not be adjusted, avoiding unnecessary data adjustments. If there are fitted feature values that do not exist in the historical data, it may be due to calculation or other reasons that cause a certain deviation from the historical data. By comparing the fitted feature set with the historical data and dynamically adjusting the fitted feature set, it is possible to capture changes in the social environment in real time while maintaining a certain degree of synchronization with historical social data. This effectively compensates for the problem of feature mismatch, ensuring that the target fitted feature set can accurately reflect the data patterns, thereby improving the accuracy of the fitting results and thus getting closer to the user's real identity, ensuring the reliability and stability of the generated user identity.
[0070] When determining the adjustment of the fitted feature set and identifying the target fitted feature set based on the adjustment results, the processing module records fitted feature values that do not exist in the historical data as replacement feature values. These replacement feature values and the historical data are used as the dataset to be aggregated. The expected number of clusters, k, is set to 2. The parameters of the Gaussian distribution are initialized, and the probability of each data point in the dataset belonging to each Gaussian distribution is calculated to obtain a responsibility value. Based on this responsibility value, the cluster corresponding to the Gaussian distribution to which the replacement feature value is assigned is determined. All historical fitted feature values contained in that cluster are used as a similar feature set, and the arithmetic mean of the historical fitted feature values in this similar feature set (historical fitted feature mean) is calculated. This historical fitted feature mean is then used to replace the replacement feature values in the fitted feature set. The target fitted feature set is determined based on the replacement result.
[0071] By analyzing historical data using clustering algorithms, the dataset most closely matching the current replacement feature value is identified, improving the accuracy of feature replacement. Data-driven automated adjustments enhance the reliability of the social robot's integration into the group. Furthermore, the process of replacing feature values by comprehensively utilizing a large amount of historical data provides rich reference information, allowing for continuous accumulation and ensuring the reliability and stability of user identity generation.
[0072] In one embodiment of the present invention, the generation module performs feature fitting based on the target fitting feature set, determines a comprehensive fitting value based on the feature fitting result, and determines the user identity based on the comprehensive fitting value.
[0073] Specifically, when performing feature fitting based on the target fitting feature set and determining the comprehensive fitting value based on the feature fitting results, the generation module determines the median and mean of the target fitting feature set, extracts the values to be fitted from the target fitting feature set that are greater than the median and constructs the first fitting set, extracts the values to be fitted from the target fitting feature set that are greater than the mean and constructs the second fitting set, and determines whether there is an intersection between the first fitting set and the second fitting set. If there is, the fitting dataset is constructed based on the intersection values and the mean of the fitting dataset is used as the comprehensive fitting value. If not, the mean of the target fitting feature set is used as the comprehensive fitting value, and the user identity is determined by traversing the database based on the comprehensive fitting value.
[0074] Regardless of whether a replacement operation is performed (i.e., regardless of whether replacement feature values exist), the processing module ultimately outputs a target fitting feature set, where each data item is denoted as the value to be fitted. By determining the median and mean of the target fitting feature set and constructing different fitting sets accordingly, the fitting features between data can be mined from multiple perspectives, thereby deriving a comprehensive fitting value. This effectively distinguishes different distribution states in the original data (video data, text data), allowing for targeted fitting. Data values greater than the median and mean often reflect relatively prominent or representative features in the data. Through screening and analysis of these data, the fitting results are made to closely match the user's actual situation. If there is an intersection, it indicates that these data have a certain degree of typicality under different measurement standards. Constructing a fitting dataset from the intersection values and taking the mean as the comprehensive fitting value enhances the representativeness of the fitting results. If there is no intersection, the mean of the target fitting feature set is used as the comprehensive fitting value, ensuring the integrity of the fitting process. This provides a clear and unified basis for subsequent database traversal to determine user identity, improving the efficiency and accuracy of data processing and ensuring the reliability and stability of user identity generation.
[0075] It's important to note that in the database, each user's identity information is associated with a series of feature data, stored in the form of feature vectors, numerical sets, etc. The comprehensive fit value, as the key feature value to be matched, is the basis for determining the user's identity. This value needs to be logically consistent or correlated with the user feature data stored in the database (usually organized as feature vectors or numerical sets). The generation module determines the corresponding user identity by traversing the database to find the user feature data that best matches the comprehensive fit value. The database is determined by a Database Management System (DBMS), and there is a certain mapping relationship between user identities and comprehensive fit values in the database. This mapping relationship ensures that the generated user identity closely matches the user's true characteristics.
[0076] This invention collects massive amounts of user data from social networking platforms at preset intervals, prioritizing the extraction of all data from the same network address. After preprocessing with data denoising, deduplication, and missing value interpolation, a group dataset is formed, ensuring complete capture of user behavior within the same network environment and avoiding omissions. The classification module extracts text and video feature vectors from the preprocessed group data. NLP techniques are used for word segmentation, stop word filtering, stemming, noun phrase recognition, and keyword counting of the text. The video data undergoes grayscale conversion and grayscale co-occurrence matrix processing. Subsequently, a random forest model is used as the vector model to calculate the vector parameter values of each group feature vector, thereby dividing the group response sequence into group non-response sequences. The processing module extracts feature vectors of the same type from the group response sequence to establish... The fitted feature series is further divided into first, second, and third feature series based on a preset two-level threshold. Fitted feature values are calculated according to a predetermined formula to construct a fitted feature set. This set is then compared with historical fitted feature values. If there are replacement feature values that have not appeared before, Gaussian mixture clustering is used to find similar historical averages to replace them, forming a target fitted feature set. The generation module calculates the median and average of the target fitted feature set, constructs the first and second fitted sets respectively, and takes the intersection of the two or the average of the entire set as the comprehensive fitted value. Finally, the database is traversed and matched to determine the user's identity, thereby realizing multimodal data fusion, real-time updating of common group features, and dynamic identity generation. This ensures that the output identity matches the user's real characteristics and improves the reliability of the social robot's integration into the group.
[0077] Second Embodiment
[0078] like Figure 2 As shown, based on the aforementioned identity generation system, the second embodiment of the present invention provides an identity generation method for fitting social network group characteristics, comprising at least the following steps:
[0079] S1: Acquire all user data from the social network platform according to the preset collection time, preprocess all user data to determine the group dataset, traverse the group dataset, and extract all group data from the same network address.
[0080] S2: Determine the population feature vector based on the population data, construct the population feature sequence based on all population feature vectors, determine the vector parameter value of each population feature vector in the population feature sequence based on the vector model, divide the population feature sequence according to the vector parameter value, and construct the population response sequence and the population non-response sequence.
[0081] S3: Extract the feature vectors of the same type of group in each group response sequence and establish a fitting feature sequence. Determine the fitting feature value based on the fitting feature sequence, and determine the corresponding fitting feature value according to the remaining feature vectors of the same type of group in the group response sequence. Establish a fitting feature set for all fitting feature values, compare the fitting feature set with historical data, and determine whether to adjust the fitting feature set based on the comparison results. Determine the target fitting feature set.
[0082] S4: Perform feature fitting based on the target feature set, determine the comprehensive fitting value based on the feature fitting results, and determine the user identity based on the comprehensive fitting value.
[0083] It should be noted that the above embodiments are merely illustrative examples. The technical solutions of each embodiment can be combined, and all are within the protection scope of this invention.
[0084] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0085] The above provides a detailed description of the identity generation system and method for fitting social network group characteristics provided by this invention. Any obvious modifications made by those skilled in the art without departing from the essence of this invention will constitute an infringement of the patent rights of this invention and will incur corresponding legal liability.
Claims
1. An identity generation system that fits the characteristics of social network groups, characterized in that... The system includes a data acquisition module, a classification module, a processing module, and a generation module. The data acquisition module sends a preprocessed group dataset and group data from the same network address to the classification module. The classification module sends group feature vectors, group feature sequences, group response sequences, and non-response sequences to the processing module. The processing module sends a target fitting feature set to the generation module. The generation module determines a comprehensive fitting value and identifies the user based on this value. The acquisition module acquires all user data from the social network platform according to a preset acquisition time, preprocesses all user data to determine a group dataset, and traverses the group dataset to extract all group data from the same network address. The classification module determines the group feature vector based on the group data, constructs the group feature sequence based on all the group feature vectors, determines the vector parameter value of each group feature vector in the group feature sequence based on the vector model, and divides the group feature sequence according to the vector parameter value to construct the group response sequence and the group non-response sequence. The processing module extracts the same type of group feature vectors from each group response sequence and establishes a fitting feature sequence. Based on the fitting feature sequence, it determines the fitting feature value and determines the corresponding fitting feature value according to the remaining same type of group feature vectors in the group response sequence. It establishes a fitting feature set for all the fitting feature values, compares the fitting feature set with historical data, and determines whether to adjust the fitting feature set based on the comparison result, and determines the target fitting feature set. The generation module performs feature fitting based on the target fitting feature set, determines a comprehensive fitting value based on the feature fitting result, and determines the user identity based on the comprehensive fitting value.
2. The identity generation system as described in claim 1, characterized in that... The preprocessing includes data noise reduction, removal of duplicate data, and handling of missing values; The missing value processing uses a preset interpolation algorithm to interpolate the values of adjacent data.
3. The identity generation system as described in claim 2, characterized in that... The group feature vector includes text feature vectors and video feature vectors; The classification module uses NLP technology to analyze all group data, performs word segmentation, stop word removal and stemming on the text data in all group data, determines the target text data, analyzes the relationship between words and extracts noun phrases; Noun phrases that conform to syntactic structure are used as keywords and the number of keywords is counted. The data size, data creation time and the number of keywords of each target text data are determined as text features. Each parameter in the text features is normalized to determine the text feature vector. The video data in all group data is converted into corresponding grayscale images, and a grayscale co-occurrence matrix is determined based on the grayscale relationship between each pixel and other pixels in the grayscale image. The video feature vector is then determined based on the grayscale co-occurrence matrix.
4. The identity generation system as described in claim 3, characterized in that: The classification module acquires a vector dataset and divides the vector dataset into a training set and a test set. A random forest model is pre-selected, and iterative training is performed on the random forest model based on the training set. The iteratively trained random forest model is then tested based on the test set. If the test value of the random forest model after the current iteration is greater than or equal to the test value of the random forest model after the previous iteration, the iterative training is stopped, and the random forest model after the current iteration is used as the vector model. Otherwise, a training set is determined based on the test set and the training set, and iterative training continues based on the training set until the preset number of iterations is reached. The vector model outputs the vector parameter value of each group feature vector. When the vector parameter value is greater than or equal to the preset vector parameter value, the group feature vector corresponding to the vector parameter value is assigned to the group response sequence; otherwise, it is assigned to the group non-response sequence.
5. The identity generation system as described in claim 4, characterized in that... The classification module extracts samples from the test set and each training set respectively, extracts features from all samples according to the pre-trained recognition model, estimates the probability density based on the extracted features, and determines the feature distribution of the test set and each training set based on the probability density estimation results. A loss function is constructed based on the feature distribution, and the extraction probability of each training set is determined according to the loss function. Samples are extracted from the training set according to the extraction probability and formed into the training set.
6. The identity generation system as described in claim 5, characterized in that... The processing module determines a preset group feature vector range corresponding to the fitted feature sequence. The preset group feature vector range includes a first preset group feature vector and a second preset group feature vector, wherein the first preset group feature vector is greater than the second preset group feature vector. The processing module divides the same type of group feature vectors that are greater than the first preset group feature vector into the first feature sequence; The processing module divides the same type of group feature vectors in the fitted feature sequence that are less than or equal to the first preset group feature vector and greater than or equal to the second preset group feature vector into the second feature sequence; The processing module divides the same type of group feature vectors in the fitted feature sequence that are smaller than the second preset group feature vector into a third feature sequence; The fitted feature values are determined based on the first feature sequence, the second feature sequence, and the third feature sequence.
7. The identity generation system as described in claim 6, characterized in that... The historical data includes multiple historical fitted feature values; When all fitting feature values in the fitting feature set exist in the historical data, the processing module determines not to adjust the fitting feature set and determines the target fitting feature set based on the fitting feature set; otherwise, the fitting feature set is adjusted and the target fitting feature set is determined based on the adjustment result. Each data point in the target fitting feature set is denoted as the value to be fitted.
8. The identity generation system as described in claim 7, characterized in that... The processing module records the fitted feature values in the fitted feature set that do not exist in the historical data as replacement feature values; Take all historical data containing the replacement feature values as the dataset to be aggregated, determine the expected number of clusters k as 2, initialize the parameters of the Gaussian distribution, calculate the probability that each data in the dataset to be aggregated belongs to each Gaussian distribution, and obtain the responsibility value. Based on the responsibility value, determine the cluster to which the replacement feature value belongs, and take all historical fitting feature values contained in the cluster as a similar feature set, and obtain the mean value of historical fitting features in the similar set; The historical fitting feature mean is replaced with the replacement feature value in the fitting feature set, and the target fitting feature set is determined based on the replacement result.
9. The identity generation system as described in claim 8, characterized in that... The generation module determines the median and mean of the target fitted feature set; Extract the values to be fitted that are greater than the median from the target fitting feature set and construct a first fitting set; Extract the values to be fitted from the target fitting feature set that are greater than the average value and construct a second fitting set; The generation module determines whether there is an intersection between the first fitting set and the second fitting set; If yes, then construct a fitted dataset based on the intersection values and use the mean of the fitted dataset as the comprehensive fitted value; if no, then use the mean of the target fitted feature set as the comprehensive fitted value. The user's identity is determined by iterating through the database based on the comprehensive fitted values.
10. A method for generating identities by fitting the characteristics of social network groups, implemented based on the identity generation system described in any one of claims 1 to 9, characterized in that... Includes the following steps: S1: Obtain all user data from the social network platform according to the preset collection time, preprocess all user data to determine the group dataset, traverse the group dataset, and extract all group data from the same network address; S2: Determine the population feature vector based on the population data, construct the population feature sequence based on all population feature vectors, determine the vector parameter value of each population feature vector in the population feature sequence based on the vector model, divide the population feature sequence according to the vector parameter value, and construct the population response sequence and the population non-response sequence; S3: Extract the feature vectors of the same type of group in each group response sequence and establish a fitting feature sequence. Determine the fitting feature value based on the fitting feature sequence, and determine the corresponding fitting feature value according to the remaining feature vectors of the same type of group in the group response sequence. Establish a fitting feature set for all fitting feature values, compare the fitting feature set with historical data, and determine whether to adjust the fitting feature set based on the comparison results. Determine the target fitting feature set. S4: Perform feature fitting based on the target fitting feature set, determine the comprehensive fitting value based on the feature fitting result, and determine the user identity based on the comprehensive fitting value.
Citation Information
Patent Citations
Big data analysis method based on decision tree
CN117056834A
Social media data aggregation analysis system and method
CN118013022A
Commercial social personalized recommendation method and system based on data analysis
CN119166903A
Multi-modality classification for one-class classification in social networks
US20110103682A1