A User Information Sharing Method and System Based on Big Data
By performing feature clustering analysis and dynamic credibility threshold screening of false evaluation information from target platforms and multiple similar platforms, the problem of insufficient identification of false evaluations in the existing technology is solved, and more accurate evaluation information filtering and user experience optimization are achieved.
Patent Information
- Application Number
- CN202510405316.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The prior art cannot effectively identify the authenticity of user evaluation information, resulting in the failure to effectively filter false evaluations, affecting users' trust in platform evaluation information.
By performing feature clustering analysis on the false evaluation information of the target platform, the first abnormal evaluation index is determined, and feature clustering analysis is performed on the false evaluation information of multiple similar platforms based on big data, the abnormal evaluation index is determined in conjunction with the dynamic credibility threshold, and reliable evaluation data is screened out.
Accurately identify and filter out unreliable evaluation data, ensure that real evaluation information is fed back to customers, reduce the negative impact of false evaluations on the platform's reputation, optimize user experience, and improve customer satisfaction.
Smart Images

Figure CN119917890B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information sharing, and particularly to a method and system for sharing user information based on big data. Background Art
[0002] With the rapid development of e-commerce and online platforms, user evaluation information has become an important reference basis for consumers to make purchase decisions. The authenticity of online evaluations directly affects consumers' choices of products or services. Therefore, ensuring the authenticity and credibility of evaluation information has become a key issue in platform operation. However, with the continuous expansion of the platform and the increase in the number of users, the problem of false evaluations has gradually emerged and had a significant impact on the credibility of the platform.
[0003] False evaluations usually include false positive evaluations (such as brushing orders, buyer shows, etc.) and malicious negative evaluations (such as malicious competition, slander, etc.). These evaluations are often posted through manual or automated means with the aim of misleading consumers, manipulating product sales, or damaging the reputation of competitors. False positive evaluations and malicious negative evaluations not only mislead consumers' decisions but also make the evaluation system of the platform lose its fairness, affecting the transparency of the platform and consumers' trust.
[0004] Currently, many platforms adopt traditional techniques such as rule detection and keyword matching for the authenticity identification of evaluation information. Most of these traditional methods rely on manually set rules and lack a dynamic adjustment mechanism, thus having many problems. Summary of the Invention
[0005] In view of the technical problem that the existing methods cannot effectively identify the authenticity of user evaluation information, resulting in ineffective filtering of false evaluations and affecting users' trust in the evaluation information of the platform, the present invention provides a method and system for sharing user information based on big data to solve this problem.
[0006] The technical solution of the present invention for solving the above technical problem is as follows:
[0007] In a first aspect, the present invention provides a method for sharing user information based on big data, including: performing feature clustering analysis on false evaluation information of a target platform to determine a first abnormal evaluation index; based on big data, performing feature clustering analysis on false evaluation information of multiple similar platforms of the target platform to determine multiple associated abnormal evaluation indexes; fusing the first abnormal evaluation index and the multiple associated abnormal evaluation indexes to determine an abnormal evaluation index; performing data credibility analysis on the evaluation information set of the target platform according to the abnormal evaluation index, setting the evaluation information that meets the dynamic credibility threshold as reliable evaluation data to obtain a reliable evaluation data set, and performing sharing of user evaluation information of the target platform.
[0008] Preferably, the method for sharing user information based on big data further includes: obtaining a historical false evaluation information set of a target platform, where the false evaluation information includes false positive evaluation data and malicious negative evaluation data; extracting features and performing clustering analysis on the historical false positive evaluation data set according to a predetermined evaluation feature attribute to obtain a first false positive evaluation index; extracting features and performing clustering analysis on the historical malicious negative evaluation data set according to a predetermined evaluation feature attribute to obtain a first malicious negative evaluation index; and using the first false positive evaluation index and the first malicious negative evaluation index as the first abnormal evaluation index.
[0009] Preferably, the method for sharing user information based on big data further includes: obtaining a predetermined evaluation feature attribute, where the predetermined evaluation feature attribute at least includes the number of evaluation words, keywords, modal particles, the number of pictures, and the picture quality; extracting features from multiple historical false positive evaluation data in the historical false positive evaluation data set respectively according to the predetermined evaluation feature attribute to obtain multiple evaluation feature data sets; and using the K-means clustering algorithm to perform feature clustering on the multiple evaluation feature data sets to obtain the first false positive evaluation index.
[0010] Preferably, the method for sharing user information based on big data further includes: taking the platform attribute characteristics of the target platform as a constraint, performing associated retrieval based on big data to determine multiple similar platforms, and obtaining multiple similar historical false evaluation information sets of the multiple similar platforms, where the similar historical false evaluation information includes similar false positive evaluation data and similar malicious negative evaluation data; extracting features and performing clustering analysis on multiple similar historical false positive evaluation data sets and similar historical malicious negative evaluation data sets respectively according to a predetermined evaluation feature attribute to obtain multiple associated abnormal evaluation indexes.
[0011] Preferably, the method for sharing user information based on big data further includes: performing an associated degree analysis on the target platform and multiple similar platforms respectively based on the platform attribute characteristics to determine multiple platform associated degrees; calculating the ratio of each platform associated degree to the sum of the multiple platform associated degrees respectively, which is set as an index fusion weight, to obtain multiple index fusion weights; performing weighted fusion on the multiple associated abnormal evaluation indexes according to the multiple index fusion weights to determine a first associated abnormal evaluation index; and performing secondary weighted fusion on the first abnormal evaluation index and the first associated abnormal evaluation index to obtain the abnormal evaluation index, where the fusion weight of the first abnormal evaluation index is 0.7 and the fusion weight of the first associated abnormal evaluation index is 0.3.
[0012] Preferably, the method for sharing user information based on big data further includes: obtaining the abnormal evaluation index, where the abnormal evaluation index includes a plurality of abnormal evaluation attribute data; randomly selecting a first evaluation information from the evaluation information set, and respectively performing similarity analysis on the first evaluation information according to the plurality of abnormal evaluation attribute data to obtain a first similarity set, and calculating the mean value to obtain a first similarity coefficient; subtracting the first similarity coefficient from 1 to obtain a first data credibility coefficient, and sequentially analyzing to obtain a plurality of data credibility coefficients of a plurality of evaluation information in the evaluation information set.
[0013] Preferably, the method for sharing user information based on big data further includes: based on big data, obtaining the sequence of the number of false evaluation information and the sequence of the proportion of false evaluation information of the target platform in a preset historical time period, as well as the sequences of the number of false evaluation information of a plurality of similar platforms and the sequences of the proportion of false evaluation information of a plurality of similar platforms in a preset time period; pre-training an evaluation trend analysis plug-in, and inputting the sequence of the number of false evaluation information, the sequence of the proportion of false evaluation information, the plurality of platform correlation degrees, the sequences of the number of false evaluation information of a plurality of similar platforms and the sequences of the proportion of false evaluation information of a plurality of similar platforms into the evaluation trend analysis plug-in to output an evaluation trend coefficient; optimizing and adjusting a predetermined credibility threshold according to the evaluation trend coefficient to obtain a dynamic credibility threshold; judging the plurality of data credibility coefficients according to the dynamic credibility threshold, and setting the evaluation information corresponding to the data credibility coefficient greater than the dynamic credibility threshold as reliable evaluation data to obtain a reliable evaluation data set.
[0014] Preferably, the method for sharing user information based on big data further includes: based on big data, collecting a sample training data set, where the sample training data includes a sequence of the number of sample false evaluation information, a sequence of the proportion of sample false evaluation information, a set of sample platform correlation degrees, a set of sequences of the number of sample false evaluation information of similar platforms, a set of sequences of the proportion of sample false evaluation information of similar platforms, and a sample evaluation trend coefficient, and the sample evaluation trend coefficient can be obtained through manual annotation; using the sample training data set to perform supervised training and testing on a BP neural network until the loss function converges to obtain a trained evaluation trend analysis plug-in.
[0015] Second aspect, the present invention provides a user information sharing system based on big data, including: a first abnormal evaluation index determination module, configured to perform feature clustering analysis on false evaluation information of a target platform to determine a first abnormal evaluation index; an associated abnormal evaluation index determination module, configured to perform feature clustering analysis on false evaluation information of multiple similar platforms of the target platform based on big data to determine multiple associated abnormal evaluation indexes; an abnormal evaluation index fusion module, configured to fuse and determine an abnormal evaluation index according to the first abnormal evaluation index and multiple associated abnormal evaluation indexes; a data credibility analysis module, configured to perform data credibility analysis on an evaluation information set of the target platform according to the abnormal evaluation index, set evaluation information that meets a dynamic credibility threshold as reliable evaluation data, obtain a reliable evaluation data set, and perform user evaluation information sharing of the target platform.
[0016] The beneficial effects of the present invention are as follows: by performing feature clustering analysis on false evaluation information of the target platform, a first abnormal evaluation index is determined; then, based on big data, feature clustering analysis is performed on false evaluation information of multiple similar platforms of the target platform to determine multiple associated abnormal evaluation indexes; further, an abnormal evaluation index is determined by fusing the first abnormal evaluation index and multiple associated abnormal evaluation indexes; then, data credibility analysis is performed on the evaluation information set of the target platform according to the abnormal evaluation index, evaluation information that meets the dynamic credibility threshold is set as reliable evaluation data, a reliable evaluation data set is obtained, and user evaluation information sharing of the target platform is performed; that is to say, by performing clustering analysis on false evaluation features of the target platform and multiple similar platforms, false positive reviews and malicious negative reviews can be identified more accurately, unreliable evaluation data can be filtered out, real evaluation information can be ensured to be fed back to customers, thereby reducing the negative impact of false evaluations on the platform's reputation, optimizing the user experience, and improving customer satisfaction. Description of the Drawings
[0017] Figure 1 It is a schematic flowchart of a user information sharing method based on big data provided by the present invention;
[0018] Figure 2 It is a schematic structural diagram of a user information sharing system based on big data provided by the present invention.
[0019] In the drawings, the components represented by each reference numeral are described as follows:
[0020] The first abnormal evaluation index determination module 11, the associated abnormal evaluation index determination module 12, the abnormal evaluation index fusion module 13, and the data credibility analysis module 14. Detailed Embodiments
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0022] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0023] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or having more advantages than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid unnecessary details from obscuring the description of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope that conforms to the principles and features disclosed in the present invention.
[0024] Embodiment 1, as Figure 1 shown, the embodiment of the present invention provides a method for sharing user information based on big data, which specifically includes the following steps:
[0025] S100: Perform feature clustering analysis on the false evaluation information of the target platform to determine the first abnormal evaluation index.
[0026] Further, step S100 of the present invention further includes:
[0027] S110: Obtain the historical false evaluation information set of the target platform, where the false evaluation information includes false positive evaluation data and malicious negative evaluation data.
[0028] Specifically, all historical user evaluation data is obtained from the target platform, which includes evaluation information of all products, services, merchants, or service providers. Then, according to the existing data on the platform or manual annotation, a set of historical false evaluation information of the target platform is obtained. Among them, false evaluation information includes false positive evaluation data and malicious negative evaluation data. False positive evaluation data usually manifests as overly flattering evaluations without actual experience or obviously unreasonable positive evaluations (such as overly consistent evaluation content, extremely positive scores, etc.). Malicious negative evaluation data manifests as negative evaluations without actual usage experience, extreme and malicious negative evaluations, or obvious signs of malicious attacks by competitors. By obtaining the set of historical false evaluation information of the target platform and classifying it into false positive evaluation data and malicious negative evaluation data, this provides basic data support for subsequent false evaluation identification, feature analysis, and credibility assessment.
[0029] S120: Extract features and perform clustering analysis on the historical false positive evaluation data set according to the predetermined evaluation feature attributes to obtain the first false positive evaluation index.
[0030] Furthermore, step S120 of the present invention further includes:
[0031] S121: Obtain the predetermined evaluation feature attributes, where the predetermined evaluation feature attributes at least include the number of evaluation words, keywords, mood words, the number of pictures, and the quality of pictures; S122: Extract features from multiple historical false positive evaluation data in the historical false positive evaluation data set according to the predetermined evaluation feature attributes to obtain multiple evaluation feature data sets; S123: Use the K-means clustering algorithm to perform feature clustering on the multiple evaluation feature data sets to obtain the first false positive evaluation index.
[0032] Specifically, first, obtain the predetermined evaluation feature attributes, where the predetermined evaluation feature attributes at least include the number of evaluation words, keywords, mood words, the number of pictures, and the quality of pictures. The number of evaluation words refers to the number of words in each evaluation. False evaluations may have certain patterns in terms of the number of words (such as being too short or too long), which do not conform to the natural evaluation habits of real users. Keywords refer to specific words included in the evaluation, and these words can reveal the authenticity of the evaluation. For example, frequently appearing keywords such as "good", "perfect", "very satisfied", etc. may be characteristics of false positive evaluations. Mood words refer to words or phrases used in the evaluation to express emotions, such as mood words like "very" and "extremely". The frequency and intensity of the use of these mood words may be more extreme in false evaluations. The number of pictures refers to the number of pictures attached to the evaluation. False positive evaluations often may include a large number of irrelevant pictures or a small number of but irrelevant pictures. The quality of pictures refers to the clarity, authenticity, etc. of the pictures attached to the evaluation. Low-quality pictures or pictures that do not match the evaluation content may be characteristics of false evaluations.
[0033] Next, according to the above-mentioned predetermined evaluation feature attributes, feature extraction is respectively performed on multiple historical false positive review data in the historical false positive review dataset, including extracting the number of words in the review, extracting keywords, extracting modal particles, extracting the number of pictures, and extracting the picture quality. For example, perform character counting on the content of each review and record the number of words; use natural language processing (NLP) techniques (such as word segmentation, part-of-speech tagging, etc.) to process the review content and extract the high-frequency keywords therein; through regular expressions or sentiment analysis tools, extract the modal particles in the review (such as "extremely", "very", "perfect", etc.); according to the picture field in the review data, count the number of pictures attached to each review; through image processing techniques (such as clarity detection, content analysis, etc.) evaluate the quality of the pictures attached to each review, evaluate factors such as the resolution of the pictures, whether they are clear, and whether they are relevant to the review content, and generate a picture quality score; for each piece of historical false positive review data, extract the corresponding features according to the above steps to generate an evaluation feature dataset containing all predetermined feature attributes, and obtain multiple evaluation feature datasets. By extracting predetermined evaluation feature attributes such as the number of words in the review, keywords, modal particles, the number of pictures, and picture quality from the historical false positive review dataset, an evaluation dataset containing multi-dimensional features can be constructed, which provides basic data for subsequent false review identification, clustering analysis, and credibility evaluation, and helps to accurately identify false reviews.
[0034] K-means clustering is an unsupervised learning method that divides data into K clusters, and each cluster contains samples with similar features. Then, use the K-means clustering algorithm to perform feature clustering on the multiple evaluation feature datasets. First, randomly select K initial centroids; then, for each data point, calculate its distance from all centroids and assign it to the nearest centroid; further update the centroid of each cluster, and the new centroid is the mean of all data points within the cluster; then repeat the above steps until the cluster assignment no longer changes or reaches the maximum number of iterations; after clustering, the feature values of each cluster can represent some common patterns of false positive reviews. Extract the features that frequently appear in false positive reviews from the clustering results. For example, certain specific word count ranges, keyword frequencies, the use of specific modal particles, excessive number of pictures, etc. If the centroid of a certain cluster indicates that the review word count is generally long and the frequency of using modal particles is high, this may indicate that this cluster belongs to typical false positive reviews. Then, through the analysis of the clustering results, it can be determined which features appear frequently in false positive reviews, and set the high-frequency features as the first false positive review indicators. For example, frequently appearing positive words such as "perfect", "absolutely good", "most satisfied", etc.; low-quality pictures or irrelevant pictures may be signs of false positive reviews, etc. These indicators will be used for subsequent false review identification and filtering.
[0035] S130: Extract features and perform clustering analysis on the historical malicious negative review dataset according to the predetermined evaluation feature attributes to obtain the first malicious negative review metric; S140: Use the first false positive review metric and the first malicious negative review metric as the first abnormal evaluation metric.
[0036] Specifically, extract features from the historical malicious negative review dataset according to the predetermined evaluation feature attributes, that is, extract relevant feature data from the historical malicious negative review dataset and generate a feature dataset, including extracting the review word count, extracting keywords (using natural language processing techniques to analyze the keywords in each negative review, especially negative sentiment words such as "bad", "terrible", "disappointed", etc.), and using regular expressions or sentiment analysis tools to extract the modal particles in the review, such as "the worst", "will never come again", etc.; then use the K-means clustering algorithm to perform feature clustering on the feature dataset to obtain the first malicious negative review metric, that is, the most representative malicious negative review features, which will be used for subsequent malicious negative review identification and filtering. Finally, use the first false positive review metric and the first malicious negative review metric as the first abnormal evaluation metric.
[0037] By performing feature extraction and K-means clustering analysis on the historical malicious negative review dataset, typical features of malicious negative reviews can be identified, such as overly short review word counts, frequently occurring negative keywords, extreme modal particles, abnormal number of pictures, and low-quality pictures, etc.; the first malicious negative review metric obtained through clustering analysis will help identify and filter malicious negative reviews in the future and improve the credibility of platform evaluations.
[0038] S200: Based on big data, perform feature clustering analysis on the false evaluation information of multiple similar platforms of the target platform to determine multiple associated abnormal evaluation metrics.
[0039] Furthermore, step S200 of the present invention further includes:
[0040] S210: With the platform attribute features of the target platform as a constraint, perform associated retrieval based on big data to determine multiple similar platforms, and obtain multiple sets of similar historical false evaluation information of the multiple similar platforms, where the similar historical false evaluation information includes similar false positive review data and similar malicious negative review data; S220: According to the predetermined evaluation feature attributes, perform feature extraction and clustering analysis on multiple sets of similar historical false positive review datasets and similar historical malicious negative review datasets respectively to obtain multiple associated abnormal evaluation metrics.
[0041] Specifically, first, it is necessary to clarify the platform attribute characteristics of the target platform, including but not limited to platform type, user activity, product or service category, user group characteristics, etc.; then, with the platform attribute characteristics of the target platform as a constraint, based on big data for associated retrieval, multiple similar platforms are determined, that is, through big data analysis and associated retrieval technologies, similar platforms with similar attributes to the target platform are found, and by calculating the similarity of the various attributes of the target platform with the attributes of other platforms, similar platforms with high similarity are identified. Then for each similar platform, a historical false evaluation information set of this platform is obtained, and multiple similar historical false evaluation information sets of the multiple similar platforms are obtained, where the similar historical false evaluation information includes similar false positive evaluation data and similar malicious negative evaluation data.
[0042] Next, according to the predetermined evaluation characteristic attributes, feature extraction and clustering analysis are respectively performed on multiple similar historical false positive evaluation data sets and similar historical malicious negative evaluation data sets. The feature extraction and clustering methods are the same as the methods for obtaining the above first false positive evaluation index, and are not elaborated here. Multiple associated abnormal evaluation indexes are obtained, and these indexes can help distinguish the characteristics of false positive evaluations and malicious negative evaluations.
[0043] S300: Determine the abnormal evaluation index according to the fusion of the first abnormal evaluation index and multiple associated abnormal evaluation indexes.
[0044] Furthermore, step S300 of the present invention further includes:
[0045] S310: Based on the platform attribute characteristics, perform associated degree analysis on the target platform and multiple similar platforms respectively to determine multiple platform associated degrees; S320: Calculate the ratio of the associated degree of each platform to the sum of the associated degrees of multiple platforms respectively, and set it as the index fusion weight to obtain multiple index fusion weights; S330: Perform weighted fusion on the multiple associated abnormal evaluation indexes according to the multiple index fusion weights to determine the first associated abnormal evaluation index; S340: Perform secondary weighted fusion on the first abnormal evaluation index and the first associated abnormal evaluation index to obtain the abnormal evaluation index, where the fusion weight of the first abnormal evaluation index is 0.7 and the fusion weight of the first associated abnormal evaluation index is 0.3.
[0046] Specifically, first, based on the platform attribute characteristics (platform type, user activity, product or service category, user group characteristics), the correlation analysis is respectively carried out on the target platform and multiple similar platforms to calculate the similarity or relationship degree between each platform. For example, by calculating the cosine similarity of the attribute characteristics between platforms, the correlation degree between the target platform and similar platforms is obtained, and the correlation degrees of multiple platforms are determined. Then, the ratio of the correlation degree of each platform to the sum of the correlation degrees of multiple platforms is calculated respectively, which is set as the index fusion weight, and multiple index fusion weights are obtained. Further, the multiple associated abnormal evaluation indicators are weighted and fused according to the multiple index fusion weights, that is, according to the calculated index fusion weight of each platform, the multiple associated abnormal evaluation indicators are weighted and calculated to obtain the first associated abnormal evaluation indicator. Obtain the fusion weight of the first abnormal evaluation indicator and the fusion weight of the first associated abnormal evaluation indicator, where the fusion weight of the first abnormal evaluation indicator is 0.7 and the fusion weight of the first associated abnormal evaluation indicator is 0.3; then, the first abnormal evaluation indicator and the first associated abnormal evaluation indicator are subjected to secondary weighted fusion to obtain the abnormal evaluation indicator. Through the weighted fusion of the first abnormal evaluation indicator and the first associated abnormal evaluation indicator, the final abnormal evaluation indicator can be obtained. This indicator comprehensively considers the characteristics and historical evaluation data of the target platform and multiple similar platforms, and can more accurately identify and filter false evaluations or malicious negative reviews. Through the correlation analysis of platform attribute characteristics and combined with the clustering results of false evaluation characteristics of multiple platforms, a more accurate abnormal evaluation indicator can be obtained by weighted fusion.
[0047] S400: According to the abnormal evaluation indicator, perform data credibility analysis on the evaluation information set of the target platform, set the evaluation information that meets the dynamic credibility threshold as reliable evaluation data, obtain a reliable evaluation data set, and perform user evaluation information sharing of the target platform.
[0048] Further, step S400 of the present invention further includes:
[0049] S410: Obtain the abnormal evaluation indicator, where the abnormal evaluation indicator includes multiple abnormal evaluation attribute data; S420: Randomly select the first evaluation information in the evaluation information set, and respectively perform similarity analysis on the first evaluation information according to the multiple abnormal evaluation attribute data to obtain a first similarity set, and calculate the mean value to obtain a first similarity coefficient; S430: Subtract the first similarity coefficient from 1 to obtain a first data credibility coefficient, and sequentially analyze the multiple data credibility coefficients of multiple evaluation information in the evaluation information set.
[0050] Specifically, obtain the abnormal evaluation index, where the abnormal evaluation index includes multiple abnormal evaluation attribute data, including the number of evaluation words, the frequency of keyword occurrences, the frequency of using filler words, the number and quality of pictures, etc.; then, randomly select any evaluation information in the evaluation information set as the first evaluation information, and perform similarity analysis on the first evaluation information respectively according to the multiple abnormal evaluation attribute data, that is, compare the characteristic data of the first evaluation information with other evaluations in the evaluation information set, calculate the similarity, such as calculating the similarity between evaluation information through the cosine similarity of characteristic data, and then form a first similarity set with the similarity values between the first evaluation information and other evaluation information, that is, it contains the similarity scores between the first evaluation information and all other evaluation information; further calculate the mean value of all similarity values in the first similarity set to obtain the first similarity coefficient. Among them, a high similarity means that the evaluation information is highly similar to other evaluation information and may be a false evaluation, so the credibility is low; a low similarity means that the evaluation information is relatively independent and the credibility is high. Then subtract the first similarity coefficient from 1, and use the difference between the two as the first data credibility coefficient. This coefficient represents the credibility of the evaluation information and can be used as a basis for screening true evaluations and false evaluations. Then use the same method for calculating the first data credibility coefficient to calculate the data credibility coefficients of multiple evaluation information in the evaluation information set in turn, and obtain multiple data credibility coefficients of multiple evaluation information.
[0051] Further, step S400 of the present invention further includes:
[0052] S440: Based on big data, obtain the sequence of the number of false evaluation information and the sequence of the proportion of false evaluation information of the target platform in a preset historical time zone, as well as the sequences of the number of multiple similar false evaluation information and the sequences of the proportion of multiple similar false evaluation information of multiple similar platforms in a preset time zone.
[0053] Specifically, configure a preset historical time zone, that is, select an appropriate time range according to requirements (for example, the past week, month, quarter, etc.), and this time interval will be used as the basis for analyzing false evaluation information. Then, based on big data, obtain the sequence of the number of false evaluation information and the sequence of the proportion of false evaluation information in the target platform within the preset historical time zone, that is, count the number of false evaluations in each time period (such as every hour, day, week, etc.) within the preset time zone of the target platform to form a sequence of the number of false evaluations; calculate the proportion of false evaluations in each time period within the preset time zone of the target platform. On the other hand, obtain multiple sequences of the number of false evaluation information of multiple similar platforms and multiple sequences of the proportion of false evaluation information of multiple similar platforms within the preset time zone. By obtaining the historical false evaluation information of the target platform and multiple similar platforms and analyzing their quantity and proportion sequences, the performance and change trend of false evaluations on each platform can be clearly understood, providing effective support for further optimizing the false evaluation filtering algorithm and improving the credibility of platform data.
[0054] S450: Pre-train an evaluation trend analysis plugin, and input the sequence of the number of false evaluation information, the sequence of the proportion of false evaluation information, the association degrees of multiple platforms, the sequences of the number of false evaluation information of multiple similar platforms, and the sequences of the proportion of false evaluation information of multiple similar platforms into the evaluation trend analysis plugin to output an evaluation trend coefficient.
[0055] Furthermore, step S450 of the present invention further includes:
[0056] S451: Based on big data, collect a sample training data set, where the sample training data includes a sequence of the number of sample false evaluation information, a sequence of the proportion of sample false evaluation information, a set of sample platform association degrees, a set of sequences of the number of sample false evaluation information of similar platforms, a set of sequences of the proportion of sample false evaluation information of similar platforms, and a sample evaluation trend coefficient, and the sample evaluation trend coefficient can be obtained through manual annotation; S452: Use the sample training data set to perform supervised training and testing on a BP neural network until the loss function converges to obtain a trained evaluation trend analysis plugin.
[0057] Specifically, first, based on big data, a sample training dataset is collected. The sample training data includes the sample false evaluation information quantity sequence, the sample false evaluation information proportion sequence, the sample platform correlation set, the sample similar false evaluation information quantity sequence set, the sample similar false evaluation information proportion sequence set, and the sample evaluation trend coefficient. The evaluation trend coefficient represents the change trend of false evaluations on the target platform and similar platforms within a preset time interval, which can be determined through manual annotation. The annotator determines the rising, falling, or stable degree of the trend by observing the fluctuation trends of the false evaluation quantity sequence and proportion sequence of the platform. For example, if the false evaluation quantity of a certain platform shows an upward trend, the trend coefficient is marked as 1.1; if it shows a downward trend, it can be marked as 0.95.
[0058] Next, an evaluation trend analysis plug-in is constructed based on the BP neural network. The evaluation trend analysis plug-in is a BP neural network model in machine learning that can be iteratively optimized. The evaluation trend analysis plug-in includes an input layer, multiple hidden layers, and an output layer. Among them, the input data of the input layer is the sample false evaluation information quantity sequence, the sample false evaluation information proportion sequence, the sample platform correlation set, the sample similar false evaluation information quantity sequence set, and the sample similar false evaluation information proportion sequence set, and the output data is the sample evaluation trend coefficient. Then, the sample training dataset is divided into a training set and a validation set according to a predetermined ratio. For example, the training set accounts for 85% and the validation set accounts for 15%. Then, the training set and the validation set are used to perform supervised training and validation on the evaluation trend analysis plug-in. During the training process, first, the input data in the training set is propagated through each layer of the neural network to calculate the output of each node, and finally the predicted evaluation trend coefficient is obtained. Then, the loss function is used to calculate the gap between the predicted value and the true value. Then, the gradient of the loss function with respect to each weight is calculated, and the weights and biases in the neural network are updated through the gradient descent method. Iterate continuously until the loss function converges. Each iteration optimizes the weights through forward propagation and backward propagation, and finally makes the loss function as small as possible. Among them, after each round of training, the validation set is used to verify the model, and the error between the predicted result and the true result on the validation set is calculated. By comparing the errors on the training set and the validation set, the training progress of the model can be monitored to avoid overfitting. An evaluation trend analysis plug-in that has completed training is obtained.
[0059] Next, the false evaluation information quantity sequence, the false evaluation information proportion sequence, multiple platform correlations, multiple similar false evaluation information quantity sequences, and multiple similar false evaluation information proportion sequences are input into the evaluation trend analysis plug-in that has completed training, and the evaluation trend coefficient is output.
[0060] S460: Optimize and adjust a predetermined credibility threshold according to the evaluation trend coefficient to obtain a dynamic credibility threshold; S470: Judge the multiple data credibility coefficients according to the dynamic credibility threshold, and set the evaluation information corresponding to the data credibility coefficient greater than the dynamic credibility threshold as reliable evaluation data to obtain a reliable evaluation data set.
[0061] Specifically, according to the evaluation trend coefficient, the credibility threshold is dynamically adjusted so as to increase the threshold when the number of false evaluations increases to reduce the influence of false evaluations; when the number of false evaluations decreases, the threshold is decreased to more accurately screen out true evaluations, that is, multiply the evaluation trend coefficient by the predetermined credibility threshold to obtain the dynamic credibility threshold. Further judge the multiple data credibility coefficients according to the dynamic credibility threshold. If the data credibility coefficient is greater than the dynamic credibility threshold, it indicates that the credibility of this evaluation information is relatively high and can be considered as reliable evaluation data; if the data credibility coefficient is less than or equal to the dynamic credibility threshold, it means that the credibility of this evaluation information is relatively low and can be considered as false evaluation data and needs to be filtered out; then set the evaluation information corresponding to the data credibility coefficient greater than the dynamic credibility threshold as reliable evaluation data to obtain a reliable evaluation data set.
[0062] By optimizing and adjusting the credibility threshold based on the evaluation trend coefficient and screening out reliable evaluation data according to the dynamic credibility threshold, the platform can more accurately identify and filter false evaluations, improve the quality and accuracy of evaluation data; ultimately, the platform can provide users with true and reliable evaluation information, enhance users' trust, and thus improve the overall reputation and user experience of the platform.
[0063] A method for sharing user information based on big data provided by an embodiment of the present invention has at least the following technical effects:
[0064] By performing feature clustering analysis on false evaluation information of a target platform to determine a first abnormal evaluation index; then, based on big data, performing feature clustering analysis on false evaluation information of multiple similar platforms of the target platform to determine multiple associated abnormal evaluation indexes; further determining an abnormal evaluation index by fusing the first abnormal evaluation index and the multiple associated abnormal evaluation indexes; then performing data credibility analysis on the evaluation information set of the target platform according to the abnormal evaluation index, and setting the evaluation information that meets the dynamic credibility threshold as reliable evaluation data to obtain a reliable evaluation data set for sharing user evaluation information of the target platform; that is to say, by performing clustering analysis on the false evaluation characteristics of the target platform and multiple similar platforms, false positive reviews and malicious negative reviews can be more accurately identified, unreliable evaluation data can be filtered out, and true evaluation information can be ensured to be fed back to customers, thereby reducing the negative impact of false evaluations on the platform's reputation, optimizing the user experience, and improving customer satisfaction.
[0065] Embodiment 2, as follows Figure 2 As shown, based on the same inventive concept as the user information sharing method based on big data provided in Embodiment 1, the embodiment of the present invention further provides a user information sharing system based on big data, including: a first abnormal evaluation index determination module 11, configured to perform feature clustering analysis on the false evaluation information of the target platform to determine the first abnormal evaluation index; an associated abnormal evaluation index determination module 12, configured to perform feature clustering analysis on the false evaluation information of multiple similar platforms of the target platform based on big data to determine multiple associated abnormal evaluation indexes; an abnormal evaluation index fusion module 13, configured to fuse and determine the abnormal evaluation index according to the first abnormal evaluation index and the multiple associated abnormal evaluation indexes; a data credibility analysis module 14, configured to perform data credibility analysis on the evaluation information set of the target platform according to the abnormal evaluation index, set the evaluation information that meets the dynamic credibility threshold as reliable evaluation data, obtain a reliable evaluation data set, and perform user evaluation information sharing of the target platform.
[0066] Furthermore, the user information sharing system based on big data is further configured to: obtain a historical false evaluation information set of the target platform, where the false evaluation information includes false positive evaluation data and malicious negative evaluation data; perform feature extraction and clustering analysis on the historical false positive evaluation data set according to a predetermined evaluation feature attribute to obtain a first false positive evaluation index; perform feature extraction and clustering analysis on the historical malicious negative evaluation data set according to a predetermined evaluation feature attribute to obtain a first malicious negative evaluation index; and use the first false positive evaluation index and the first malicious negative evaluation index as the first abnormal evaluation index.
[0067] Furthermore, the user information sharing system based on big data is further configured to: obtain a predetermined evaluation feature attribute, where the predetermined evaluation feature attribute at least includes the number of evaluation words, keywords, modal particles, the number of pictures, and the picture quality; perform feature extraction on multiple historical false positive evaluation data in the historical false positive evaluation data set according to the predetermined evaluation feature attribute to obtain multiple evaluation feature data sets; and use the K-means clustering algorithm to perform feature clustering on the multiple evaluation feature data sets to obtain the first false positive evaluation index.
[0068] Furthermore, the user information sharing system based on big data is further configured to: perform associated retrieval based on big data with the platform attribute characteristics of the target platform as a constraint to determine multiple similar platforms, and obtain multiple similar historical false evaluation information sets of the multiple similar platforms, where the similar historical false evaluation information includes similar false positive evaluation data and similar malicious negative evaluation data; perform feature extraction and clustering analysis on multiple similar historical false positive evaluation data sets and similar historical malicious negative evaluation data sets according to a predetermined evaluation feature attribute to obtain multiple associated abnormal evaluation indexes.
[0069] Further, the user information sharing system based on big data is also used for: based on platform attribute characteristics, respectively performing correlation analysis on the target platform and multiple similar platforms to determine the correlation degrees of the multiple platforms; respectively calculating the ratios of each platform's correlation degree to the sum of the correlation degrees of the multiple platforms, which are set as index fusion weights, to obtain multiple index fusion weights; performing weighted fusion on the multiple correlation anomaly evaluation indexes according to the multiple index fusion weights to determine the first correlation anomaly evaluation index; performing secondary weighted fusion on the first anomaly evaluation index and the first correlation anomaly evaluation index to obtain the anomaly evaluation index, where the fusion weight of the first anomaly evaluation index is 0.7 and the fusion weight of the first correlation anomaly evaluation index is 0.3.
[0070] Further, the user information sharing system based on big data is also used for: obtaining the anomaly evaluation index, where the anomaly evaluation index includes multiple anomaly evaluation attribute data; randomly selecting the first evaluation information in the evaluation information set, and respectively performing similarity analysis on the first evaluation information according to the multiple anomaly evaluation attribute data to obtain a first similarity set, and calculating the mean value to obtain a first similarity coefficient; subtracting the first similarity coefficient from 1 to obtain a first data credibility coefficient, and sequentially analyzing the multiple data credibility coefficients of the multiple evaluation information in the evaluation information set.
[0071] Further, the user information sharing system based on big data is also used for: based on big data, obtaining the false evaluation information quantity sequence and false evaluation information proportion sequence of the target platform in a preset historical time zone, as well as the multiple false evaluation information quantity sequences and multiple false evaluation information proportion sequences of multiple similar platforms in a preset time zone; pre-training an evaluation trend analysis plug-in, and inputting the false evaluation information quantity sequence, false evaluation information proportion sequence, multiple platform correlation degrees, multiple false evaluation information quantity sequences and multiple false evaluation information proportion sequences of multiple similar platforms into the evaluation trend analysis plug-in to output an evaluation trend coefficient; optimizing and adjusting a predetermined credibility threshold according to the evaluation trend coefficient to obtain a dynamic credibility threshold; judging the multiple data credibility coefficients according to the dynamic credibility threshold, and setting the evaluation information corresponding to the data credibility coefficient greater than the dynamic credibility threshold as reliable evaluation data to obtain a reliable evaluation data set.
[0072] Further, the user information sharing system based on big data is further configured to: collect a sample training data set based on big data, wherein the sample training data includes a sample false evaluation information quantity sequence, a sample false evaluation information proportion sequence, a sample platform association degree set, a sample same-kind false evaluation information quantity sequence set, a sample same-kind false evaluation information proportion sequence set, and a sample evaluation trend coefficient, and the sample evaluation trend coefficient can be obtained through manual annotation; use the sample training data set to perform supervised training and testing on a BP neural network until the loss function converges, so as to obtain a trained evaluation trend analysis plug-in.
[0073] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept.
[0074] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. A method for sharing user information based on big data, characterized in that, The method includes: Performing feature clustering analysis on the false evaluation information of the target platform to determine the first abnormal evaluation index; Based on big data, performing feature clustering analysis on the false evaluation information of multiple similar platforms of the target platform to determine multiple associated abnormal evaluation indexes; Fusing and determining the abnormal evaluation index according to the first abnormal evaluation index and the multiple associated abnormal evaluation indexes; Performing data credibility analysis on the evaluation information set of the target platform according to the abnormal evaluation index, setting the evaluation information that meets the dynamic credibility threshold as reliable evaluation data, obtaining a reliable evaluation data set, and sharing the user evaluation information of the target platform; Among them, setting the evaluation information that meets the dynamic credibility threshold as reliable evaluation data and obtaining a reliable evaluation data set includes: Based on big data, obtaining the false evaluation information quantity sequence and false evaluation information proportion sequence of the target platform in a preset historical time zone, as well as the multiple similar false evaluation information quantity sequences and multiple similar false evaluation information proportion sequences of multiple similar platforms in a preset time zone; Pre-training an evaluation trend analysis plugin, and inputting the false evaluation information quantity sequence, false evaluation information proportion sequence, multiple platform correlation degrees, multiple similar false evaluation information quantity sequences, and multiple similar false evaluation information proportion sequences into the evaluation trend analysis plugin to output an evaluation trend coefficient; Optimally adjusting the predetermined credibility threshold according to the evaluation trend coefficient to obtain a dynamic credibility threshold; Judging multiple data credibility coefficients according to the dynamic credibility threshold, and setting the evaluation information corresponding to the data credibility coefficient greater than the dynamic credibility threshold as reliable evaluation data to obtain a reliable evaluation data set; Among them, pre-training the evaluation trend analysis plugin includes: Based on big data, collecting a sample training data set, where the sample training data includes a sample false evaluation information quantity sequence, a sample false evaluation information proportion sequence, a sample platform correlation degree set, a sample similar false evaluation information quantity sequence set, a sample similar false evaluation information proportion sequence set, and a sample evaluation trend coefficient, and the sample evaluation trend coefficient can be obtained through manual annotation; Using the sample training data set to perform supervised training and testing on a BP neural network until the loss function converges to obtain a trained evaluation trend analysis plugin.
2. The method for sharing user information based on big data according to claim 1, wherein Performing feature clustering analysis on the false evaluation information of the target platform to determine the first abnormal evaluation index, including: Obtaining the historical false evaluation information set of the target platform, where the false evaluation information includes false positive evaluation data and malicious negative evaluation data; Performing feature extraction and clustering analysis on the historical false positive evaluation data set according to a predetermined evaluation feature attribute to obtain a first false positive evaluation index; Performing feature extraction and clustering analysis on the historical malicious negative evaluation data set according to a predetermined evaluation feature attribute to obtain a first malicious negative evaluation index; Taking the first false positive evaluation index and the first malicious negative evaluation index as the first abnormal evaluation index.
3. A method for sharing user information based on big data according to claim 2, characterized in that, Performing feature extraction and clustering analysis on the historical false positive evaluation data set according to a predetermined evaluation feature attribute to obtain a first false positive evaluation index, including: Obtain predetermined evaluation feature attributes, where the predetermined evaluation feature attributes at least include the number of evaluation words, keywords, modal particles, the number of pictures, and the picture quality; According to the predetermined evaluation feature attributes, respectively extract features from multiple historical false positive review data in the historical false positive review dataset to obtain multiple evaluation feature datasets; Use the K-means clustering algorithm to perform feature clustering on the multiple evaluation feature datasets to obtain the first false positive review index.
4. A method for sharing user information based on big data according to claim 2, characterized in that, Based on big data, perform feature clustering analysis on false evaluation information of multiple similar platforms of the target platform to determine multiple associated abnormal evaluation indexes, including: With the platform attribute characteristics of the target platform as a constraint, perform associated retrieval based on big data to determine multiple similar platforms, and obtain multiple sets of similar historical false evaluation information of the multiple similar platforms, where the similar historical false evaluation information includes similar false positive review data and similar malicious negative review data; According to the predetermined evaluation feature attributes, respectively perform feature extraction and clustering analysis on multiple sets of similar historical false positive review datasets and similar historical malicious negative review datasets to obtain multiple associated abnormal evaluation indexes.
5. A method for sharing user information based on big data according to claim 4, characterized in that, Fuse and determine the abnormal evaluation index according to the first abnormal evaluation index and multiple associated abnormal evaluation indexes, including: Based on the platform attribute characteristics, respectively perform association degree analysis on the target platform and multiple similar platforms to determine multiple platform association degrees; Calculate the ratio of each platform association degree to the sum of multiple platform association degrees respectively, set it as the index fusion weight, and obtain multiple index fusion weights; Perform weighted fusion on the multiple associated abnormal evaluation indexes according to the multiple index fusion weights to determine the first associated abnormal evaluation index; Perform secondary weighted fusion on the first abnormal evaluation index and the first associated abnormal evaluation index to obtain the abnormal evaluation index, where the fusion weight of the first abnormal evaluation index is 0.7 and the fusion weight of the first associated abnormal evaluation index is 0.
3.
6. A method for sharing user information based on big data according to claim 5, characterized in that Perform data credibility analysis on the evaluation information set of the target platform according to the abnormal evaluation index, including: Obtain the abnormal evaluation index, where the abnormal evaluation index includes multiple abnormal evaluation attribute data; Randomly select the first evaluation information in the evaluation information set, and perform similarity analysis on the first evaluation information respectively according to the multiple abnormal evaluation attribute data to obtain the first similarity set, and calculate the mean value to obtain the first similarity coefficient; Subtract the first similarity coefficient from 1 to obtain the first data credibility coefficient, and sequentially analyze the multiple data credibility coefficients of multiple evaluation information in the evaluation information set.
7. A user information sharing system based on big data, characterized in that, Steps for implementing the method for sharing user information based on big data according to any one of claims 1 to 6, including: The first abnormal evaluation index determination module is used to perform feature clustering analysis on the false evaluation information of the target platform to determine the first abnormal evaluation index; The associated abnormal evaluation index determination module is used to perform feature clustering analysis on the false evaluation information of multiple similar platforms of the target platform based on big data to determine multiple associated abnormal evaluation indexes; An abnormal evaluation index fusion module, which is used to fuse and determine an abnormal evaluation index according to the first abnormal evaluation index and multiple associated abnormal evaluation indexes; A data credibility analysis module, which is used to perform data credibility analysis on the evaluation information set of the target platform according to the abnormal evaluation index, set the evaluation information that meets the dynamic credibility threshold as reliable evaluation data, obtain a reliable evaluation data set, and share the user evaluation information of the target platform.
Citation Information
Patent Citations
Cross-platform e-commerce fraud detection method and system based on comment data
CN109145187A