Merchant early warning monitoring method and system based on multi-dimensional data fusion and terminal equipment
By conducting abnormal detection and cluster analysis on the historical operation data and user evaluation text of historical merchants, a risk prediction model is established, and the problem of inconsistent multi-dimensional data format is solved, and the accuracy and reliability of merchant early warning analysis is improved.
Patent Information
- Application Number
- CN202510447186.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The multidimensional data collected in the prior art lacks a unified format, resulting in a significant reduction in the accuracy and reliability of merchant early warning analysis results.
By obtaining historical operation data and user evaluation text of historical merchants, conducting abnormal detection and clustering analysis, determining historical relationships, establishing a risk prediction model, and using this model to conduct early warning and monitoring of target merchants.
Through multi-dimensional data fusion, the correlation between historical operation data and user evaluation text is accurately determined, which improves the accuracy of risk prediction, promptly discovers potential risk signals, and solves the problem of inconsistent multi-dimensional data formats.
Smart Images

Figure CN119963236A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data fusion, and in particular to a merchant early warning monitoring method, system and terminal device based on multi-dimensional data fusion. Background Art
[0002] In today's digital business era, merchants have the ability to collect multi-dimensional data during operations, and these data sources are extensive and rich. However, the lack of a unified format for the collected multi-dimensional data directly leads to the fragmentation and decentralization of the data. When faced with the critical task of early warning analysis, the differences in data dimensions make it difficult to apply existing early warning analysis methods, which greatly reduces the accuracy and reliability of early warning analysis results. Summary of the invention
[0003] The main purpose of the embodiments of the present invention is to provide a merchant early warning monitoring method, system and terminal device based on multi-dimensional data fusion, aiming to solve the problem that the multi-dimensional data collected in the related technology lacks a unified format, thereby greatly reducing the accuracy and reliability of the early warning analysis results.
[0004] In a first aspect, an embodiment of the present invention provides a merchant early warning monitoring method based on multi-dimensional data fusion, comprising:
[0005] Obtaining historical operation data and user evaluation text corresponding to historical merchants, and performing anomaly detection on the historical operation data to obtain first anomaly data and a first anomaly distribution corresponding to the first anomaly data;
[0006] Determine a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution;
[0007] Performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data;
[0008] Determine a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution;
[0009] Determine, according to the first description text and the second description text, a historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text;
[0010] Establishing a risk prediction model based on the historical operation data, the user evaluation text and the historical association relationship;
[0011] The current operation data and current evaluation text corresponding to the target merchant are obtained, and the target merchant is monitored for early warning according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result.
[0012] In a second aspect, an embodiment of the present invention provides a merchant early warning monitoring system based on multi-dimensional data fusion, including:
[0013] A data collection module, used to obtain historical operation data and user evaluation text corresponding to historical merchants, and perform anomaly detection on the historical operation data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data;
[0014] A first generating module, configured to determine a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution;
[0015] An abnormality analysis module, used for performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data;
[0016] A second generating module, configured to determine a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution;
[0017] A relationship determination module, configured to determine, according to the first description text and the second description text, a historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text;
[0018] A model building module, used to build a risk prediction model based on the historical operation data, the user evaluation text and the historical association relationship;
[0019] The early warning monitoring module is used to obtain the current operating data and current evaluation text corresponding to the target merchant, and to perform early warning monitoring on the target merchant according to the risk prediction model in combination with the current operating data and the current evaluation text to obtain the target monitoring result.
[0020] In the third aspect, an embodiment of the present invention further provides a terminal device, comprising a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of any one of the merchant early warning monitoring methods based on multi-dimensional data fusion provided in the specification of the present invention are realized.
[0021] The embodiment of the present invention provides a merchant early warning monitoring method, system and terminal device based on multi-dimensional data fusion, the method comprising: obtaining historical operation data and user evaluation text corresponding to historical merchants, and performing anomaly detection on the historical operation data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data; determining a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution; performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data; determining a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution; determining the historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text according to the first description text and the second description text; establishing a risk prediction model according to the historical operation data, the user evaluation text and the historical association relationship; obtaining current operation data and current evaluation text corresponding to the target merchant, and performing early warning monitoring on the target merchant according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result. The method determines the first description text corresponding to the historical operation data of the historical merchant and the second description text corresponding to the user evaluation text, thereby accurately determining the historical correlation between the historical operation data and the user evaluation text based on the first description text and the second description text, thereby revealing the intrinsic connection between the operation data and the user evaluation, and then providing richer information support for decision-making, thereby combining the historical operation data, user evaluation text and historical correlation to establish a risk prediction model, which can more accurately predict the risks that merchants may face, and then use the risk prediction model to conduct early warning monitoring of the current operation data and current evaluation text of the target merchant, and can timely discover potential risk signals. It also solves the problem that the multi-dimensional data collected in the related technology lacks a unified format, thereby greatly reducing the accuracy and reliability of the early warning analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 A flowchart of a merchant early warning monitoring method based on multi-dimensional data fusion provided by an embodiment of the present invention;
[0024] Figure 2 A schematic diagram of the module structure of a merchant early warning monitoring system based on multi-dimensional data fusion provided by an embodiment of the present invention;
[0025] Figure 3A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0028] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0029] The embodiment of the present invention provides a merchant early warning monitoring method, system and terminal device based on multi-dimensional data fusion. The merchant early warning monitoring method based on multi-dimensional data fusion can be applied to a terminal device, which can be an electronic device such as a tablet computer, a laptop computer, a desktop computer, a personal digital assistant and a wearable device. The terminal device can be a server or a server cluster.
[0030] Some embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0031] Please refer to Figure 1 , Figure 1 A flowchart of a merchant early warning monitoring method based on multi-dimensional data fusion provided in an embodiment of the present invention.
[0032] like Figure 1 As shown, the merchant early warning monitoring method based on multi-dimensional data fusion includes steps S101 to S107.
[0033] Step S101: Obtain historical operating data and user evaluation text corresponding to historical merchants, and perform anomaly detection on the historical operating data to obtain first anomaly data and a first anomaly distribution corresponding to the first anomaly data.
[0034] Exemplarily, historical operation data corresponding to historical merchants and user evaluation texts of consumers after historical consumption are obtained from the database. The historical operation data is the customer flow, inventory turnover rate, etc. corresponding to historical merchants at different times. The historical operation data is a discrete data type, and the user evaluation text is a text data type.
[0035] Exemplarily, statistical analysis is performed on historical operating data based on statistical methods to obtain the degree of deviation corresponding to each sub-operating data in the historical operating data, thereby determining the abnormal value corresponding to the sub-operating data according to the degree of deviation, and then obtaining the first abnormal data from the historical operating data according to the abnormal value and the preset value.
[0036] Exemplarily, a statistical analysis is performed on the first abnormal data, for example, the distribution of the first abnormal data in different time periods is calculated to obtain a first abnormal distribution corresponding to the first abnormal data.
[0037] In some embodiments, the anomaly detection on the historical operating data to obtain first anomaly data and a first anomaly distribution corresponding to the first anomaly data includes: determining a target window, and segmenting the historical operating data according to the target window to obtain a plurality of segmented operating data; determining time information corresponding to each first sub-operating data in the segmented operating data, and determining a loss mapping value corresponding to the first sub-operating data according to the time information; determining a window center corresponding to the segmented operating data according to the loss mapping value and the number of data corresponding to the first sub-operating data in the segmented operating data; filtering out a target center associated with each second sub-operating data in the historical operating data from the window center; and filtering out a target center associated with each second sub-operating data in the historical operating data according to the target center and the second sub-operating data. The distance information between the operation data determines the first distribution characterization value corresponding to the second sub-operation data, and the first distribution characterization value is used to represent the distribution status corresponding to the second sub-operation data; the nearest center corresponding to the second sub-operation data is screened out from the window center, and the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operation data are determined according to the nearest center and the target center combined with the window width corresponding to the target window; the second sub-operation data is subjected to anomaly detection according to the first maximum distribution value and the first minimum distribution value combined with the first distribution characterization value to obtain the first abnormal data; and the historical operation data is subjected to data statistics according to the first abnormal data to obtain the first abnormal distribution corresponding to the first abnormal data.
[0038] For example, a suitable time window is determined as the target window according to business needs. For example, the target window can be one week, half a month, or one month, etc., so that the historical operation data is segmented in chronological order according to the target window to obtain multiple segmented operation data. There is no data overlap between the segmented operation data.
[0039] Exemplarily, the time information corresponding to the first sub-operation data in each segmented operation data is obtained from the database, and then the base corresponding to the exponential function is determined, and the base is between 0 and 1, and then the time information corresponding to the first sub-operation data in the segmented operation data is determined as the reference time, so as to calculate the difference information between the time information and the reference time, and then substitute the difference information into the exponential function to obtain the loss mapping value corresponding to the first sub-operation data. The loss mapping value is used to characterize the degree of difference in the acquisition time of different first sub-operation data under the same target window.
[0040] Exemplarily, the number of data corresponding to each first sub-operation data in the split operation data is counted to obtain the number of data corresponding to the first sub-operation data in the split operation data. According to the data characteristics of the historical operation data, it is known that it is very difficult to have completely equal data in the split operation data. Therefore, when counting the number of data corresponding to each first sub-operation data in the split operation data, the present application considers that the other data is the same data as the first sub-operation data when the data gap between other data in the split operation data and the first sub-operation data is within a preset gap range.
[0041] Exemplarily, the distribution weight corresponding to the first sub-operation data is determined according to the amount of data, and the first data is obtained by multiplying the distribution weight and the loss mapping value and then multiplying the resultant value by the first sub-operation data. Furthermore, the first data corresponding to all the first sub-operation data are summed up and determined as the window center corresponding to the split operation data.
[0042] Exemplarily, each sub-data in the historical operation data is determined as the second sub-operation data, and then the center distance between the second sub-operation data and each window center is calculated according to calculation methods such as Euclidean distance and Manhattan distance, so as to compare the center distance with the preset distance, and then determine the window center whose center distance is less than or equal to the preset distance as the target center corresponding to the second sub-operation data.
[0043] Exemplarily, the distance information between the second sub-operation data and each target center is calculated according to calculation methods such as Euclidean distance and Manhattan distance, and the segmented operation data corresponding to the target center is determined, and then the number of targets corresponding to the target center in the segmented operation data corresponding to it is determined, wherein the number of targets is calculated taking into account the great difficulty of having data that is completely equal to the target center in the segmented operation data. Therefore, when counting the number of targets corresponding to each target center in the segmented operation data corresponding to it, the present application considers that when the data gap between other sub-data in the segmented operation data corresponding to the target center and the target center is within a preset gap range, the other sub-data is considered to be the same data as the target center, thereby determining the number of targets corresponding to the target center in the segmented operation data corresponding to it based on the data gap between other sub-data and the target center and the preset gap range.
[0044] Exemplarily, the target kernel function is determined, and then the target value corresponding to the distance information under the target kernel function is determined, so as to obtain the first distribution characterization value corresponding to the second sub-operation data according to the following formula, wherein the first distribution characterization value is used to represent the distribution status corresponding to the second sub-operation data.
[0045] ;
[0046] in, represents the first distribution representation value corresponding to the t-th second sub-operation data, represents the number of the target centers corresponding to the t-th second sub-operation data, represents the number of targets corresponding to the kth target center in the corresponding split operation data corresponding to the tth second sub-operation data, represents the number of targets corresponding to the yth target center in the corresponding split operation data corresponding to the tth second sub-operation data, Represents the target kernel function, which can be a Gaussian kernel function or an improved function of the Gaussian kernel function. Represents the distance information between the tth second sub-operation data and the kth target center.
[0047] Exemplarily, the window center corresponding to when the distance information between the window center and the second sub-operation data is minimum is determined as the nearest center corresponding to the second sub-operation data, and then based on the nearest center and the target center, combined with the window width corresponding to the target window, the range is adjusted based on the nearest center and the window width and the target center to obtain the boundary value corresponding to the first distribution characterization value, that is, the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operation data are obtained.
[0048] Exemplarily, the first distribution characterization value of each second sub-operation data is compared with its corresponding first maximum distribution value and first minimum distribution value. If the first distribution characterization value exceeds the first maximum distribution value or is lower than the first minimum distribution value, the second sub-operation data is determined to be abnormal data, and then these abnormal data are aggregated to obtain the first abnormal data.
[0049] Exemplarily, the first abnormal data is statistically analyzed according to the time dimension under the historical operation data to determine the first abnormal distribution corresponding to the first abnormal data. The first abnormal distribution is used to characterize the time point or time distribution information of the first abnormal data in the historical operation data.
[0050] Specifically, the time information of the first sub-operation data is determined and the loss mapping value is calculated. The loss mapping value can be used to quantify the deviation of the data from the business goal in the time dimension, so as to more accurately characterize the normal distribution range of the data by calculating the window center, screening the target center and the nearest center, and combining the distance information to determine the distribution characterization value and the maximum distribution value and the minimum distribution value. This anomaly detection method based on multi-dimensional information can more accurately identify real abnormal data compared to simple threshold judgment or single statistical indicator methods, thereby providing good support for subsequent early warning monitoring.
[0051] In some embodiments, the determining of the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operation data according to the nearest center and the target center in combination with the window width corresponding to the target window includes: determining the center distance between the nearest center and the target center, and when the center distance is greater than the window width, determining that the first parameter corresponding to the first minimum distribution value is the data difference between the center distance and the window width; when the center distance is less than or equal to the window width, determining that the first parameter corresponding to the first minimum distribution value is preset data; determining the second parameter corresponding to the first maximum distribution value according to the center distance and the window width; determining the segmented operation data corresponding to the target center as the associated operation data corresponding to the target center, and determining the quantity information corresponding to the target center in the associated operation data; determining the target kernel function, and determining the first minimum distribution value corresponding to the second sub-operation data according to the target kernel function and the first parameter in combination with the target center and the quantity information; determining the first maximum distribution value corresponding to the second sub-operation data according to the target kernel function and the second parameter in combination with the target center and the quantity information; wherein the first minimum distribution value and the first maximum distribution value are obtained according to the following formulas:
[0052] ;
[0053] ;
[0054] in, represents the first minimum distribution value corresponding to the i-th second sub-operation data, represents the first maximum distribution value corresponding to the i-th second sub-operation data, represents the number of the target centers corresponding to the i-th second sub-operation data, represents the quantity information corresponding to the jth target center corresponding to the i-th second sub-operation data, represents the quantity information corresponding to the mth target center corresponding to the i-th second sub-operation data, represents the target kernel function, The first parameter representing the distance between the nearest center corresponding to the i-th second sub-operation data and the j-th target center and the window width; The second parameter corresponding to the center distance between the nearest center corresponding to the i-th second sub-operation data and the j-th target center and the window width.
[0055] Exemplarily, the center distance between the nearest center and the target center is calculated according to the Euclidean distance calculation formula, and then the center distance and the window width are compared. When the center distance is greater than the window width, the window width is subtracted from the center distance to obtain the data difference, and the data difference is determined as the first parameter required for the subsequent calculation of the first minimum distribution value; when the center distance is less than or equal to the window width, the first parameter required for calculating the first minimum distribution value is determined as the preset data, where the preset data is equal to zero.
[0056] Exemplarily, the center distance and the window width are summed and the sum result is determined as the second parameter required for subsequent calculation of the first maximum distribution value.
[0057] Exemplarily, the segmented operation data corresponding to the target center is determined as the associated operation data corresponding to the target center, and then the quantity information corresponding to the target center in the associated operation data is counted, wherein, when calculating the target quantity, it is considered that there is data in the associated operation data that is completely equal to the target center, which has great difficulty. Therefore, when counting the quantity information corresponding to each target center in the associated operation data corresponding to it, the present application considers that when the data gap between other sub-data in the associated operation data corresponding to the target center and the target center is within a preset gap range, the other sub-data is considered to be the same data as the target center, thereby determining the quantity information corresponding to the target center in the associated operation data corresponding to it based on the data gap between the other sub-data and the target center and the preset gap range.
[0058] Exemplarily, the target kernel function is determined as needed, and the target kernel function may be a Gaussian kernel function or an improved function of the Gaussian kernel function, and then the first minimum distribution value corresponding to the second sub-operation data is determined according to the following formula in combination with the target kernel function, the first parameter, the target center, and the quantity information:
[0059] ;
[0060] in, represents the first minimum distribution value corresponding to the i-th second sub-operation data, represents the number of target centers corresponding to the i-th second sub-operation data, Indicates the quantity information corresponding to the jth target center corresponding to the i-th second sub-operation data, Indicates the quantity information corresponding to the mth target center corresponding to the i-th second sub-operation data, represents the target kernel function, Represents the first parameter corresponding to the center distance between the nearest center corresponding to the i-th second sub-operation data and the j-th target center and the window width.
[0061] Exemplarily, the first maximum distribution value corresponding to the second sub-operation data is determined according to the following formula in combination with the target kernel function, the second parameter, the quantity information, and the target center:
[0062] ;
[0063] in, represents the first maximum distribution value corresponding to the i-th second sub-operation data, represents the number of target centers corresponding to the i-th second sub-operation data, Indicates the quantity information corresponding to the jth target center corresponding to the i-th second sub-operation data, Indicates the quantity information corresponding to the mth target center corresponding to the i-th second sub-operation data, represents the target kernel function, Represents the second parameter corresponding to the center distance between the nearest center corresponding to the i-th second sub-operation data and the j-th target center and the window width.
[0064] Specifically, the present application maps the original data to a high-dimensional feature space through a target kernel function (such as a Gaussian kernel function or its improved function), thereby making the characteristics of the data richer and more distinctive, and uses the quantity information of the target center corresponding to the second sub-operation data to adjust the mapping data corresponding to the target kernel function to obtain an accurate first maximum distribution value and a first minimum distribution value, thereby providing reliable support for subsequent anomaly identification.
[0065] In some embodiments, the method of performing anomaly detection on the second sub-operation data according to the first maximum distribution value and the first minimum distribution value in combination with the first distribution characterization value to obtain the first anomaly data includes: determining the nearest neighbor data corresponding to the second sub-operation data from the historical operation data, and determining the second distribution characterization value corresponding to the nearest neighbor data; determining the first loss value corresponding to the second sub-operation data and the second loss value corresponding to the nearest neighbor data; determining the third distribution characterization value corresponding to the second sub-operation data according to the first distribution characterization value and the second distribution characterization value in combination with the first loss value and the second loss value; determining the second minimum distribution value and the second maximum distribution value corresponding to the nearest neighbor data, and determining the target threshold value according to the first minimum distribution value, the first maximum distribution value, the second minimum distribution value and the second maximum distribution value; performing anomaly detection on the second sub-operation data according to the target threshold value and the third distribution characterization value to obtain the first anomaly data.
[0066] Exemplarily, in the historical operation data, the data closest to each second sub-operation data is found based on a distance measurement method (such as Euclidean distance, Manhattan distance, etc.), and these data are the nearest neighbor data corresponding to the second sub-operation data.
[0067] Exemplarily, for each nearest neighbor data, the second distribution characterization value corresponding to each nearest neighbor data is calculated in the same manner as the first distribution characterization value is determined previously.
[0068] Exemplarily, the first loss value corresponding to the second sub-operation data and the second loss value corresponding to the nearest neighbor data are determined in the same manner as the loss mapping value corresponding to each second sub-operation in the split operation data.
[0069] Exemplarily, the first distribution characterization value and the second distribution characterization value are weightedly summed according to the first loss value and the second loss value to determine the third distribution characterization value corresponding to the second sub-operation data.
[0070] Exemplarily, for each nearest neighbor data, the second minimum distribution value and the second maximum distribution value corresponding to each nearest neighbor data are determined according to the method for determining the first minimum distribution value and the first maximum distribution value.
[0071] Exemplarily, the first window position corresponding to each nearest neighbor data in the split operation data and the second window position corresponding to each second sub-operation data are obtained, and then the integers from 1 to the first window position are summed to obtain a first value, and then the first window position and the first value are divided to obtain a first weight corresponding to the first sub-operation data, and the second value is obtained by summing the integers from 1 to the second window position, and then the second window position and the second value are divided to obtain a second weight corresponding to the second sub-operation data.
[0072] Exemplarily, the first minimum distribution value and the first maximum distribution value are averaged to obtain a first average value, and the second minimum distribution value and the second maximum distribution value are averaged to obtain a second average value, thereby weightedly summing the first average value and the second average value according to the first weight and the second weight to obtain the target threshold.
[0073] Exemplarily, the third distribution characterization value of each second sub-operation data is compared with the target threshold, and if the third distribution characterization value exceeds the target threshold, the second sub-operation data is determined to be abnormal data. All second sub-operation data determined to be abnormal are aggregated to obtain first abnormal data.
[0074] Step S102: Determine a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution.
[0075] Exemplarily, the adjacent data corresponding to the first abnormal data is obtained from the historical operation data, so as to determine the change trend corresponding to the first abnormal data based on the adjacent data, wherein the change trend is such as an upward trend, a downward trend, a fluctuating trend or a stable trend, etc., so as to determine the first description text corresponding to the historical merchant based on the change trend in combination with the abnormal time corresponding to the first abnormal distribution using the sentence generation rule. For example, it can be stipulated that the abnormal time corresponding to the first abnormal distribution is described first, and then the change trend is explained. For example, according to the sentence generation rule, the determined descriptive elements are organized into text. For example: "During the [abnormal time corresponding to the first abnormal distribution], the [business indicators] of the historical merchant showed a [changing trend], which was highly correlated with the time when the abnormality occurred, which had a [degree of impact] on the [business aspects] of the merchant."
[0076] Step S103: performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data.
[0077] Exemplarily, cluster analysis is performed on user evaluation texts using a clustering algorithm such as k-means clustering to obtain text clustering results corresponding to the user evaluation texts, and then text rational classification is performed on each subclass cluster in the text clustering results to obtain the text type corresponding to the subclass cluster, and then the subclass cluster when the text type is the user negative review type is determined as the second abnormal data, thereby obtaining the time information corresponding to the second abnormal data and then determining the second abnormal distribution corresponding to the second abnormal data based on the time information.
[0078] In some embodiments, the clustering analysis of the user evaluation text to obtain the second abnormal data and the second abnormal distribution corresponding to the second abnormal data includes: using multiple clustering algorithms to perform data clustering on the user evaluation text to obtain multiple initial clustering results; determining a first evaluation text from the user evaluation text, and obtaining remaining evaluation texts after excluding the first evaluation text from the user evaluation text; calculating a first similarity between each second evaluation text in the remaining evaluation text and the first evaluation text, and determining a second similarity corresponding to the first evaluation text and the remaining evaluation text based on the first similarity; respectively obtaining relevant clusters corresponding to the first evaluation text from the initial clustering results, and determining the number of texts corresponding to the relevant clusters; and calculating a first similarity between the first evaluation text and the remaining evaluation text based on the second similarity and the text; This quantity determines the text weight corresponding to the first evaluation text; determines the cluster weight corresponding to each first subclass cluster in the initial clustering result according to the text weight; calculates the classification similarity between any two of the initial clustering results, and determines the cluster weight corresponding to the initial clustering result according to the classification similarity and the cluster weight; performs data clustering on the user evaluation text according to the text weight, the cluster weight and the cluster weight to obtain a target clustering result; performs sentiment analysis on each second subclass cluster in the target clustering result to obtain a target sentiment type corresponding to the second subclass cluster; determines the second abnormal data from the second subclass cluster according to the target sentiment type; performs data statistics on the user evaluation text according to the second abnormal data to obtain the second abnormal distribution corresponding to the second abnormal data.
[0079] Exemplarily, multiple algorithms suitable for text clustering such as K-Means clustering, hierarchical clustering, DBSCAN, etc. are determined, and then the user evaluation texts are clustered using the selected multiple clustering algorithms respectively, thereby obtaining multiple initial clustering results.
[0080] Exemplarily, a first evaluation text is arbitrarily selected from the user evaluation texts, and then the first evaluation text is excluded from the user evaluation texts, and the remaining text is the remaining evaluation text. A second evaluation text is arbitrarily selected from the remaining evaluation texts, and a first similarity between the second evaluation text and the first evaluation text is calculated using a similarity calculation method such as cosine similarity, edit distance, etc.
[0081] Exemplarily, a first similarity between the first evaluation text and each second evaluation text in the remaining evaluation texts is obtained, and then all the first similarities are summed to obtain the corresponding second similarities between the first evaluation text and the remaining evaluation texts.
[0082] Exemplarily, in each initial clustering result, clusters containing the first evaluation text are found. These clusters are the relevant clusters corresponding to the first evaluation text, and then the number of texts contained in each relevant cluster is counted, so that the total number of all texts is summed up to obtain the total number, and then the second similarity is divided by the total number to determine the text weight corresponding to the first evaluation text. In this step, it is considered that the clustering results corresponding to different clustering methods may be different, and there may be an imbalance in the number of clusters in each clustering result, and then the bias information on the cluster size is eliminated according to the number of texts contained in each relevant cluster.
[0083] Exemplarily, the text weight corresponding to each sub-text data in each first sub-class cluster in the initial clustering result is obtained, and then the text weight corresponding to each sub-text data in the first sub-class cluster is summed to obtain the cluster weight corresponding to the first sub-class cluster.
[0084] Exemplarily, an initial clustering result is arbitrarily determined from multiple initial clustering results as the current clustering result, and then the cluster weight corresponding to each second subclass cluster in the current clustering result is determined, and then the cluster weights corresponding to all second subclass clusters are summed to obtain the weight and value corresponding to the current clustering result, and then the classification similarity between the current clustering result and any one of the remaining initial clustering results is calculated using algorithms such as the Rand Index and the Adjusted Rand Index, so as to obtain the classification similarity between the current clustering result and each of the remaining initial clustering results, and then all the classification similarities are summed and then multiplied by the weight sum to obtain the cluster weight corresponding to the current clustering result, and then the above steps are performed on each initial clustering result to obtain the cluster weight corresponding to each initial clustering result.
[0085] Exemplarily, the user evaluation text is re-clustered according to the text weight, cluster weight and cluster weight. For example, based on the original clustering, the initial clustering result is adjusted and optimized according to the text weight, cluster weight and cluster weight to obtain the target clustering result.
[0086] Exemplarily, sentiment analysis methods such as dictionary-based sentiment analysis, machine learning sentiment analysis, etc. are used to perform sentiment analysis on each second subclass cluster in the target clustering result to determine the target sentiment type corresponding to each second subclass cluster, such as positive, negative, neutral, etc. Thus, according to the target sentiment type, texts with negative target sentiment type are screened out from the second subclass cluster as the second abnormal data.
[0087] Exemplarily, data statistics are performed on the second abnormal data in the user evaluation text, and the distribution of the second abnormal data in different aspects (such as time, subject, evaluation source, etc.) is analyzed to obtain a second abnormal distribution corresponding to the second abnormal data.
[0088] In some embodiments, the data clustering of the user evaluation text according to the text weights, the cluster weights and the clustering weights to obtain a target clustering result includes: constructing a first diagonal matrix and a first column matrix according to the text weights, and determining a first text matrix according to the first diagonal matrix and the first column matrix; constructing a second diagonal matrix according to the cluster weights, and determining a second text matrix according to the first text matrix and the second diagonal matrix; constructing a third diagonal matrix according to the clustering weights, and determining a third text matrix according to the second text matrix and the third diagonal matrix; determining a target supply-cooperation matrix corresponding to the user evaluation text according to the third text matrix and the transpose of the third text matrix; and data clustering the user evaluation text according to the target supply-cooperation matrix to obtain the target clustering result.
[0089] Exemplarily, first, the text weights are arranged in order. The text weight is the weight value corresponding to each user evaluation text calculated previously. With these text weight values as diagonal elements, a diagonal matrix is constructed, which is called the first diagonal matrix. The non-diagonal elements of the first diagonal matrix are all 0, and the elements on the diagonal correspond to the text weights of each user evaluation text in turn. And the user evaluation texts are sorted in order to obtain a column matrix to obtain the corresponding first column matrix.
[0090] Exemplarily, the first diagonal matrix is multiplied by the first column matrix by matrix multiplication. The matrix multiplication rule is: the row elements of the first matrix are multiplied by the column elements of the second matrix correspondingly and then summed to obtain the elements at the corresponding positions in the result matrix. The matrix obtained after multiplication is the first text matrix.
[0091] Exemplarily, these weight values are arranged in order according to the previously calculated cluster weights. A second diagonal matrix is constructed with the cluster weight values as diagonal elements, and the non-diagonal elements are also 0. This diagonal matrix reflects the importance weight of each cluster, thereby obtaining a second diagonal matrix.
[0092] Exemplarily, the first text matrix is multiplied by the second diagonal matrix. Through matrix multiplication, each element in the first text matrix is further weighted and adjusted according to the cluster weight to obtain the second text matrix. This step takes into account the influence of the importance of the cluster on the text matrix.
[0093] Exemplarily, these weight values are arranged in order according to the previously calculated cluster weights. A third diagonal matrix is constructed with the cluster weight values as diagonal elements and the non-diagonal elements as 0. This matrix reflects the importance of each clustering result, and then the second text matrix is multiplied by the third diagonal matrix. Through this matrix multiplication, the text weight, cluster weight and cluster weight are comprehensively considered, and the second text matrix is finally weighted and adjusted to obtain the third text matrix.
[0094] Exemplarily, according to the rule of matrix transposition, the rows and columns of the third text matrix are interchanged to obtain its transposed matrix, and then the third text matrix is multiplied by the transposed matrix of the third text matrix to obtain the target supply-cooperation matrix corresponding to the user evaluation text. The target supply-cooperation matrix reflects the relationship between the correlation and weighted comprehensive influences between the user evaluation texts.
[0095] Exemplarily, a clustering algorithm such as a spectral clustering algorithm is used to perform clustering operations using the eigenvalues and eigenvectors of the target supply-cooperation matrix. The target supply-cooperation matrix can provide similarity information between data, so that the target supply-cooperation matrix is used as input data, and the spectral clustering algorithm is used to cluster the user evaluation texts. The clustering algorithm will then divide the user evaluation texts into different clusters according to the relationships between the texts reflected in the matrix, and finally obtain the target clustering results.
[0096] Specifically, the cluster weights reflect the importance of different clusters in the overall clustering structure. By constructing the second diagonal matrix and multiplying it with the first text matrix, the text matrix can be further adjusted according to the importance of the clusters. The cluster weights take into account the reliability and importance of the results obtained by different clustering algorithms. By constructing the third diagonal matrix and multiplying it with the second text matrix, the advantages of different clustering algorithms are integrated. Different clustering algorithms may divide data from different perspectives. Comprehensively considering the cluster weights can make full use of the advantages of various algorithms, avoid the limitations of a single algorithm, and improve the accuracy and comprehensiveness of the clustering results. Therefore, by multiplying the third text matrix with its transposed matrix, the target supply-cooperation matrix is obtained, which can effectively capture the correlation between user evaluation texts. The target supply-cooperation matrix reflects the similarity and mutual relationship between texts. This relationship is the result of comprehensive consideration of text weights, cluster weights, and cluster weights, and is more comprehensive and accurate than simple text similarity calculation.
[0097] Step S104: Determine a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution.
[0098] Exemplarily, an abnormality type analysis is performed on the second abnormal data to obtain a target abnormality type corresponding to each second abnormal data, the target abnormality type includes but is not limited to commodity abnormalities, service abnormalities, and after-sales abnormalities, and the second abnormal data is classified by abnormality degree to obtain a target abnormality degree corresponding to each second abnormal data, the target abnormality degree includes but is not limited to minor abnormalities and severe abnormalities.
[0099] Exemplarily, the second abnormal data is statistically analyzed according to the target abnormality type and the target abnormality degree to obtain high-frequency abnormal words corresponding to the second abnormal data, and the time information of the occurrence of the high-frequency abnormal words is determined according to the second abnormal distribution, and then the corresponding second description text is obtained according to the high-frequency abnormal words and the abnormal occurrence time corresponding to the second abnormal data and the second abnormal distribution according to the text generation rules.
[0100] For example, the text generation rule is: "Under [target anomaly type], the proportion of [high-frequency anomaly words] reported by recent customers has increased significantly, and generally occurs mainly under [anomaly occurrence time]."
[0101] Step S105: determining the historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text according to the first description text and the second description text.
[0102] Exemplarily, the historical operation data and the user evaluation text are converted to the same dimension to obtain the corresponding first description text and second description text respectively, and then the text similarity value between the first description text and the second description text is calculated according to the text similarity algorithm, so that when the text similarity value meets the preset value, it is determined that there is an association between the historical operation data and the user evaluation text, and the historical association relationship corresponding to the historical operation data and the user evaluation text is determined to be positively correlated; however, when the text similarity value does not meet the preset value, it means that the direct similarity between the historical operation data and the user evaluation text is low, and their association relationship cannot be simply determined. Further in-depth analysis is required at this time. According to the first description text, the first change trend corresponding to the historical merchant under the first abnormal data is determined, and this change trend can be divided into two situations of getting better or worse. Similarly, for the user evaluation text, the second change trend corresponding to the historical merchant under the second abnormal data is determined according to the second description text, and then the historical association relationship between the historical operation data and the user evaluation text is determined by comparing the first change trend and the second change trend. If the first change trend and the second change trend are the same, that is, both are either getting better or worse, then it can be determined that the historical association relationship corresponding to the historical operation data and the user evaluation text is positively correlated. On the contrary, when the first change trend and the second change trend are opposite, it is determined that the corresponding historical correlation relationship between the historical operation data and the user evaluation text is negatively correlated.
[0103] In some embodiments, determining the historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text based on the first description text and the second description text includes: performing keyword recognition on the first description text to obtain a first keyword and performing keyword recognition on the second description text to obtain a second keyword; performing vector representation on the first keyword using a text representation model to obtain multiple first representation vectors and performing vector representation on the second keyword using the text representation model to obtain multiple second representation vectors; obtaining a first target vector corresponding to the first keyword from multiple first representation vectors through a maximum pooling operation and a neural network, and performing splicing processing based on the first target vector to obtain a first sentence vector corresponding to the first description text; obtaining a second target vector corresponding to the second keyword from multiple second representation vectors through the maximum pooling operation and the neural network, and performing splicing processing based on the second target vector to obtain a second sentence vector corresponding to the second description text. a second sentence vector; performing keyword similarity calculation on the first keyword and the second keyword through the first target vector and the second target vector to obtain a first vector result; performing sentence similarity calculation on the first description text and the second description text through the first sentence vector and the second sentence vector to obtain a second vector result; performing similarity calculation on the first keyword and the second description text according to the first target vector and the second sentence vector to obtain a third vector result; performing similarity calculation on the second keyword and the first description text according to the second target vector and the first sentence vector to obtain a fourth vector result; fusing the first vector result, the second vector result, the third vector result and the fourth vector result to determine the corresponding text similarity between the first description text and the second description text; determining a preset threshold, and determining the historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text according to the text similarity and the preset threshold.
[0104] Exemplarily, the TF-IDF algorithm is used to perform keyword recognition on the first description text to obtain the first keyword, and then the same method is used to analyze the second description text to identify the second keyword therein.
[0105] Exemplarily, a text representation model such as Word2Vec, GloVe, etc. is used to input the first keyword into the model, and the model converts each first keyword into a corresponding vector according to its pre-trained word vector library, thereby obtaining multiple first representation vectors. The second keyword is processed using the same text representation model, and each second keyword is converted into a vector to obtain multiple second representation vectors.
[0106] Exemplarily, a maximum pooling operation is performed on multiple first characterization vectors, that is, the maximum value is selected from the corresponding dimension of each first characterization vector to form a new vector. This new vector is then input into a neural network for further processing to obtain a first target vector corresponding to the first keyword, so that the first target vectors are spliced according to certain rules, such as splicing in sequence, and finally the first sentence vector corresponding to the first description text is obtained. Repeat the above steps, perform maximum pooling operations and neural network processing on multiple second characterization vectors, obtain a second target vector corresponding to the second keyword, and splice the second target vectors to obtain a second sentence vector corresponding to the second description text.
[0107] Exemplarily, the first target vector and the second target vector are multiplied to obtain a first result, and the first vector modulus corresponding to the first target vector and the second vector modulus corresponding to the second target vector are obtained, thereby multiplying the first vector modulus and the second vector modulus to obtain a second result, and then dividing the first result and the second result to obtain a third result, and then obtaining a third result between multiple first keywords and second keywords, thereby splicing the third results into a related vector, and then adjusting the related vector using the first weight matrix and the first deviation parameter between the keywords to obtain the target result between the first keyword and the second keyword, and then summing the target results using the sigma function to obtain the first vector result.
[0108] Exemplarily, the cosine similarity is used to calculate the similarity between the first sentence vector and the second sentence vector to obtain the sentence similarity, and then the first sentence vector and the second sentence vector are subjected to a dot product operation to obtain a first operation result, and then the first sentence vector and the second sentence vector are subjected to a vector difference calculation and the absolute value of the difference is calculated to obtain a second operation result, and then the first sentence vector and the second sentence vector are subjected to a dot addition operation and then adjusted using the second weight parameter and the second deviation parameter to obtain a third operation result, and finally, the first operation result, the second operation result, the third operation result and the fourth operation result are subjected to a dot addition operation and then adjusted using the corresponding third weight matrix and the third deviation matrix in the sentence vector processing process to obtain a target operation result, and then the target operation result is summed using the sigma function to obtain a second vector result.
[0109] Exemplarily, the similarity between the second sentence vector and each first target vector is calculated to obtain a first similarity matrix, and then the first similarity matrix is adjusted using the fourth weight parameter and the fourth deviation parameter to obtain a fifth operation result, and then the fifth operation result is summed using the sigma function to obtain a third vector result. The similarity between the first sentence vector and each second target vector is calculated to obtain a second similarity matrix, and then the second similarity matrix is adjusted using the fifth weight parameter and the fifth deviation parameter to obtain a sixth operation result, and then the sixth operation result is summed using the sigma function to obtain a fourth vector result.
[0110] Exemplarily, the first vector result, the second vector result, the third vector result and the fourth vector result are added to obtain a target summation result, which is then adjusted using the sixth weight parameter and the sixth deviation parameter and then summed using the sigma function to obtain a data fusion result, which is then converted using the Softmax function to obtain the corresponding text similarity between the first description text and the second description text.
[0111] Exemplarily, according to actual business needs and experience, a preset threshold is set, and then the calculated text similarity is compared with the preset threshold. Thus, when the text similarity value meets the preset threshold, it is determined that there is an association between the historical operation data and the user evaluation text, and the corresponding historical association relationship between the historical operation data and the user evaluation text is determined to be positively correlated; however, when the text similarity value does not meet the preset threshold, it means that the direct similarity between the historical operation data and the user evaluation text is low, and their association relationship cannot be simply determined. Further in-depth analysis is required at this time. According to the first description text, the first change trend corresponding to the historical merchant under the first abnormal data is determined, and this change trend can be divided into two situations of getting better or worse. Similarly, for the user evaluation text, the second change trend corresponding to the historical merchant under the second abnormal data is determined according to the second description text, and then the historical association relationship between the historical operation data and the user evaluation text is determined by comparing the first change trend and the second change trend. If the first change trend and the second change trend are the same, that is, both are either getting better or worse, then it can be determined that the historical association relationship corresponding to the historical operation data and the user evaluation text is positively correlated. On the contrary, when the first change trend and the second change trend are opposite, it is determined that the corresponding historical correlation relationship between the historical operation data and the user evaluation text is negatively correlated.
[0112] Specifically, a multi-dimensional comparison and analysis of the first description text and the second description text can unearth the potential correlation between historical operation data and user evaluation texts, thereby providing good support for subsequent early warning analysis.
[0113] Step S106: establishing a risk prediction model based on the historical operation data, the user evaluation text and the historical association relationship.
[0114] For example, through a professional annotation system, professional annotators will carefully analyze and annotate historical operation data and user evaluation texts according to established risk assessment standards and rules, and then assign corresponding risk labels, such as "high risk", "medium risk" or "low risk", and then obtain the risk annotation results corresponding to the historical operation data and user evaluation texts from the professional annotation system. Thus, appropriate machine learning or deep learning algorithms are used to integrate historical operation data and user evaluation texts using historical association relationships. During the training process, the model will learn the mapping relationship between historical operation data and user evaluation texts and risks, and gradually improve the ability to predict risks by continuously adjusting internal parameters. Finally, the model will output the risk prediction results. In order to evaluate the performance of the risk prediction model, the accuracy of the risk prediction results is calculated according to the risk labels. When the calculated accuracy does not meet the preset results, the risk prediction model is adjusted, such as grid search, random search, etc. By constantly trying different parameter combinations and observing the changes in accuracy, the optimal parameter settings are gradually found. After repeated adjustments and optimizations, until the accuracy meets the preset results, a reliable risk prediction model is obtained.
[0115] Step S107, obtaining current operation data and current evaluation text corresponding to the target merchant, and performing early warning monitoring on the target merchant according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result.
[0116] For example, the current operating data and current evaluation text corresponding to the target merchant are obtained from the database in real time, and then the current operating data and current evaluation text are input into the previously established risk prediction model. The risk prediction model will analyze and calculate the input data according to its internal algorithms and parameters, and then the risk prediction model will output the risk level of the target merchant, such as low risk, medium risk, and high risk. Thus, the risk level determines the target monitoring result corresponding to the target merchant.
[0117] In some embodiments, after obtaining the target monitoring result, the method also includes: when the target monitoring result is a target preset result, generating target warning information based on the target monitoring result; sending the target warning information to the target terminal corresponding to the target merchant, so that after the target terminal receives the target warning information, the target terminal reminds the target merchant according to the target warning information.
[0118] Exemplarily, when the target preset result is either medium risk or high risk, the corresponding target warning information is generated according to the target monitoring result and the generation rules. For example, [target merchant] Hello, according to the recent current operating data and current evaluation text, your customer has a problem with [target monitoring result]. Please solve the problem in time.
[0119] Exemplarily, the target terminal can be the mobile phone, tablet computer, office computer within the enterprise, etc. of the person in charge of the merchant. The collected information includes mobile phone number, email address, instant messaging tool account, etc., to ensure that the target warning information can be accurately sent to the target terminal, and then select the appropriate sending method according to the type and characteristics of the target terminal. For mobile terminals, you can choose SMS, instant messaging applications (such as WeChat, DingTalk, etc.); for office computers, you can send emails through the company's internal email system. At the same time, you can choose a single sending method or a combination of multiple methods according to the severity and urgency of the risk. For example, for high-risk target warning information, it can be sent by SMS and email at the same time to ensure that the merchant can receive it in time, so that the generated target warning information can be sent to the target terminal according to the selected sending method.
[0120] For example, a corresponding reminder mechanism is set on the target terminal to ensure that the target merchant can be reminded in time after receiving the target warning information. For mobile terminals, SMS reminder sound and vibration reminder can be set; for instant messaging applications, the message reminder function can be turned on; for email systems, an email reminder pop-up window can be set.
[0121] For example, after sending the target warning information, confirm whether the target merchant has received and noticed the warning information through appropriate means. The merchant can be prompted to reply and confirm in the target warning information, or in subsequent communication, the merchant can be asked whether he has received the warning and understand his handling of the warning information.
[0122] See also Figure 2 , Figure 2A merchant early warning monitoring system 200 based on multi-dimensional data fusion is provided in an embodiment of the present application. The merchant early warning monitoring system 200 based on multi-dimensional data fusion includes a data acquisition module 201, a first generation module 202, an abnormality analysis module 203, a second generation module 204, a relationship determination module 205, a model building module 206, and an early warning monitoring module 207, wherein the data acquisition module 201 is used to obtain historical operating data and user evaluation texts corresponding to historical merchants, and perform abnormality detection on the historical operating data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data; the first generation module 202 is used to determine the first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution; the abnormality analysis module 203 is used to cluster the user evaluation texts Analyze and obtain the second abnormal data and the second abnormal distribution corresponding to the second abnormal data; a second generation module 204 is used to determine the second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution; a relationship determination module 205 is used to determine the historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text according to the first description text and the second description text; a model establishment module 206 is used to establish a risk prediction model based on the historical operation data, the user evaluation text and the historical association relationship; an early warning monitoring module 207 is used to obtain the current operation data and the current evaluation text corresponding to the target merchant, and perform early warning monitoring on the target merchant according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result.
[0123] In some implementations, the merchant early warning monitoring system 200 based on multi-dimensional data fusion can be applied to a terminal device.
[0124] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the merchant early warning monitoring system 200 based on multi-dimensional data fusion described above can refer to the corresponding process in the aforementioned merchant early warning monitoring method based on multi-dimensional data fusion embodiment, and will not be repeated here.
[0125] See also Figure 3 , Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.
[0126] like Figure 3 As shown, the terminal device 300 includes a processor 301 and a memory 302 , and the processor 301 and the memory 302 are connected via a bus 303 , such as an I2C (Inter-integrated Circuit) bus.
[0127] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), and the processor 301 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0128] Specifically, the memory 302 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk.
[0129] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a partial structure related to the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the embodiment of the present invention is applied. The specific server may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0130] The processor is used to run a computer program stored in the memory, and implement any one of the merchant early warning monitoring methods based on multi-dimensional data fusion provided by the embodiments of the present invention when executing the computer program.
[0131] In one embodiment, the processor is used to run a computer program stored in the memory, and implements the following steps when executing the computer program:
[0132] Obtaining historical operation data and user evaluation text corresponding to historical merchants, and performing anomaly detection on the historical operation data to obtain first anomaly data and a first anomaly distribution corresponding to the first anomaly data;
[0133] Determine a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution;
[0134] Performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data;
[0135] Determine a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution;
[0136] Determine, according to the first description text and the second description text, a historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text;
[0137] Establishing a risk prediction model based on the historical operation data, the user evaluation text and the historical association relationship;
[0138] The current operation data and current evaluation text corresponding to the target merchant are obtained, and the target merchant is monitored for early warning according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result.
[0139] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the terminal device described above can refer to the corresponding process in the aforementioned merchant early warning monitoring method embodiment based on multi-dimensional data fusion, and will not be repeated here.
[0140] An embodiment of the present invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement any step of the merchant early warning monitoring method based on multi-dimensional data fusion provided in the description of the embodiment of the present invention.
[0141] The storage medium may be an internal storage unit of the terminal device described in the foregoing embodiment, such as a hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc., equipped on the terminal device.
[0142] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0143] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.
[0144] The serial numbers of the embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. The above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A merchant early warning monitoring method based on multi-dimensional data fusion, characterized in that: The method comprises: Obtaining historical operation data and user evaluation text corresponding to historical merchants, and performing anomaly detection on the historical operation data to obtain first anomaly data and a first anomaly distribution corresponding to the first anomaly data; Determine a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution; Performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data; Determine a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution; Determine, according to the first description text and the second description text, a historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text; Establishing a risk prediction model based on the historical operation data, the user evaluation text and the historical association relationship; The current operation data and current evaluation text corresponding to the target merchant are obtained, and the target merchant is monitored for early warning according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result.
2. The method according to claim 1, characterized in that The performing anomaly detection on the historical operation data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data includes: Determine a target window, and segment the historical operation data according to the target window to obtain a plurality of segmented operation data; Determine time information corresponding to each first sub-operation data in the split operation data, and determine a loss mapping value corresponding to the first sub-operation data according to the time information; Determine the window center corresponding to the split operation data according to the loss mapping value and the number of data corresponding to the first sub-operation data in the split operation data; Filtering out a target center associated with each second sub-operation data in the historical operation data from the window center; Determine the time information corresponding to each first sub-operation data in the split operation data according to the target center and the second sub-operation data, and determine the first distribution characterization value corresponding to the second sub-operation data according to the distance information between the loss mapping values corresponding to the first sub-operation data according to the time information, wherein the first distribution characterization value is used to represent the distribution status corresponding to the second sub-operation data; Filtering out the nearest center corresponding to the second sub-operation data from the center of the window, and determining a first maximum distribution value and a first minimum distribution value corresponding to the second sub-operation data according to the nearest center and the target center combined with a window width corresponding to the target window; Performing anomaly detection on the second sub-operation data according to the first maximum distribution value and the first minimum distribution value combined with the first distribution characterization value to obtain the first abnormal data; The first abnormal distribution corresponding to the first abnormal data is obtained by performing data statistics on the historical operation data according to the first abnormal data.
3. The method according to claim 2, characterized in that The determining the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operation data according to the nearest center and the target center combined with the window width corresponding to the target window includes: Determine a center distance between the nearest center and the target center, and when the center distance is greater than the window width, determine that a first parameter corresponding to the first minimum distribution value is a data difference between the center distance and the window width; When the center distance is less than or equal to the window width, determining that the first parameter corresponding to the first minimum distribution value is preset data; Determine a second parameter corresponding to the first maximum distribution value according to the center distance and the window width; Determine the segmented operation data corresponding to the target center as the associated operation data corresponding to the target center, and determine the quantity information corresponding to the target center in the associated operation data; Determine a target kernel function, and determine the first minimum distribution value corresponding to the second sub-operation data according to the target kernel function and the first parameter combined with the target center and the quantity information; Determine the first maximum distribution value corresponding to the second sub-operation data according to the target kernel function and the second parameter combined with the target center and the quantity information; The first minimum distribution value and the first maximum distribution value are obtained according to the following formula: ; ; in, represents the first minimum distribution value corresponding to the i-th second sub-operation data, represents the first maximum distribution value corresponding to the i-th second sub-operation data, represents the number of the target centers corresponding to the i-th second sub-operation data, represents the quantity information corresponding to the jth target center corresponding to the i-th second sub-operation data, represents the quantity information corresponding to the mth target center corresponding to the i-th second sub-operation data, represents the target kernel function, The first parameter representing the distance between the nearest center corresponding to the i-th second sub-operation data and the j-th target center and the window width; The second parameter corresponding to the center distance between the nearest center corresponding to the i-th second sub-operation data and the j-th target center and the window width.
4. The method according to claim 2, characterized in that: The performing anomaly detection on the second sub-operation data according to the first maximum distribution value and the first minimum distribution value combined with the first distribution characterization value to obtain the first abnormal data includes: Determine the nearest neighbor data corresponding to the second sub-operation data from the historical operation data, and determine a second distribution representation value corresponding to the nearest neighbor data; Determine a first loss value corresponding to the second sub-operation data and a second loss value corresponding to the nearest neighbor data; Determine a third distribution characterization value corresponding to the second sub-operation data according to the first distribution characterization value and the second distribution characterization value combined with the first loss value and the second loss value; Determine a second minimum distribution value and a second maximum distribution value corresponding to the nearest neighbor data, and determine a target threshold value according to the first minimum distribution value, the first maximum distribution value, the second minimum distribution value, and the second maximum distribution value; The first abnormal data is obtained by performing anomaly detection on the second sub-operation data according to the target threshold and the third distribution characterization value.
5. The method according to claim 1, characterized in that The step of performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data includes: Using multiple clustering algorithms to perform data clustering on the user evaluation text to obtain multiple initial clustering results; Determining a first evaluation text from the user evaluation texts, and obtaining remaining evaluation texts after excluding the first evaluation text from the user evaluation texts; Calculating a first similarity between each second evaluation text in the remaining evaluation text and the first evaluation text, and determining a corresponding second similarity between the first evaluation text and the remaining evaluation text according to the first similarity; Obtaining relevant clusters corresponding to the first evaluation texts from the initial clustering results respectively, and determining the number of texts corresponding to the relevant clusters; Determine a text weight corresponding to the first evaluation text according to the second similarity and the number of texts; Determine the cluster weight corresponding to each first subclass cluster in the initial clustering result according to the text weight; Calculating the classification similarity between any two of the initial clustering results, and determining the clustering weight corresponding to the initial clustering result according to the classification similarity and the cluster weight; Performing data clustering on the user evaluation text according to the text weight, the cluster weight and the cluster weight to obtain a target clustering result; Performing sentiment analysis on each second subclass cluster in the target clustering result to obtain a target sentiment type corresponding to the second subclass cluster; determining the second abnormal data from the second subclass cluster according to the target emotion type; According to the second abnormal data, data statistics are performed on the user evaluation text to obtain the second abnormal distribution corresponding to the second abnormal data.
6. The method according to claim 5, characterized in that The step of performing data clustering on the user evaluation text according to the text weight, the cluster weight and the cluster weight to obtain a target clustering result includes: Constructing a first diagonal matrix and a first column matrix according to the text weights, and determining a first text matrix according to the first diagonal matrix and the first column matrix; constructing a second diagonal matrix according to the cluster weights, and determining a second text matrix according to the first text matrix and the second diagonal matrix; constructing a third diagonal matrix according to the clustering weights, and determining a third text matrix according to the second text matrix and the third diagonal matrix; Determining a target supply-cooperation matrix corresponding to the user evaluation text according to the third text matrix and the transpose of the third text matrix; Data clustering is performed on the user evaluation text according to the target supply-cooperation matrix to obtain the target clustering result.
7. The method according to claim 1, characterized in that The determining, according to the first description text and the second description text, the historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text includes: Performing keyword recognition on the first description text to obtain a first keyword and performing keyword recognition on the second description text to obtain a second keyword; Performing vector representation on the first keyword using a text representation model to obtain a plurality of first representation vectors, and performing vector representation on the second keyword using the text representation model to obtain a plurality of second representation vectors; Obtaining a first target vector corresponding to the first keyword from the plurality of first representation vectors through a maximum pooling operation and a neural network, and performing concatenation processing according to the first target vectors to obtain a first sentence vector corresponding to the first description text; Obtaining a second target vector corresponding to the second keyword from the plurality of second representation vectors through the maximum pooling operation and the neural network, and obtaining a second sentence vector corresponding to the second description text through concatenation processing according to the second target vectors; Performing keyword similarity calculation on the first keyword and the second keyword by using the first target vector and the second target vector to obtain a first vector result; Calculating sentence similarity between the first description text and the second description text by using the first sentence vector and the second sentence vector to obtain a second vector result; Calculate the similarity between the first keyword and the second description text according to the first target vector and the second sentence vector to obtain a third vector result; Calculate the similarity between the second keyword and the first description text according to the second target vector and the first sentence vector to obtain a fourth vector result; Fusion of the first vector result, the second vector result, the third vector result and the fourth vector result to determine the text similarity between the first description text and the second description text; A preset threshold is determined, and the historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text is determined according to the text similarity and the preset threshold.
8. The method according to any one of claims 1 to 7, characterized in that After obtaining the target monitoring result, the method further includes: When the target monitoring result is a target preset result, target warning information is generated according to the target monitoring result; The target warning information is sent to a target terminal corresponding to the target merchant, so that after the target terminal receives the target warning information, the target terminal reminds the target merchant according to the target warning information.
9. A merchant early warning monitoring system based on multi-dimensional data fusion, characterized in that: include: A data collection module, used to obtain historical operation data and user evaluation text corresponding to historical merchants, and perform anomaly detection on the historical operation data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data; A first generating module, configured to determine a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution; An abnormality analysis module, used for performing cluster analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data; A second generating module, configured to determine a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution; A relationship determination module, configured to determine, according to the first description text and the second description text, a historical association relationship corresponding to the historical merchant under the historical operation data and the user evaluation text; A model building module, used to build a risk prediction model based on the historical operation data, the user evaluation text and the historical association relationship; The early warning monitoring module is used to obtain the current operating data and current evaluation text corresponding to the target merchant, and to perform early warning monitoring on the target merchant according to the risk prediction model in combination with the current operating data and the current evaluation text to obtain the target monitoring result.
10. A terminal device, characterized in that: The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and implement the merchant early warning monitoring method based on multi-dimensional data fusion as described in any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Information generation method and device, computer equipment and medium
CN114169796A
E-commerce data monitoring method and system based on multi-modal information fusion
CN118229330A
AI-based customer behavior analysis and prediction system and method
CN119250889A
Description evaluation method, description evaluation device, and description evaluation program
JP2019028937A