Merchant early warning monitoring method, system and terminal device based on multi-dimensional data fusion
By integrating and analyzing merchant operation data and user review texts from multiple dimensions, a risk prediction model was established, which solved the problem of inaccurate early warning analysis caused by inconsistent merchant data formats and achieved more accurate early warning monitoring.
Patent Information
- Application Number
- CN202510447186.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The lack of a unified format for multi-dimensional merchant data in existing technologies has significantly reduced the accuracy and reliability of early warning analysis results.
By performing anomaly detection on historical merchant operational data and cluster analysis on user review texts, a risk prediction model is established, which is then combined with current data for early warning and monitoring.
It enables the timely detection of potential risks to merchants, improving the accuracy and reliability of early warning analysis.
Smart Images

Figure CN119963236B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data fusion, and particularly relates to a merchant early warning monitoring method and system based on multi-dimensional data fusion and a terminal device. BACKGROUND
[0002] In today's digital business era, merchants have the ability to collect multi-dimensional data in the operation process, and the data is extensive and rich in sources. However, due to the lack of unified format of the collected multi-dimensional data, the data directly presents the characteristics of fragmentation and decentralization. When facing the key task of early warning analysis, the difference in data dimensions makes the existing early warning analysis method difficult to apply, thereby greatly reducing the accuracy and reliability of the early warning analysis result. SUMMARY
[0003] The main purpose of the embodiment of the present application is to provide a merchant early warning monitoring method and system based on multi-dimensional data fusion and a terminal device, which aims to solve the problem that the collected multi-dimensional data lacks a unified format and thus greatly reduces the accuracy and reliability of the early warning analysis result in the related art.
[0004] In a first aspect, the embodiment of the present application provides a merchant early warning monitoring method based on multi-dimensional data fusion, comprising:
[0005] Obtaining historical operation data and user evaluation text corresponding to a historical merchant, and performing abnormality detection on the historical operation data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data;
[0006] Determining a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution;
[0007] Performing clustering analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data;
[0008] Determining a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution;
[0009] Determining a historical correlation corresponding to the historical merchant under the historical operation data and the user evaluation text according to the first description text and the second description text;
[0010] Establishing a risk prediction model according to the historical operation data, the user evaluation text and the historical correlation;
[0011] obtain current operation data and current evaluation text corresponding to a target merchant, and perform early warning monitoring on the target merchant according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result.
[0012] In a second aspect, an embodiment of the present application provides a merchant early warning monitoring system based on multi-dimensional data fusion, comprising:
[0013] a data acquisition module configured to obtain historical operation data and user evaluation text corresponding to a historical merchant, and perform anomaly detection on the historical operation data to obtain first anomaly data and a first anomaly distribution corresponding to the first anomaly data;
[0014] a first generation module configured to determine first description text corresponding to the historical merchant according to the first anomaly data and the first anomaly distribution;
[0015] an anomaly analysis module configured to perform clustering analysis on the user evaluation text to obtain second anomaly data and a second anomaly distribution corresponding to the second anomaly data;
[0016] a second generation module configured to determine second description text corresponding to the historical merchant according to the second anomaly data and the second anomaly distribution;
[0017] a relationship determination module configured to determine a historical correlation corresponding to the historical merchant under the historical operation data and the user evaluation text according to the first description text and the second description text;
[0018] a model establishment module configured to establish a risk prediction model according to the historical operation data, the user evaluation text, and the historical correlation;
[0019] a warning monitoring module configured to obtain current operation data and current evaluation text corresponding to a target merchant, and perform early warning monitoring on the target merchant according to the risk prediction model in combination with the current operation data and the current evaluation text to obtain a target monitoring result.
[0020] In a third aspect, an embodiment of the present application further provides a terminal device, comprising a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory, wherein the computer program is executed by the processor to realize the steps of any one of the merchant early warning monitoring methods based on multi-dimensional data fusion provided in the specification of the present application.
[0021] The embodiment of the present application provides a merchant early warning monitoring method, system and terminal device based on multi-dimensional data fusion, which comprises the following steps: obtaining historical operation data and user evaluation text corresponding to a historical merchant, and performing abnormality detection on the historical operation data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data; determining a first description text corresponding to the historical merchant according to the first abnormal data and the first abnormal distribution; performing clustering analysis on the user evaluation text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data; determining a second description text corresponding to the historical merchant according to the second abnormal data and the second abnormal distribution; determining a historical correlation corresponding to the historical merchant under the historical operation data and the user evaluation text according to the first description text and the second description text; establishing a risk prediction model according to the historical operation data, the user evaluation text and the historical correlation; obtaining current operation data and current evaluation text corresponding to a target merchant, and performing early warning monitoring on the target merchant according to the risk prediction model, the current operation data and the current evaluation text to obtain a target monitoring result. The method determines a first description text corresponding to the historical operation data and a second description text corresponding to the user evaluation text of the historical merchant, so as to accurately determine a historical correlation corresponding to the historical operation data and the user evaluation text according to the first description text and the second description text, so as to reveal the internal relationship between the operation data and the user evaluation, and further provide more abundant information support for decision-making, so as to establish a risk prediction model according to the historical operation data, the user evaluation text and the historical correlation, accurately predict the risks that the merchant may face, and further perform early warning monitoring on the current operation data and the current evaluation text of the target merchant by using the risk prediction model, so as to timely find potential risk signals. The method also solves the problem that the multi-dimensional data collected in the related art lacks a unified format, and the accuracy and reliability of early warning analysis results are greatly reduced. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0023] Figure 1 A flowchart of a merchant early warning monitoring method based on multi-dimensional data fusion provided by the embodiment of the present application is shown in the figure.
[0024] Figure 2 A module structure diagram of a merchant early warning monitoring system based on multi-dimensional data fusion provided by the embodiment of the present application is shown in the figure.
[0025] Figure 3A structural schematic block diagram of a terminal device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of the present application.
[0027] The flowchart shown in the drawings is only an example and does not necessarily include all the contents and operations / steps, nor does it necessarily be executed in the described order. For example, some operations / steps can be further decomposed, combined or partially merged, so that the actual execution order can be changed according to the actual situation.
[0028] It should be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0029] The embodiments of the present application provide a merchant early warning monitoring method and system based on multi-dimensional data fusion and a terminal device. The merchant early warning monitoring method based on multi-dimensional data fusion can be applied to the terminal device, which can be an electronic device such as a tablet computer, a notebook computer, a desktop computer, a personal digital assistant and a wearable device. The terminal device can be a server or a server cluster.
[0030] Some embodiments of the present application will be described in detail below with reference to the drawings. In the case of no conflict, the following embodiments and features in the embodiments can be combined with each other.
[0031] Please refer to Figure 1 , Figure 1 A flowchart of a merchant early warning monitoring method based on multi-dimensional data fusion provided by an embodiment of the present application is shown.
[0032] As shown in Figure 1 , the merchant early warning monitoring method based on multi-dimensional data fusion includes steps S101 to S107.
[0033] Step S101, obtaining historical operation data corresponding to a historical merchant and user evaluation text, and performing abnormality detection on the historical operation data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data.
[0034] Exemplarily, historical operation data corresponding to the historical merchant and user evaluation texts of the consumers after historical user consumption are obtained from the database. The historical operation data are corresponding to the customer flow and inventory turnover rate of the historical merchant at different times. The historical operation data are discrete data types, and the user evaluation texts are text data types.
[0035] Exemplarily, statistical analysis is performed on the historical operation data based on a statistical method to obtain a deviation degree corresponding to each sub-operation data in the historical operation data, so as to determine an abnormal value corresponding to the sub-operation data according to the deviation degree, and further obtain first abnormal data from the historical operation data according to the abnormal value and a preset value.
[0036] Exemplarily, statistical analysis is performed on the first abnormal data, for example, the distribution of the first abnormal data at different time periods is calculated to obtain a first abnormal distribution corresponding to the first abnormal data.
[0037] In some embodiments, the abnormal detection on the historical operation data to obtain the first abnormal data and the first abnormal distribution corresponding to the first abnormal data comprises: determining a target window, and segmenting the historical operation data according to the target window to obtain a plurality of segmented operation data; determining time information corresponding to each first sub-operation data in the segmented operation data, and determining a loss mapping value corresponding to the first sub-operation data according to the time information; determining a window center corresponding to the segmented operation data according to the loss mapping value and a data quantity corresponding to the first sub-operation data in the segmented operation data; screening out a target center associated with each second sub-operation data in the historical operation data from the window center; determining a first distribution representation value corresponding to the second sub-operation data according to distance information between the target center and the second sub-operation data, the first distribution representation value being used to represent a distribution condition corresponding to the second sub-operation data; screening out a nearest center corresponding to the second sub-operation data from the window center, and determining a first maximum distribution value and a first minimum distribution value corresponding to the second sub-operation data according to the nearest center and the target center in combination with a window width corresponding to the target window; performing abnormal detection on the second sub-operation data according to the first maximum distribution value and the first minimum distribution value in combination with the first distribution representation value to obtain the first abnormal data; and performing data statistics on the historical operation data according to the first abnormal data to obtain the first abnormal distribution corresponding to the first abnormal data.
[0038] Exemplarily, a suitable time window is determined as the target window according to the business requirements. For example, the target window can be a week, half a month, or a month, etc., so that the historical operation data are segmented according to the time sequence according to the target window to obtain a plurality of segmented operation data. There is no data overlap between the segmented operation data.
[0039] For example, the time information corresponding to the first sub-operational data in each segmented operational data is obtained from the database. The base of the exponential function is then determined, with the base between 0 and 1. The time information corresponding to the first sub-operational data in the segmented operational data is then used as a reference time. The difference between the time information and the reference time is calculated, and this difference is substituted into the exponential function to obtain the loss mapping value corresponding to the first sub-operational data. This loss mapping value is used to characterize the degree of difference in the acquisition time of different first sub-operational data within the same target window.
[0040] For example, the number of data corresponding to each first sub-operational data in the segmented operation data is counted to obtain the data quantity of the first sub-operational data in the segmented operation data. However, based on the characteristics of historical operation data, it is known that it is difficult to find completely equal data in the segmented operation data. Therefore, when counting the data quantity of each first sub-operational data in the segmented operation data, this application considers other data in the segmented operation data to be the same data as the first sub-operational data when the data difference between other data in the segmented operation data and the first sub-operational data is within a preset difference range.
[0041] For example, the distribution weight corresponding to the first sub-operational data is determined according to the amount of data. Then, the distribution weight and the loss mapping value are multiplied together and then multiplied with the first sub-operational data to obtain the first data. Finally, the first data corresponding to all the first sub-operational data are summed to determine the window center corresponding to the segmented operation data.
[0042] For example, each sub-data in the historical operation data is determined as the second sub-operation data. Then, the center distance between the second sub-operation data and the center of each window is calculated according to calculation methods such as Euclidean distance and Manhattan distance. The center distance is compared with a preset distance, and the window center with a center distance less than or equal to the preset distance is determined as the target center corresponding to the second sub-operation data.
[0043] For example, the distance information between the second sub-operational data and each target center is calculated using methods such as Euclidean distance and Manhattan distance, and the segmented operation data corresponding to the target center is determined. Then, the target quantity corresponding to the target center in its corresponding segmented operation data is determined. When calculating the target quantity, it is considered that it is difficult to find data in the segmented operation data that is exactly the same as the target center. Therefore, when calculating the target quantity corresponding to each target center in its corresponding segmented operation data, this application considers other sub-data in the segmented operation data corresponding to the target center as the same data as the target center when the data difference between other sub-data and the target center is within a preset difference range. Thus, the target quantity corresponding to the target center in its corresponding segmented operation data is determined based on the data difference between other sub-data and the target center and the preset difference range.
[0044] For example, a target kernel function is determined, and then the target value corresponding to the distance information under the target kernel function is determined. Thus, the first distribution representation value corresponding to the second sub-operation data is obtained according to the following formula, wherein the first distribution representation value is used to represent the distribution status corresponding to the second sub-operation data.
[0045] ;
[0046] in, This represents the first distribution representation value corresponding to the t-th second sub-operational data. This represents the number of target centers corresponding to the t-th second sub-operational data. This indicates the number of targets corresponding to the k-th target center in its corresponding segmented operational data, which is the t-th second sub-operational data. This indicates the number of targets corresponding to the y-th target center in its corresponding segmented operational data, which is the t-th second sub-operational data. This represents the target kernel function, which can be a Gaussian kernel function or a modified Gaussian kernel function. This represents the distance information between the t-th second sub-operational data and the k-th target center.
[0047] For example, the window center corresponding to the minimum distance information between the center of the window and the second sub-operation data is determined as the nearest center corresponding to the second sub-operation data. Then, based on the nearest center and the target center, and combined with the window width corresponding to the target window, the range is adjusted according to the window width and the target center with the nearest center as the reference to obtain the boundary value corresponding to the first distribution representation value, that is, to obtain the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operation data.
[0048] For example, the first distribution representation value of each second sub-operational data is compared with its corresponding first maximum distribution value and first minimum distribution value. If the first distribution representation value exceeds the first maximum distribution value or is lower than the first minimum distribution value, the second sub-operational data is determined to be abnormal data, and these abnormal data are then aggregated to obtain the first abnormal data.
[0049] For example, statistical analysis is performed on the first abnormal data based on the time dimension under historical operational data to determine the first abnormal distribution corresponding to the first abnormal data. The first abnormal distribution is used to characterize the time point or time distribution information of the first abnormal data in historical operational data.
[0050] Specifically, the time information of the first sub-operational data is determined and the loss mapping value is calculated. This loss mapping value quantifies the deviation of the data from business objectives over time. By calculating the window center, filtering the target center and the nearest center, and combining distance information, the distribution representation value, maximum distribution value, and minimum distribution value are determined, enabling a more accurate depiction of the normal distribution range of the data. This anomaly detection method based on multi-dimensional information, compared to simple threshold judgments or single statistical indicators, can more accurately identify truly abnormal data, thus providing strong support for subsequent early warning and monitoring.
[0051] In some implementations, determining the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operational data based on the nearest center and the target center combined with the window width corresponding to the target window includes: determining the center distance between the nearest center and the target center, and when the center distance is greater than the window width, determining the first parameter corresponding to the first minimum distribution value as the data difference between the center distance and the window width; when the center distance is less than or equal to the window width, determining the first parameter corresponding to the first minimum distribution value as preset data; determining the second parameter corresponding to the first maximum distribution value based on the center distance and the window width; determining the segmented operation data corresponding to the target center as the associated operation data corresponding to the target center, and determining the quantity information corresponding to the target center in the associated operation data; determining the target kernel function, and determining the first minimum distribution value corresponding to the second sub-operational data based on the target kernel function and the first parameter combined with the target center and the quantity information; determining the first maximum distribution value corresponding to the second sub-operational data based on the target kernel function and the second parameter combined with the target center and the quantity information; wherein, the first minimum distribution value and the first maximum distribution value are obtained according to the following formulas:
[0052] ;
[0053] ;
[0054] in, This represents the first minimum distribution value corresponding to the i-th second sub-operational data. This represents the first maximum distribution value corresponding to the i-th second sub-operational data. This represents the number of target centers corresponding to the i-th second sub-operational data. This represents the quantity information corresponding to the j-th target center corresponding to the i-th second sub-operation data. This represents the quantity information corresponding to the m-th target center corresponding to the i-th second sub-operation data. This represents the target kernel function. The first parameter represents the relationship between the center distance between the nearest center and the target center corresponding to the i-th second sub-operation data and the window width; The second parameter represents the relationship between the center distance between the nearest center corresponding to the i-th second sub-operational data and the j-th target center and the window width.
[0055] For example, the center distance between the nearest center and the target center is calculated according to the Euclidean distance calculation formula. The center distance is then compared with the window width. When the center distance is greater than the window width, the center distance is subtracted from the window width to obtain the data difference, which is then used as the first parameter required for the subsequent calculation of the first minimum distribution value. When the center distance is less than or equal to the window width, the first parameter required for the calculation of the first minimum distribution value is determined as a preset data, where the preset data is equal to zero.
[0056] For example, the center distance and the window width are summed to determine the sum as a second parameter required for subsequent calculation of the first maximum distribution value.
[0057] For example, the segmented operational data corresponding to the target center is determined as the associated operational data corresponding to the target center. Then, the quantity information corresponding to the target center in its associated operational data is counted. When calculating the target quantity, it is considered that it is difficult to find data in the associated operational data that is exactly the same as the target center. Therefore, when counting the quantity information corresponding to each target center in its associated operational data, this application considers that other sub-data in the associated operational data corresponding to the target center as the same data as the target center when the data difference between the other sub-data and the target center is within a preset difference range. Thus, the quantity information corresponding to the target center in its associated operational data is determined based on the data difference between the other sub-data and the target center and the preset difference range.
[0058] For example, the target kernel function is determined as needed. The target kernel function can be a Gaussian kernel function or a modified function of the Gaussian kernel function. Then, the first minimum distribution value corresponding to the second sub-operational data is determined according to the following formula, combining the target kernel function, the first parameter, the target center, and the quantity information:
[0059] ;
[0060] in, This represents the first minimum distribution value corresponding to the i-th second sub-operational data. This represents the number of target centers corresponding to the i-th second sub-operational data. This represents the quantity information corresponding to the j-th target center for the i-th second sub-operational data. This represents the quantity information corresponding to the m-th target center for the i-th second sub-operational data. Represents the target kernel function. The first parameter represents the relationship between the center distance between the nearest center corresponding to the i-th second sub-operational data and the j-th target center and the window width.
[0061] For example, the first maximum distribution value corresponding to the second sub-operational data is determined according to the following formula, combining the target kernel function, the second parameter, quantity information, and the target center:
[0062] ;
[0063] in, This represents the first maximum distribution value corresponding to the i-th second sub-operational data. This represents the number of target centers corresponding to the i-th second sub-operational data. This represents the quantity information corresponding to the j-th target center for the i-th second sub-operational data. This represents the quantity information corresponding to the m-th target center for the i-th second sub-operational data. Represents the target kernel function. The second parameter represents the relationship between the center distance between the nearest center corresponding to the i-th second sub-operational data and the j-th target center and the window width.
[0064] Specifically, this application maps the original data to a high-dimensional feature space using a target kernel function (such as a Gaussian kernel function or an improved function thereof), thereby making the data features richer and more distinctive. It also uses the quantity information of the target centers corresponding to the second sub-operation data to adjust the mapping data corresponding to the target kernel function to obtain accurate first maximum distribution value and first minimum distribution value, thus providing reliable support for subsequent anomaly identification.
[0065] In some implementations, the step of obtaining the first abnormal data by performing anomaly detection on the second sub-operational data based on the first maximum distribution value and the first minimum distribution value combined with the first distribution characterization value includes: determining the nearest neighbor data corresponding to the second sub-operational data from the historical operation data, and determining the second distribution characterization value corresponding to the nearest neighbor data; determining the first loss value corresponding to the second sub-operational data and the second loss value corresponding to the nearest neighbor data; determining the third distribution characterization value corresponding to the second sub-operational data based on the first distribution characterization value and the second distribution characterization value combined with the first loss value and the second loss value; determining the second minimum distribution value and the second maximum distribution value corresponding to the nearest neighbor data, and determining the target threshold based on the first minimum distribution value, the first maximum distribution value, the second minimum distribution value, and the second maximum distribution value; and performing anomaly detection on the second sub-operational data based on the target threshold and the third distribution characterization value to obtain the first abnormal data.
[0066] For example, in historical operational data, the data closest to each second sub-operational data is found based on distance measurement methods (such as Euclidean distance, Manhattan distance, etc.), and these data are the nearest neighbor data corresponding to that second sub-operational data.
[0067] For example, for each nearest neighbor data, the second distribution representation value corresponding to each nearest neighbor data is calculated in the same way as the first distribution representation value was previously determined.
[0068] For example, the first loss value corresponding to the second sub-operation data and the second loss value corresponding to the nearest neighbor data are determined in the same way as determining the loss mapping value corresponding to each second sub-operation in the segmented operation data.
[0069] For example, the third distribution representation value corresponding to the second sub-operation data is determined by weighted summation of the first distribution representation value and the second distribution representation value based on the first loss value and the second loss value.
[0070] For example, for each nearest neighbor data, the second minimum distribution value and the second maximum distribution value corresponding to each nearest neighbor data are determined according to the method previously used to determine the first minimum distribution value and the first maximum distribution value.
[0071] For example, the first window position corresponding to each nearest neighbor data in the segmented operation data and the second window position corresponding to each second sub-operation data are obtained. Then, the integers from 1 to the first window position are summed to obtain a first value. Then, the first window position and the first value are divided to obtain a first weight corresponding to the first sub-operation data. The second value is obtained by summing the integers from 1 to the second window position. Then, the second window position and the second value are divided to obtain a second weight corresponding to the second sub-operation data.
[0072] For example, the first minimum distribution value and the first maximum distribution value are averaged to obtain a first average value, and the second minimum distribution value and the second maximum distribution value are averaged to obtain a second average value, thereby obtaining a target threshold by weighted summation of the first average value and the second average value according to the first weight and the second weight.
[0073] For example, the third distribution representation value of each second sub-operational data is compared with a target threshold. If the third distribution representation value exceeds the target threshold, the second sub-operational data is determined to be abnormal data. All the second sub-operational data determined to be abnormal are aggregated to obtain the first abnormal data.
[0074] Step S102: Determine the first descriptive text corresponding to the historical merchant based on the first abnormal data and the first abnormal distribution.
[0075] For example, adjacent data corresponding to the first abnormal data is obtained from historical operational data. The changing trend corresponding to the first abnormal data is then determined based on this adjacent data. This changing trend may include an upward trend, a downward trend, a fluctuating trend, or a stable trend. Based on this changing trend and the abnormal time corresponding to the first abnormal distribution, a statement generation rule is used to determine the first descriptive text corresponding to the historical merchant. For example, it can be stipulated that the abnormal time corresponding to the first abnormal distribution is described first, followed by the changing trend. For example, according to the statement generation rule, the determined descriptive elements are organized into text. For example: "During [the abnormal time corresponding to the first abnormal distribution], the historical merchant's [business indicators] showed a [changing trend]. This trend is highly correlated with the time of the abnormality, which has an [degree of impact] on the merchant's [business aspects]."
[0076] Step S103: Perform cluster analysis on the user review text to obtain the second abnormal data and the second abnormal distribution corresponding to the second abnormal data.
[0077] For example, clustering algorithms such as k-means clustering are used to perform cluster analysis on user review texts to obtain the text clustering results corresponding to the user review texts. Then, each sub-cluster in the text clustering results is classified into text rationality to obtain the text type corresponding to the sub-cluster. Then, the sub-cluster with the text type of negative user reviews is identified as the second abnormal data, thereby obtaining the time information corresponding to the second abnormal data and determining the second abnormal distribution corresponding to the second abnormal data based on the time information.
[0078] In some implementations, the step of performing cluster analysis on the user review text to obtain second abnormal data and a second abnormal distribution corresponding to the second abnormal data includes: using multiple clustering algorithms to cluster the user review text to obtain multiple initial clustering results; determining a first review text from the user review text, and obtaining remaining review texts after excluding the first review text from the user review text; calculating a first similarity between each second review text in the remaining review text and the first review text, and determining a corresponding second similarity between the first review text and the remaining review text based on the first similarity; obtaining relevant clusters corresponding to the first review text from the initial clustering results, and determining the number of texts corresponding to the relevant clusters; and determining the number of texts corresponding to the second similarity and the second abnormal distribution based on the second similarity and the second abnormal distribution. This quantity determines the text weight corresponding to the first evaluation text; based on the text weight, it determines the cluster weight corresponding to each first sub-cluster in the initial clustering result; it calculates the classification similarity between any two initial clustering results, and determines the cluster weight corresponding to the initial clustering result based on the classification similarity and the cluster weight; it performs data clustering on the user evaluation text based on the text weight, the cluster weight, and the cluster weight to obtain the target clustering result; it performs sentiment analysis on each second sub-cluster in the target clustering result to obtain the target sentiment type corresponding to the second sub-cluster; it determines the second abnormal data from the second sub-cluster based on the target sentiment type; and it performs data statistics on the user evaluation text based on the second abnormal data to obtain the second abnormal distribution corresponding to the second abnormal data.
[0079] For example, several suitable algorithms for text clustering, such as K-Means clustering, hierarchical clustering, and DBSCAN, are identified, and then the selected clustering algorithms are used to cluster user review texts to obtain multiple initial clustering results.
[0080] For example, a first review text is arbitrarily selected from the user review texts, and then excluded from the user review texts. The remaining text is the residual review text. A second review text is arbitrarily selected from the residual review texts, and a similarity calculation method such as cosine similarity or edit distance is used to calculate the first similarity between the second review text and the first review text.
[0081] For example, a first similarity is obtained between each second evaluation text in the first evaluation text and the remaining evaluation texts, and then all first similarities are summed to obtain the corresponding second similarity between the first evaluation text and the remaining evaluation texts.
[0082] For example, in each initial clustering result, clusters containing the first evaluation text are found. These clusters are the relevant clusters corresponding to the first evaluation text. The number of texts contained in each relevant cluster is then counted, and the total number of texts is summed. The second similarity score is then divided by the total number of texts to determine the text weight corresponding to the first evaluation text. This step takes into account that different clustering methods may produce different clustering results, and that there may be an imbalance in the number of clusters within each clustering result. Therefore, the bias information regarding cluster size is eliminated based on the number of texts contained in each relevant cluster.
[0083] For example, the text weights corresponding to each sub-text data in each first sub-cluster in the initial clustering result are obtained, and then the text weights corresponding to each sub-text data in the first sub-cluster are summed to obtain the cluster weights corresponding to the first sub-cluster.
[0084] For example, one initial clustering result is arbitrarily selected from multiple initial clustering results as the current clustering result. Then, the cluster weight corresponding to each second sub-cluster in the current clustering result is determined. Then, the cluster weights corresponding to all second sub-clusters are summed to obtain the weight sum value corresponding to the current clustering result. Then, algorithms such as the RAND index and the adjusted RAND index are used to calculate the classification similarity between the current clustering result and any one of the remaining initial clustering results, thereby obtaining the classification similarity between the current clustering result and each of the remaining initial clustering results. Then, all classification similarities are summed and multiplied by the weight sum value to obtain the cluster weight corresponding to the current clustering result. The above steps are performed on each initial clustering result to obtain the cluster weight corresponding to each initial clustering result.
[0085] For example, user review texts are re-clustered based on text weights, cluster weights, and grouping weights. For instance, the initial clustering results are adjusted and optimized based on text weights, cluster weights, and grouping weights to obtain the target clustering result.
[0086] For example, sentiment analysis methods such as dictionary-based sentiment analysis and machine learning sentiment analysis are used to perform sentiment analysis on each second sub-cluster in the target clustering results to determine the target sentiment type corresponding to each second sub-cluster, such as positive, negative, or neutral. Based on the target sentiment type, texts with a negative target sentiment type are then selected from the second sub-clusters as second outlier data.
[0087] For example, statistical analysis is performed on the second abnormal data in the user review text to analyze the distribution of the second abnormal data in different aspects (such as time, topic, review source, etc.), thereby obtaining the second abnormal distribution corresponding to the second abnormal data.
[0088] In some implementations, the step of performing data clustering on the user review text based on the text weight, the cluster weight, and the grouping weight to obtain a target clustering result includes: constructing a first diagonal matrix and a first column matrix based on the text weight, and determining a first text matrix based on the first diagonal matrix and the first column matrix; constructing a second diagonal matrix based on the cluster weight, and determining a second text matrix based on the first text matrix and the second diagonal matrix; constructing a third diagonal matrix based on the grouping weight, and determining a third text matrix based on the second text matrix and the third diagonal matrix; determining a target covariance matrix corresponding to the user review text based on the third text matrix and the transpose of the third text matrix; and performing data clustering on the user review text based on the target covariance matrix to obtain the target clustering result.
[0089] For example, firstly, the text weights are arranged in order. Each text weight is the weight value corresponding to each user review text calculated earlier. Using these text weight values as diagonal elements, a diagonal matrix is constructed, called the first diagonal matrix. The off-diagonal elements of the first diagonal matrix are all 0, and the elements on its diagonal correspond sequentially to the text weights of each user review text. The user review texts are then sorted sequentially to obtain a column matrix, thus obtaining the corresponding first column matrix.
[0090] For example, matrix multiplication is used to multiply the first diagonal matrix by the first column matrix. The matrix multiplication rule is: multiply the row elements of the first matrix by the column elements of the second matrix, then sum them to obtain the corresponding elements in the resulting matrix. The resulting matrix is the first text matrix.
[0091] For example, arrange these weight values in order according to the previously calculated cluster weights. Construct a second diagonal matrix with the cluster weight values as diagonal elements, and again, the off-diagonal elements are 0. This diagonal matrix reflects the importance weight of each cluster, thus obtaining the second diagonal matrix.
[0092] For example, the first text matrix is multiplied by the second diagonal matrix. Through matrix multiplication, each element of the first text matrix is further weighted according to its cluster weight to obtain the second text matrix. This step considers the impact of cluster importance on the text matrix.
[0093] For example, based on the previously calculated cluster weights, these weight values are arranged in order. A third diagonal matrix is constructed with the cluster weight values as diagonal elements and 0 as off-diagonal elements. This matrix reflects the importance of each cluster result, and then the second text matrix is multiplied by the third diagonal matrix. Through this matrix multiplication, the text weights, cluster weights, and group weights are comprehensively considered, and the second text matrix is finally weighted and adjusted to obtain the third text matrix.
[0094] For example, according to the rules of matrix transpose, the rows and columns of the third text matrix are interchanged to obtain its transpose matrix. Then, the third text matrix is multiplied by its transpose matrix to obtain the target supply and demand matrix corresponding to the user review texts. The target supply and demand matrix reflects the relationship between user review texts under the combined influence of correlation and weight.
[0095] For example, clustering algorithms such as spectral clustering can be used to perform clustering operations using the eigenvalues and eigenvectors of the target covariance matrix. The target covariance matrix can provide similarity information between data. Thus, the target covariance matrix is used as input data, and the spectral clustering algorithm is used to cluster user review texts. Then, the clustering algorithm will divide the user review texts into different clusters according to the relationship between the texts reflected by the matrix, and finally obtain the target clustering result.
[0096] Specifically, cluster weights reflect the importance of different clusters in the overall clustering structure. Constructing a second diagonal matrix and multiplying it by the first text matrix allows for further adjustment of the text matrix based on the importance of the clusters. Cluster weights consider the reliability and importance of the results obtained from different clustering algorithms. By constructing a third diagonal matrix and multiplying it by the second text matrix, the advantages of different clustering algorithms are combined. Different clustering algorithms may divide the data from different perspectives; comprehensively considering cluster weights can fully utilize the advantages of various algorithms, avoid the limitations of a single algorithm, and improve the accuracy and comprehensiveness of the clustering results. Thus, multiplying the third text matrix by its transpose yields the target covariance matrix, which can effectively capture the correlation between user review texts. The target covariance matrix reflects the similarity and interrelationships between texts; this relationship is the result of comprehensively considering text weights, cluster weights, and group weights, making it more comprehensive and accurate than simple text similarity calculations.
[0097] Step S104: Determine the second descriptive text corresponding to the historical merchant based on the second abnormal data and the second abnormal distribution.
[0098] For example, an anomaly type analysis is performed on the second abnormal data to obtain the target anomaly type corresponding to each second abnormal data. The target anomaly type includes, but is not limited to, product anomaly, service anomaly, and after-sales anomaly. The anomaly severity classification is performed on the second abnormal data to obtain the target anomaly severity corresponding to each second abnormal data. The target anomaly severity includes, but is not limited to, minor anomaly and severe anomaly.
[0099] For example, the high-frequency abnormal words corresponding to the second abnormal data are obtained by statistically analyzing the target abnormality type and the target abnormality degree, and the time information of the occurrence of the high-frequency abnormal words is determined according to the second abnormality distribution. Then, the corresponding second descriptive text is obtained according to the text generation rules based on the high-frequency abnormal words and the time of occurrence of the abnormality corresponding to the second abnormal data and the second abnormality distribution.
[0100] For example, the text generation rule is: "Under the [Target Anomaly Type], the proportion of [High-Frequency Anomaly Words] in recent customer feedback has increased significantly, generally occurring mainly under the [Anomaly Occurrence Time]."
[0101] Step S105: Determine the historical association relationship of the historical merchant under the historical operation data and the user review text based on the first description text and the second description text.
[0102] For example, historical operational data and user review texts are transformed to the same dimension to obtain corresponding first and second descriptive texts. Then, a text similarity algorithm is used to calculate the text similarity value between the first and second descriptive texts. When the text similarity value meets a preset value, a correlation is determined between the historical operational data and user review texts, and the historical correlation between them is determined to be positive. However, when the text similarity value does not meet the preset value, it indicates that the direct similarity between the historical operational data and user review texts is low, and their correlation cannot be simply determined. Further in-depth analysis is required. Based on the first descriptive text, the first trend of change for the historical merchant under the first abnormal data is determined. This trend can be divided into two cases: improvement or deterioration. Similarly, for user review texts, the second trend of change for the historical merchant under the second abnormal data is determined based on the second descriptive text. Then, the historical correlation between the historical operational data and user review texts is determined by comparing the first and second trends. If the first and second trends are the same, i.e., both improve or both deteriorate, then the historical correlation between the historical operational data and user review texts can be determined to be positive. Conversely, if the first and second trends are opposite, then the historical correlation between historical operational data and user review text is determined to be negative.
[0103] In some implementations, determining the historical association relationship of the historical merchant under the historical operating data and the user review text based on the first descriptive text and the second descriptive text includes: performing keyword recognition on the first descriptive text to obtain a first keyword and performing keyword recognition on the second descriptive text to obtain a second keyword; using a text representation model to perform vector representation on the first keyword to obtain multiple first representation vectors and using the text representation model to perform vector representation on the second keyword to obtain multiple second representation vectors; obtaining a first target vector corresponding to the first keyword from the multiple first representation vectors through max pooling operation and neural network, and performing concatenation processing based on the first target vector to obtain a first sentence vector corresponding to the first descriptive text; obtaining a second target vector corresponding to the second keyword from the multiple second representation vectors through the max pooling operation and the neural network, and performing concatenation processing based on the second target vector to obtain a second sentence vector corresponding to the second descriptive text. The second sentence vector; the keyword similarity calculation of the first keyword and the second keyword using the first target vector and the second target vector to obtain the first vector result; the sentence similarity calculation of the first descriptive text and the second descriptive text using the first sentence vector and the second sentence vector to obtain the second vector result; the similarity calculation of the first keyword and the second descriptive text using the first target vector and the second sentence vector to obtain the third vector result; the similarity calculation of the second keyword and the first descriptive text using the second target vector and the first sentence vector to obtain the fourth vector result; the first vector result, the second vector result, the third vector result, and the fourth vector result are fused to determine the corresponding text similarity between the first descriptive text and the second descriptive text; a preset threshold is determined, and the historical association relationship of the historical merchant under the historical operating data and the user review text is determined based on the text similarity and the preset threshold.
[0104] For example, the TF-IDF algorithm is used to identify the first keyword in the first descriptive text, and then the same method is used to analyze and identify the second keyword in the second descriptive text.
[0105] For example, using text representation models such as Word2Vec and GloVe, the first keyword is input into the model. The model converts each first keyword into a corresponding vector based on its pre-trained word vector library, thus obtaining multiple first representation vectors. The same text representation model is used to process the second keyword, converting each second keyword into a vector, resulting in multiple second representation vectors.
[0106] For example, a max-pooling operation is performed on multiple first representation vectors, that is, the maximum value is selected from the corresponding dimension of each first representation vector to form a new vector. This new vector is then input into a neural network for further processing to obtain the first target vector corresponding to the first keyword. The first target vectors are then concatenated according to certain rules, such as sequential concatenation, to finally obtain the first sentence vector corresponding to the first descriptive text. The above steps are repeated to perform max-pooling and neural network processing on multiple second representation vectors to obtain the second target vector corresponding to the second keyword, and these second target vectors are then concatenated to obtain the second sentence vector corresponding to the second descriptive text.
[0107] For example, the first target vector and the second target vector are multiplied to obtain a first result, and the first vector modulus corresponding to the first target vector and the second vector modulus corresponding to the second target vector are obtained. The first vector modulus and the second vector modulus are then multiplied to obtain a second result. The first result and the second result are then divided to obtain a third result. This yields a third result between multiple first keywords and second keywords. The third result is then concatenated into a related vector. The related vector is then adjusted using a first weight matrix and a first deviation parameter between the keywords to obtain the target result between the first keyword and the second keyword. Finally, the target result is summed using the sigma function to obtain the first vector result.
[0108] For example, the similarity between the first sentence vector and the second sentence vector is calculated using cosine similarity to obtain sentence similarity. Then, the first sentence vector and the second sentence vector are multiplied to obtain the first operation result. Next, the vector difference between the first sentence vector and the second sentence vector is calculated, and the absolute value of the difference is calculated to obtain the second operation result. Then, the first sentence vector and the second sentence vector are summed and adjusted using the second weight parameter and the second deviation parameter to obtain the third operation result. Finally, the first operation result, the second operation result, the third operation result and the fourth operation result are summed and adjusted using the third weight matrix and the third deviation matrix corresponding to the sentence vector processing to obtain the target operation result. Finally, the target operation result is summed using the sigma function to obtain the second vector result.
[0109] For example, the similarity between the second sentence vector and each first target vector is calculated to obtain a first similarity matrix. Then, the first similarity matrix is adjusted using a fourth weight parameter and a fourth deviation parameter to obtain a fifth operation result. Finally, the sigma function is used to sum the fifth operation result to obtain a third vector result. Similarly, the similarity between the first sentence vector and each second target vector is calculated to obtain a second similarity matrix. Then, the second similarity matrix is adjusted using a fifth weight parameter and a fifth deviation parameter to obtain a sixth operation result. Finally, the sigma function is used to sum the sixth operation result to obtain a fourth vector result.
[0110] For example, the first vector result, the second vector result, the third vector result, and the fourth vector result are added together to obtain the target summation result. The target summation result is then adjusted using the sixth weight parameter and the sixth bias parameter, and then summed using the sigma function to obtain the data fusion result. The Softmax function is then used to transform the data fusion result to obtain the corresponding text similarity between the first descriptive text and the second descriptive text.
[0111] For example, based on actual business needs and experience, a preset threshold is set, and the calculated text similarity is compared with the preset threshold. When the text similarity value meets the preset threshold, it is determined that there is a correlation between historical operational data and user review text, and the corresponding historical correlation between historical operational data and user review text is positively correlated. However, when the text similarity value does not meet the preset threshold, it indicates that the direct similarity between historical operational data and user review text is low, and their correlation cannot be simply determined. Further in-depth analysis is required. Based on the first descriptive text, the first trend of change corresponding to the historical merchant under the first abnormal data is determined. This trend can be divided into two cases: improvement or deterioration. Similarly, for user review text, based on the second descriptive text, the second trend of change corresponding to the historical merchant under the second abnormal data is determined. Then, by comparing the first and second trends, the historical correlation between historical operational data and user review text is determined. If the first and second trends are the same, that is, both improve or both deteriorate, then the corresponding historical correlation between historical operational data and user review text can be determined to be positively correlated. Conversely, if the first and second trends are opposite, then the historical correlation between historical operational data and user review text is determined to be negative.
[0112] Specifically, multi-dimensional comparison and analysis of the first and second descriptive texts can uncover potential correlations between historical operational data and user review texts, thereby providing strong support for subsequent early warning analysis.
[0113] Step S106: Establish a risk prediction model based on the historical operating data, the user review text, and the historical correlation.
[0114] For example, through a professional annotation system, professional annotators meticulously analyze and annotate historical operational data and user review texts according to established risk assessment standards and rules, assigning corresponding risk labels such as "high risk," "medium risk," or "low risk." The system then yields the risk annotation results corresponding to the historical operational data and user review texts. Appropriate machine learning or deep learning algorithms are then used to integrate the historical operational data and user review texts based on historical correlations. During training, the model learns the mapping relationship between historical operational data, user review texts, and risks, gradually improving its risk prediction ability by continuously adjusting internal parameters. Finally, the model outputs risk prediction results. To evaluate the performance of the risk prediction model, the accuracy of the risk prediction results is calculated based on the risk labels. When the calculated accuracy does not meet the preset results, the parameters of the risk prediction model are adjusted, for example, using grid search or random search. By continuously trying different parameter combinations and observing the changes in accuracy, the optimal parameter settings are gradually found. Through repeated adjustments and optimizations, until the accuracy meets the preset results, a reliable risk prediction model is obtained.
[0115] Step S107: Obtain the current operating data and current evaluation text corresponding to the target merchant, and conduct early warning monitoring on the target merchant based on the risk prediction model combined with the current operating data and the current evaluation text to obtain the target monitoring results.
[0116] For example, the current operational data and current evaluation text of the target merchant are obtained in real time from the database, and then input into a previously established risk prediction model. The risk prediction model analyzes and calculates the input data based on its internal algorithms and parameters, and then outputs the risk level of the target merchant, such as low risk, medium risk, or high risk. This risk level determines the target monitoring result corresponding to the target merchant.
[0117] In some embodiments, after obtaining the target monitoring result, the method further includes: when the target monitoring result is a target preset result, generating target early warning information based on the target monitoring result; and sending the target early warning information to the target terminal corresponding to the target merchant, so that after receiving the target early warning information, the target terminal reminds the target merchant based on the target early warning information.
[0118] For example, if the target preset result is either medium risk or high risk, then the corresponding target warning information is generated according to the generation rules based on the target monitoring result. For example, "[Target Merchant] Hello, according to the recent current operation data and current evaluation text, your account has a problem with the [Target Monitoring Result]. Please resolve this problem in a timely manner."
[0119] For example, the target terminal can be the merchant's mobile phone, tablet, or internal office computer. The collected information includes mobile phone number, email address, and instant messaging account, ensuring accurate delivery of the target alert information to the target terminal. The appropriate delivery method is then selected based on the type and characteristics of the target terminal. For mobile terminals, SMS or instant messaging applications (such as WeChat or DingTalk) can be used; for office computers, emails can be sent through the company's internal email system. Furthermore, a single delivery method or a combination of methods can be selected based on the severity and urgency of the risk. For example, for high-risk target alert information, both SMS and email can be sent simultaneously to ensure the merchant receives it promptly, thus delivering the generated target alert information to the target terminal according to the selected delivery method.
[0120] For example, a corresponding reminder mechanism can be set up on the target terminal to ensure that the target merchant is promptly notified after receiving the target warning information. For mobile terminals, SMS reminder sounds and vibration reminders can be set; for instant messaging applications, message reminder functions can be enabled; and for email systems, email reminder pop-ups can be set.
[0121] For example, after sending the target alert message, confirm whether the target merchant has received and noticed the alert message through appropriate means. This can be done by prompting the merchant to reply and confirm within the target alert message itself, or by inquiring about the merchant's receipt of the alert and understanding their handling of the alert message during subsequent communication.
[0122] Please see Figure 2 , Figure 2This application provides a merchant early warning and monitoring system 200 based on multi-dimensional data fusion. The system includes a data acquisition module 201, a first generation module 202, an anomaly analysis module 203, a second generation module 204, a relationship determination module 205, a model building module 206, and an early warning monitoring module 207. The data acquisition module 201 is used to obtain historical operational data and user review text corresponding to historical merchants, and to perform anomaly detection on the historical operational data to obtain first anomaly data and a first anomaly distribution corresponding to the first anomaly data. The first generation module 202 is used to determine a first descriptive text corresponding to the historical merchant based on the first anomaly data and the first anomaly distribution. The anomaly analysis module 203 is used to cluster the user review text. The system analyzes and obtains second abnormal data and a second abnormal distribution corresponding to the second abnormal data; a second generation module 204 is used to determine the second descriptive text corresponding to the historical merchant based on the second abnormal data and the second abnormal distribution; a relationship determination module 205 is used to determine the historical association relationship of the historical merchant under the historical operating data and the user review text based on the first descriptive text and the second descriptive text; a model building module 206 is used to build a risk prediction model based on the historical operating data, the user review text, and the historical association relationship; and an early warning monitoring module 207 is used to obtain the current operating data and current review text corresponding to the target merchant, and to perform early warning monitoring on the target merchant based on the risk prediction model and the current operating data and current review text to obtain the target monitoring result.
[0123] In some implementations, the merchant early warning and monitoring system 200 based on multi-dimensional data fusion can be applied to terminal devices.
[0124] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the merchant early warning monitoring system 200 based on multi-dimensional data fusion described above can be referred to the corresponding process in the aforementioned merchant early warning monitoring method embodiment based on multi-dimensional data fusion, and will not be repeated here.
[0125] Please see Figure 3 , Figure 3 This is a schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.
[0126] like Figure 3 As shown, the terminal device 300 includes a processor 301 and a memory 302, which are connected via a bus 303, such as an I2C (Inter-integrated Circuit) bus.
[0127] Specifically, processor 301 provides computing and control capabilities to support the operation of the entire terminal device. Processor 301 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0128] Specifically, the memory 302 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.
[0129] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the embodiments of the present invention, and does not constitute a limitation on the terminal device to which the embodiments of the present invention are applied. A specific server may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0130] The processor is used to run a computer program stored in a memory, and when executing the computer program, implements any of the merchant early warning and monitoring methods based on multi-dimensional data fusion provided in the embodiments of the present invention.
[0131] In one embodiment, the processor is configured to run a computer program stored in memory, and when executing the computer program, perform the following steps:
[0132] Obtain historical operational data and user review text corresponding to historical merchants, and perform anomaly detection on the historical operational data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data;
[0133] The first descriptive text corresponding to the historical merchant is determined based on the first abnormal data and the first abnormal distribution.
[0134] Cluster analysis is performed on the user review text to obtain the second abnormal data and the second abnormal distribution corresponding to the second abnormal data;
[0135] The second descriptive text corresponding to the historical merchant is determined based on the second abnormal data and the second abnormal distribution.
[0136] Based on the first description text and the second description text, determine the historical association relationship of the historical merchant under the historical operation data and the user review text;
[0137] A risk prediction model is established based on the historical operational data, the user review text, and the historical correlations.
[0138] Obtain the current operational data and current evaluation text corresponding to the target merchant, and conduct early warning monitoring of the target merchant based on the risk prediction model combined with the current operational data and the current evaluation text to obtain the target monitoring results.
[0139] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the terminal device described above can be referred to the corresponding process in the aforementioned embodiment of the merchant early warning monitoring method based on multi-dimensional data fusion, and will not be repeated here.
[0140] This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the steps of any of the merchant early warning monitoring methods based on multi-dimensional data fusion provided in the specification of this invention.
[0141] The storage medium can be an internal storage unit of the terminal device described in the foregoing embodiments, such as the hard drive or memory of the terminal device. Alternatively, the storage medium can be an external storage device of the terminal device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device.
[0142] Those skilled in the art will understand that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0143] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0144] The sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The above descriptions are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A merchant early warning monitoring method based on multi-dimensional data fusion, characterized in that, The method includes: Obtain historical operational data and user review text corresponding to historical merchants, and perform anomaly detection on the historical operational data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data; The first descriptive text corresponding to the historical merchant is determined based on the first abnormal data and the first abnormal distribution. Cluster analysis is performed on the user review text to obtain the second abnormal data and the second abnormal distribution corresponding to the second abnormal data; The second descriptive text corresponding to the historical merchant is determined based on the second abnormal data and the second abnormal distribution. Based on the first description text and the second description text, determine the historical association relationship of the historical merchant under the historical operation data and the user review text; A risk prediction model is established based on the historical operational data, the user review text, and the historical correlations. Obtain the current operational data and current evaluation text corresponding to the target merchant, and conduct early warning monitoring of the target merchant based on the risk prediction model combined with the current operational data and the current evaluation text to obtain the target monitoring results; Wherein, the step of performing cluster analysis on the user review text to obtain the second abnormal data and the second abnormal distribution corresponding to the second abnormal data includes: Multiple clustering algorithms were used to cluster the user review text to obtain multiple initial clustering results; The first evaluation text is determined from the user evaluation text, and the remaining evaluation text is obtained after excluding the first evaluation text from the user evaluation text. Calculate the first similarity between each second evaluation text in the remaining evaluation text and the first evaluation text, and determine the corresponding second similarity between the first evaluation text and the remaining evaluation text based on the first similarity; The relevant clusters corresponding to the first evaluation text are obtained from the initial clustering results, and the number of texts corresponding to the relevant clusters is determined. The text weight corresponding to the first evaluation text is determined based on the second similarity and the number of texts. The cluster weight corresponding to each first sub-cluster in the initial clustering result is determined based on the text weight; Calculate the classification similarity between any two initial clustering results, and determine the cluster weight corresponding to the initial clustering result based on the classification similarity and the cluster weight; The user review text is clustered based on the text weight, the cluster weight, and the clustering weight to obtain the target clustering result. Sentiment analysis is performed on each second sub-cluster in the target clustering results to obtain the target sentiment type corresponding to the second sub-cluster; The second anomalous data is determined from the second sub-cluster based on the target sentiment type; The second abnormal distribution is obtained by performing data statistics on the user review text based on the second abnormal data.
2. The method according to claim 1, characterized in that, The step of performing anomaly detection on the historical operational data to obtain first anomalous data and a first anomalous distribution corresponding to the first anomalous data includes: Determine the target window, and segment the historical operation data according to the target window to obtain multiple segmented operation data; Determine the time information corresponding to each first sub-operational data in the segmented operation data, and determine the loss mapping value corresponding to the first sub-operational data based on the time information; The window center corresponding to the segmented operation data is determined based on the loss mapping value and the number of data corresponding to the first sub-operation data in the segmented operation data. Filter the target center associated with each second sub-operational data in the historical operation data from the window center; Based on the target center and the second sub-operational data, the time information corresponding to each first sub-operational data in the segmented operation data is determined, and the distance information between the loss mapping values corresponding to the first sub-operational data is determined based on the time information to determine the first distribution characterization value corresponding to the second sub-operational data. The first distribution characterization value is used to represent the distribution status corresponding to the second sub-operational data. The nearest center corresponding to the second sub-operation data is selected from the center of the window, and the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operation data are determined based on the nearest center and the target center in combination with the window width corresponding to the target window; The first abnormal data is obtained by performing anomaly detection on the second sub-operation data based on the first maximum distribution value and the first minimum distribution value combined with the first distribution characterization value; The first abnormal distribution corresponding to the first abnormal data is obtained by performing data statistics on the historical operating data based on the first abnormal data.
3. The method according to claim 2, characterized in that, The step of determining the first maximum distribution value and the first minimum distribution value corresponding to the second sub-operational data based on the nearest center and the target center combined with the window width corresponding to the target window includes: Determine the center distance between the nearest center and the target center, and when the center distance is greater than the window width, determine the first parameter corresponding to the first minimum distribution value as the data difference between the center distance and the window width; When the center distance is less than or equal to the window width, the first parameter corresponding to the first minimum distribution value is determined to be preset data. The second parameter corresponding to the first maximum distribution value is determined based on the center distance and the window width; The segmented operational data corresponding to the target center is determined as the associated operational data corresponding to the target center, and the quantity information corresponding to the target center in the associated operational data is determined; Determine the target kernel function, and based on the target kernel function and the first parameter, combine the target center and the quantity information to determine the first minimum distribution value corresponding to the second sub-operational data; The first maximum distribution value corresponding to the second sub-operation data is determined based on the target kernel function and the second parameter, combined with the target center and the quantity information; The first minimum distribution value and the first maximum distribution value are obtained according to the following formulas: ; ; in, This represents the first minimum distribution value corresponding to the i-th second sub-operational data. This represents the first maximum distribution value corresponding to the i-th second sub-operational data. This represents the number of target centers corresponding to the i-th second sub-operational data. This represents the quantity information corresponding to the j-th target center corresponding to the i-th second sub-operation data. This represents the quantity information corresponding to the m-th target center corresponding to the i-th second sub-operation data. This represents the target kernel function. The first parameter represents the relationship between the center distance between the nearest center and the jth target center corresponding to the i-th second sub-operation data and the window width; The second parameter represents the relationship between the center distance between the nearest center corresponding to the i-th second sub-operational data and the j-th target center and the window width.
4. The method according to claim 2, characterized in that, The step of obtaining the first abnormal data by performing anomaly detection on the second sub-operational data based on the first maximum distribution value and the first minimum distribution value combined with the first distribution characterization value includes: Determine the nearest neighbor data corresponding to the second sub-operation data from the historical operation data, and determine the second distribution representation value corresponding to the nearest neighbor data; Determine the first loss value corresponding to the second sub-operational data and the second loss value corresponding to the nearest neighbor data; The third distribution representation value corresponding to the second sub-operation data is determined based on the first distribution representation value and the second distribution representation value, combined with the first loss value and the second loss value; Determine the second minimum distribution value and the second maximum distribution value corresponding to the nearest neighbor data, and determine the target threshold based on the first minimum distribution value, the first maximum distribution value, the second minimum distribution value, and the second maximum distribution value; The first abnormal data is obtained by performing anomaly detection on the second sub-operation data based on the target threshold and the third distribution characterization value.
5. The method according to claim 1, characterized in that, The step of performing data clustering on the user review text based on the text weight, the cluster weight, and the clustering weight to obtain the target clustering result includes: Construct a first diagonal matrix and a first column matrix based on the text weights, and determine a first text matrix based on the first diagonal matrix and the first column matrix; Construct a second diagonal matrix based on the cluster weights, and determine the second text matrix based on the first text matrix and the second diagonal matrix; Construct a third diagonal matrix based on the clustering weights, and determine the third text matrix based on the second text matrix and the third diagonal matrix; The target supply matrix corresponding to the user evaluation text is determined based on the third text matrix and the transpose of the third text matrix; The target clustering result is obtained by performing data clustering on the user review text based on the target supply matrix.
6. The method according to claim 1, characterized in that, The step of determining the historical association relationship of the historical merchant under the historical operation data and the user review text based on the first description text and the second description text includes: The first keyword is obtained by performing keyword recognition on the first descriptive text, and the second keyword is obtained by performing keyword recognition on the second descriptive text; The first keyword is represented by a text representation model to obtain multiple first representation vectors, and the second keyword is represented by the text representation model to obtain multiple second representation vectors. The first target vector corresponding to the first keyword is obtained from multiple first representation vectors through max pooling operation and neural network, and the first sentence vector corresponding to the first description text is obtained by concatenating the first target vector. The second target vector corresponding to the second keyword is obtained from multiple second representation vectors through the max pooling operation and the neural network, and the second sentence vector corresponding to the second descriptive text is obtained by concatenating the second target vector. The first vector result is obtained by calculating the keyword similarity between the first keyword and the second keyword using the first target vector and the second target vector. The second vector result is obtained by calculating the sentence similarity between the first descriptive text and the second descriptive text using the first sentence vector and the second sentence vector; A third vector result is obtained by calculating the similarity between the first keyword and the second descriptive text based on the first target vector and the second sentence vector. A fourth vector result is obtained by calculating the similarity between the second keyword and the first descriptive text based on the second target vector and the first sentence vector. The text similarity between the first descriptive text and the second descriptive text is determined by fusing the first vector result, the second vector result, the third vector result, and the fourth vector result. A preset threshold is determined, and the historical association relationship of the historical merchant under the historical operation data and the user review text is determined based on the text similarity and the preset threshold.
7. The method according to any one of claims 1-6, characterized in that, After obtaining the target monitoring results, the method further includes: When the target monitoring result is the target preset result, target early warning information is generated based on the target monitoring result; The target warning information is sent to the target terminal corresponding to the target merchant, so that after receiving the target warning information, the target terminal will remind the target merchant according to the target warning information.
8. A merchant early warning and monitoring system based on multi-dimensional data fusion, characterized in that, The merchant early warning and monitoring system based on multi-dimensional data fusion, as described in claim 1, comprises: The data acquisition module is used to obtain historical operational data and user review text corresponding to historical merchants, and to perform anomaly detection on the historical operational data to obtain first abnormal data and a first abnormal distribution corresponding to the first abnormal data. The first generation module is used to determine the first descriptive text corresponding to the historical merchant based on the first abnormal data and the first abnormal distribution. Anomaly analysis module is used to perform cluster analysis on the user review text to obtain second abnormal data and the second abnormal distribution corresponding to the second abnormal data; The second generation module is used to determine the second descriptive text corresponding to the historical merchant based on the second abnormal data and the second abnormal distribution. The relationship determination module is used to determine the historical association relationship of the historical merchant under the historical operation data and the user evaluation text based on the first description text and the second description text. The model building module is used to build a risk prediction model based on the historical operational data, the user review text, and the historical correlations. The early warning monitoring module is used to obtain the current operating data and current evaluation text of the target merchant, and to conduct early warning monitoring of the target merchant based on the risk prediction model combined with the current operating data and the current evaluation text to obtain the target monitoring results.
9. A terminal device, characterized in that, The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and, when executing the computer program, implement the merchant early warning monitoring method based on multi-dimensional data fusion as described in any one of claims 1 to 7.
Citation Information
Patent Citations
E-commerce data monitoring method and system based on multi-modal information fusion
CN118229330A
AI-based customer behavior analysis and prediction system and method
CN119250889A