Data analysis methods, apparatus, electronic devices and readable media
By acquiring the characteristic change sequences and object change sequences of the business, and using causal models for estimation and statistics, the problem of high human cost in machine learning models is solved, and efficient data analysis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-05-24
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, data analysis for machine learning models requires extensive data annotation by experts, resulting in high labor costs and impacting overall efficiency.
By acquiring the characteristic change sequence and object change sequence of the business, and using causal models for estimation and statistics, object change estimation data can be directly generated from business data. The correlation between the characteristic change sequence and the object change sequence can be analyzed, reducing the data labeling process.
It reduces the labor costs of data analysis, improves overall efficiency, and allows for direct analysis using raw data, eliminating the need for data identification and labeling steps.
Smart Images

Figure CN114896565B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly to a method, apparatus, electronic device, and readable medium for data analysis. Background Technology
[0002] With the development of computer technology, business and services conducted through computers and the internet are increasing daily. This generates a large amount of business data. Analyzing and organizing this data to identify key factors impacting business operations is beneficial for timely adjustments to business content and management methods.
[0003] In related technologies, the method for analyzing business data is to use machine learning models. The machine model is trained using data labeled by experts, and then the trained machine model is used for data analysis.
[0004] However, the above methods require experts to analyze and label a large amount of data for training in order to obtain an accurate machine model. Therefore, a large amount of manpower is required, which increases the labor cost of the solution and affects the overall efficiency of the solution. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a data analysis method, apparatus, electronic device, and readable medium to reduce the labor costs of data analysis solutions and improve the overall efficiency of data analysis solutions.
[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0007] According to one aspect of the embodiments of this application, a data analysis method is provided, including:
[0008] Acquire a sequence of characteristic changes in the business, the sequence of characteristic changes including business data collected within a preset time period;
[0009] Obtain the object change sequence of the business, where the object change sequence represents the result obtained based on the statistical data of the business.
[0010] Based on the feature change sequence, an object change estimation data is generated, which represents the result predicted based on the business data.
[0011] Based on the estimated object change data and the object change sequence, model statistics are performed to obtain statistical data, which represents the significant relationship between the feature change sequence and the object change sequence;
[0012] Based on the statistical data, the correlation between the changes in the feature change sequence and the changes in the object change sequence is determined, and the data analysis results are obtained.
[0013] According to one aspect of the embodiments of this application, a data analysis apparatus is provided, comprising:
[0014] The feature change acquisition module is used to acquire the feature change sequence of the business, the feature change sequence including business data collected within a preset time period;
[0015] An object change acquisition module is used to acquire the object change sequence of the business, wherein the object change sequence represents the result obtained based on the statistical data of the business.
[0016] The data estimation module is used to estimate based on the feature change sequence and generate object change estimation data, wherein the object change estimation data represents the result predicted based on the business data;
[0017] The model statistics module is used to perform model statistics based on the object change estimation data and the object change sequence to obtain statistical data, which represents the significant relationship between the feature change sequence and the object change sequence.
[0018] The results analysis module is used to determine the correlation between the changes in the feature change sequence and the changes in the object change sequence based on the statistical data, and to obtain data analysis results.
[0019] In some embodiments of this application, based on the above technical solutions, the feature change sequence includes P business change sequences, and the object change sequence includes P object change data; the data estimation module includes:
[0020] The factor generation submodule is used to generate an influence factor based on the P business change sequences and the P object change data. The influence factor represents the degree to which the P-th object change data is affected by the P business change sequences or by the P-1 object change data.
[0021] The estimated data generation submodule is used to weight the P business change sequences according to the influencing factors to generate object change estimated data.
[0022] In some embodiments of this application, based on the above technical solutions, the influence factor includes a first influence parameter, each business change sequence includes M business features, where M is an integer greater than or equal to 1, the first influence parameter includes M business parameters, and the business parameters represent the degree of influence of the corresponding business features on the object change data; the factor generation submodule includes:
[0023] The business parameter determination unit is used to determine the M business parameters based on the weighted sum of the M business parameters and the corresponding M business features, as well as the object change data corresponding to the business change sequence.
[0024] The parameter generation unit is used to merge the business parameters corresponding to each business change sequence into the first influence parameter. The first influence parameter is a P×M matrix, and each element represents the degree of influence of the business features in the corresponding business change sequence on the object change data.
[0025] In some embodiments of this application, based on the above technical solutions, the estimated data generation submodule includes:
[0026] The business weighting unit is used to weight the corresponding business features in the corresponding business change sequence according to each element in the first influence parameter, so as to obtain P×M weighted business features.
[0027] The feature summation unit is used to sum the P×M weighted service features to obtain P estimated data.
[0028] In some embodiments of this application, based on the above technical solutions, the model statistics module includes:
[0029] The first discrete relation value determination submodule is used to determine discrete relation values based on the mapping relationship between the estimated object change data and the P object change data. The discrete relation values identify the degree of dispersion between the estimated object change data and the object change data.
[0030] The first statistical summation submodule is used to perform statistical summation on the discrete relation values to obtain the statistical data.
[0031] In some embodiments of this application, based on the above technical solutions, the result analysis module includes:
[0032] The first statistical distribution submodule is used to input the quantity of the object's change data into a statistical distribution function to obtain the statistical threshold.
[0033] The first association determination submodule is used to determine, if the statistical data is greater than the statistical threshold, that the changes in the P business change sequences are associated with the changes in the Pth object change data.
[0034] The first association determination submodule is further configured to determine that if the statistical data is less than or equal to the statistical threshold, then the P business change sequences are not associated with the change data of the Pth object.
[0035] In some embodiments of this application, based on the above technical solutions, the data analysis device further includes:
[0036] The first standard deviation determination module is used to determine the object standard deviation of P object change data;
[0037] The first statistical value determination module is used to determine, for the determined business change sequence, the ratio of the element in the first influence parameter corresponding to each business feature in the business change sequence to the standard deviation of the object, so as to obtain the feature statistical value of each business feature;
[0038] The first correlation determination module is used to determine the correlation between changes in each business feature and changes in object data based on the comparison results of the feature statistics value and the threshold of the feature statistics value for each business feature.
[0039] In some embodiments of this application, based on the above technical solutions, the influence factor further includes a second influence parameter, which includes P-1 object parameters, wherein the object parameters represent the degree of influence of the corresponding object change data on the Pth object change data; the business parameter determination unit includes:
[0040] The object parameter determination subunit is used to determine the M business parameters and the P-1 object parameters based on the weighted sum of the M business parameters and the corresponding M business features, the weighted sum of the P-1 object parameters and the corresponding P-1 object change data, and the object change data corresponding to the business change sequence.
[0041] The data analysis device also includes:
[0042] The second influence parameter determination subunit is used to determine the second influence parameter based on the calculated P-1 object parameters and the preset object parameters corresponding to the Pth object change data.
[0043] In some embodiments of this application, based on the above technical solutions, the data analysis device further includes:
[0044] The weighted change data determination module is used to weight the corresponding object change data according to each object parameter in the second influence parameter to obtain P weighted change data;
[0045] The change prediction data determination module is used to sum the P weighted change data to obtain change prediction data, which represents the result predicted based on the P-1 object change data.
[0046] In some embodiments of this application, based on the above technical solutions, the model statistics module includes:
[0047] The second discrete relation value determination submodule is used to determine discrete relation values based on the mapping relationship between the estimated object change data and the P object change data. The discrete relation values identify the degree of dispersion between the estimated object change data and the object change data.
[0048] The second statistical summation submodule is used to perform statistical summation on the discrete relation values to obtain the statistical data.
[0049] In some embodiments of this application, based on the above technical solutions, the result analysis module includes:
[0050] The second statistical distribution submodule inputs the quantity of the object's changing data into a statistical distribution function for calculation to obtain the statistical threshold.
[0051] The second association determination submodule is used to determine that if the statistical data is greater than the statistical threshold, there is an association between the changes in the P-1 object change data and the changes in the Pth object change data.
[0052] The second association determination submodule is further configured to determine that if the statistical data is less than or equal to the statistical threshold, the changes in the P-1 object change data are not associated with the changes in the Pth object change data.
[0053] In some embodiments of this application, based on the above technical solutions, the data analysis device further includes:
[0054] The second standard deviation determination module is used to determine the object standard deviation of the P-1 object change data;
[0055] The second statistical value determination module is used to calculate the ratio of the object parameter corresponding to each object change data to the object standard deviation for the P-1 object change data, and obtain the object statistical value of the P-1 object change data.
[0056] The second correlation determination module is used to determine the impact of the changes in the P-1 objects on the changes in the Pth object based on the comparison results of the characteristic statistical values of the P-1 object change data and the object statistical value threshold.
[0057] According to one aspect of the embodiments of this application, an electronic device is provided, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the data analysis method as described above by executing the executable instructions.
[0058] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the data analysis method as described in the above technical solutions.
[0059] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the data analysis method provided in the various optional implementations described above.
[0060] In the embodiments of this application, firstly, a feature change sequence and an object change sequence of the business are obtained. The feature change sequence is generated based on the data to be analyzed, while the object change sequence represents the result obtained from statistical analysis of the business data. Subsequently, estimation is performed based on the feature change sequence to generate estimated object change data, which represents the result predicted based on the business data. Then, model statistics are performed based on the estimated object change data and the object change sequence to obtain statistical data, which represents the significant relationship between the feature change sequence and the object change sequence. Finally, based on the statistical data, the correlation between the changes in the feature change sequence and the changes in the object change sequence is determined, resulting in the data analysis result. In the data analysis process, the estimated object change data is directly obtained from the feature change sequence obtained from the business data, and model statistics are performed. The correlation between the changes in the object change data and the changes in the feature change sequence is analyzed using the statistical data. This allows for direct use of the raw data for the analysis process, eliminating the need for data identification and labeling. This reduces the labor costs of the data analysis solution and improves its overall efficiency.
[0061] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0062] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0063] Figure 1 This is a schematic diagram of an implementation environment in one of the embodiments of this application;
[0064] Figure 2This is a schematic flowchart illustrating the overall solution in the embodiments of this application;
[0065] Figure 3 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0066] Figure 4 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0067] Figure 5 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0068] Figure 6 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0069] Figure 7 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0070] Figure 8 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0071] Figure 9 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0072] Figure 10 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0073] Figure 11 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0074] Figure 12 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0075] Figure 13 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0076] Figure 14 This is a schematic flowchart of a data analysis method according to an embodiment of this application;
[0077] Figure 15 A schematic block diagram of the data analysis apparatus in an embodiment of this application is shown.
[0078] Figure 16 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation
[0079] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0080] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0081] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0082] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily need to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0083] It should be understood that the data analysis method in this application can be applied to scenarios involving causal analysis based on business data, and specifically to scenarios and products related to discounted refueling and travel service operations in the Internet of Vehicles (IoV). Taking discounted refueling as an example, gas stations accumulate a large amount of vehicle refueling data, as well as related sales data and data communication data, through the IoV network during daily operations. During different sales cycles, changes in various sales conditions, such as promotional activities, fluctuations in oil prices, the impact of holidays, and large-scale group events, can lead to changes in business data, such as changes in sales volume or website traffic. Through the solution in this application, based on the business-related feature data collected from the IoV network, it is possible to analyze which types of business data influence the traffic or number of visitors to services such as discounted refueling, thereby enabling the analysis of the reasons for changes in visitor numbers and making corresponding adjustments. The embodiments of this invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0084] The concept of the Internet of Vehicles (IoV) originates from the Internet of Things (IoT), which uses vehicles in motion as information sensing objects and leverages next-generation information and communication technologies to achieve network connections between vehicles, people, roads, service platforms, etc., thereby improving the overall intelligent driving level of vehicles, providing users with a safe, comfortable, intelligent, and efficient driving experience and transportation services, while improving traffic operation efficiency and enhancing the level of intelligence in social transportation services.
[0085] The technical solution provided in this application will be described in detail below with reference to specific embodiments. Please refer to... Figure 1 , Figure 1 This is a schematic diagram of an implementation environment in this application embodiment. The implementation environment includes an in-vehicle terminal 110, a server 120, and a management terminal 130. The in-vehicle terminal 110 and the server 120 communicate via a wired or wireless network. The server 120 is equipped with a data analysis device that receives data sent from the in-vehicle terminal 110, analyzes the data, and generates data analysis results. Users can browse the data analysis results on the server through the management terminal 130. Specifically, during daily service operations, the in-vehicle terminal 110 sends the data it needs to collect to the server 120. After collecting sufficient data, the server 120 performs data analysis to obtain data analysis results for business personnel or managers to understand changes in the business and the reasons for these changes. For example, if a manager notices a decrease in the number of visitors during a preset time period while refueling, they can use the data analysis results to find the reason for the decrease. Based on the data analysis results, it may be found that the number of clicks on promotional activities in the business data affects sales, thus revealing that the sales decrease is due to changes in promotional activities.
[0086] Server 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This document does not impose any restrictions on this.
[0087] The vehicle-mounted terminal 110 and the management terminal 130 can be mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminal equipment, aircraft, etc., but are not limited to these. The vehicle-mounted terminal 110 and the server 120 can be directly or indirectly connected through wired or wireless communication, which is not limited herein. The number of client 110 and server 120 is also not limited.
[0088] The following section uses the Internet of Vehicles (IoV) as an example to introduce the overall process of the solution proposed in this application. Please refer to [link / reference needed]. Figure 2 , Figure 2 This is a schematic flowchart illustrating the overall solution in an embodiment of this application. Figure 2As shown, the overall scheme comprises six stages: data acquisition stage 210, causal model construction stage 220, causal model regression stage 230, statistical attribution model construction stage 240, statistical attribution model testing stage 250, and source tracing stage 260. The specific objective of data analysis is to analyze which input features influence the target variable, given that the target variable has been identified. Therefore, in data acquisition stage 210, the target variable data and the input feature data whose influence needs to be analyzed are acquired. Specifically, the data analysis device collects business data from the vehicle network for several periods of refueling operations and extracts the change in the number of transaction objects participating in the discounted refueling activity as the target variable data. Specifically, the business can be segmented by time, such as weekly, ten-day, or monthly periods. Input feature data is typically business-related data and is also acquired in stages according to the same rules as the target variable. Specifically, each period's data will include, for example, data on changes in refueling transaction volume during holidays, behavioral changes in accessing / favoriting / commenting on discounted refueling activities, data on changes in the data collection process, and price changes. In the causal model construction phase 220, a causal model based on time series and feature data is constructed using the data collected in the data acquisition phase 210. The collected data is arranged chronologically according to its period, forming a time series. The causal model can be a specific algorithm or model; the specific calculation method is determined based on the type, quantity, and number of features of the collected target variable data and input feature data, thus constructing the causal model. In the causal model regression phase 230, the collected information is input into the constructed model for regression calculation. The model is solved progressively according to the period, yielding the model parameters. Subsequently, in the statistical attribution model construction phase 240, the target data is estimated using the model constructed in the causal model construction phase 220 and the model parameters calculated in the causal model regression phase 230, based on the collected input feature data, thereby obtaining the estimated change data of the object. In the statistical attribution model testing phase 250, comparing and analyzing the correlation between the estimated data of object change and the collected target variable data allows analysis of whether the input feature data has influenced the target variable data. For example, if the input feature data includes too many useless features, the input feature data may not have sufficient influence on the target variable data. In the case of time series analysis, it is also possible to analyze whether historical target variable data has influenced subsequent target variable data, i.e., whether the target variable data in period P-1 has an impact on the target variable data in period P.For example, if trading volume is used as the target variable, a higher trading volume in period p-1 might lead to a decrease in market demand, resulting in a decrease in trading volume in period p, thus having an impact. In the source tracing phase 260, the input feature data or historical target variable data that influenced the target variable data identified in the statistical attribution model testing phase 250 can be further analyzed to determine which specific feature or which specific period of variable data caused the influence.
[0089] The data analysis methods described in the embodiments of this application are further explained below. Please refer to... Figure 3 , Figure 3 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 3 As shown, this data analysis method includes at least steps S310 to S350, which are described in detail below:
[0090] Step S310: Obtain the characteristic change sequence of the service, which includes the service data collected within a preset time period;
[0091] Step S320: Obtain the object change sequence of the business. The object change sequence represents the result obtained based on business data statistics.
[0092] Feature change sequences refer to business data typically collected within a preset time period. Changes in this data cause changes in other related data within the business data. Object change sequences refer to results obtained based on business data statistics. These changes are generally considered to be influenced by changes in other data. Therefore, object change sequences can be statistically obtained based on changes in business data. For example, when analyzing changes in the number of customers, the object change sequence can be the change in customer numbers, generated based on the change in the number of customers caused by changes in other business data. For example, if the number of customers decreased by 30, the object change sequence would include the element 30. The data to be analyzed is usually business data from the service process. During data acquisition, business data from the vehicle network storage can be divided into P periods according to time sequence and extracted. From each period of data, the corresponding feature change sequences and object change sequences are extracted according to data extraction rules. Specifically, feature change sequences usually include one or more business features, such as changes in the number of fuel transactions during holidays, changes in click / favorite / comment discount fuel behavior, link changes during data collection, and price changes. The data analysis device calculates and extracts various business features from the business data according to the rules governing each feature, thereby forming a feature change sequence. The extraction method for the object change sequence is the same as that for the feature change sequence.
[0093] Step S330: Estimate based on feature change sequence to generate object change estimation data, which represents the result predicted based on business data.
[0094] The data analysis process focuses on the object change sequence. Assuming the object change sequence includes P object change data points, and the sequence is sorted by time series, the other P-1 object change data points can be considered historical data of the Pth object change data point. The data analysis device is equipped with a causal model. This model includes P characteristic change sequences and the correlation between the P-1 object change data points and the Pth object change data point. The degree to which the P characteristic change sequences and P-1 object change data points influence the Pth object change data point is represented by parameters in the causal model. Based on the characteristic change sequences and object change sequences, relevant parameters in the causal model can be calculated to determine the causal model corresponding to the collected business data. Using the determined causal model, the data in the object change sequence can be estimated based on the characteristic change sequences, thus obtaining estimated object change data. If multiple object change data points exist in the object change sequence, multiple estimated object change data points will be estimated based on the causal model.
[0095] Step S340: Perform model statistics based on the object change estimation data and the object change sequence to obtain statistical data, which represents the significant relationship between the feature change sequence and the object change sequence.
[0096] Each feature change sequence corresponds to an estimated object change. The data analysis device statistically analyzes the estimated object change data and the corresponding object change sequence according to a pre-defined statistical model, thus obtaining statistical data on both. The statistical model typically depends on the data types in the collected feature change sequences and object change sequences, for example, based on whether the input meets the conditions for parametric and non-parametric statistics and the sample size. The statistical data directly reflects the significant relationship between the estimated object change data and the object change sequence. Since each estimated object change data corresponds to a feature change sequence, the statistical data also represents the significant relationship between the feature change sequence and the object change sequence.
[0097] Step S350: Based on statistical data, determine the correlation between the changes in the feature change sequence and the changes in the object change sequence to obtain the data analysis results.
[0098] The data analysis device uses statistical data to determine the correlation between changes in the characteristic change sequence and changes in the object change sequence, thereby obtaining data analysis results. Specifically, the data analysis device can compare the distribution of statistical data with the results of a preset distribution function to assess whether there is a correlation between the estimated object change data and the actual object change data. Correlation typically includes positive and negative correlations, as well as specific forms such as linear correlation and non-linear correlation. If there is a correlation between the estimated object change data and the actual object change data, it is determined that the corresponding characteristic change sequence has an impact on the object change data; conversely, if no correlation exists, it is determined that the characteristic change sequence has no impact on the object change data.
[0099] In the embodiments of this application, firstly, a feature change sequence and an object change sequence of the business are obtained. The feature change sequence is generated based on the data to be analyzed, while the object change sequence represents the result obtained from statistical analysis of the business data. Subsequently, estimation is performed based on the feature change sequence to generate estimated object change data, which represents the result predicted based on the business data. Then, model statistics are performed based on the estimated object change data and the object change sequence to obtain statistical data, which represents the significant relationship between the feature change sequence and the object change sequence. Finally, based on the statistical data, the correlation between the changes in the feature change sequence and the changes in the object change sequence is determined, resulting in data analysis results. During the data analysis process, estimated object change data is directly obtained from the feature change sequence obtained from the business data, and model statistics are performed. The correlation between the changes in the object change data and the changes in the feature change sequence is analyzed using the statistical data. This allows for direct use of the raw data for the analysis process, eliminating the need for data identification and labeling, reducing the manpower cost of the solution, and improving the overall efficiency of the solution.
[0100] In one embodiment of this application, please refer to Figure 4 , Figure 4 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 4 As shown, based on the above technical solution, the feature change sequence includes P business change sequences, and the object change sequence includes P object change data; step S330 above, estimating based on the feature change sequence to generate object change estimation data, includes the following steps:
[0101] Step S410: Based on P business change sequences and P object change data, generate an impact factor. The impact factor represents the degree to which the Pth object change data is affected by the P business change sequences or by the P-1 object change data.
[0102] Step S420: Weight the P business change sequences according to the influencing factors to generate object change estimation data.
[0103] In estimation based on characteristic change sequences, a causal model is used to determine the degree of influence of historical object change data and business change sequences on the object change data to be analyzed. The determined result is expressed in the form of an impact factor. The causal model can be predetermined before data analysis or determined during the analysis process. The causal model can be implemented using machine learning models or computational formulas. Taking a computational formula as an example, assuming data from period tp to period t is collected, and the object change data is the retention transaction change rate {Y}... t-i |i=0,1,...,p}, and the characteristic change sequence is {X t-i,j Given a sequence of |i=0,1,...,p;j=1,...,m}, where the feature change sequence includes j business feature vectors, the causal model can be calculated using the following formula:
[0104]
[0105] Among them, Y t-i Y represents the vector of changes in retained transactions for the discounted fuel purchase during period ti. t X represents the vector of retained transaction change rates for discounted fuel services in period t. t-i,j Let b represent the j-th feature vector in period ti (including the rate of change in the number of customers participating in promotional activities, the rate of change in fuel transactions during holidays, the rate of change in click / favorite / comment behavior related to promotional fuel purchases, the rate of change in link data during the data collection process, the rate of change in prices, etc.). 0j This represents the influence factor of the j-th business characteristic in period t, i.e., the degree of influence; a0 represents the intercept term, a i b represents the influence factor of the fuel data vector in period ti. ij Let e represent the influencing factor of the j-th business characteristic in period ti. t Let represent the residual sequence vector at period t. In this example, by substituting the obtained p characteristic change sequences and p object change data into the causal model formula for solving, the corresponding influence factor A = {a} can be obtained. i |i=0,1,...,p} and B={b ij |i=0,1,...,p; j=1,...,m}.
[0106] In the embodiments of this application, a corresponding impact factor is calculated for each business change sequence included in the feature change sequence, thereby enabling a detailed breakdown of the impact of specific business change sequences on the object change data to be analyzed, improving the granularity of data analysis, and helping to improve the accuracy of the analysis results.
[0107] In one embodiment of this application, please refer to Figure 5 , Figure 5 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 5 As shown, based on the above technical solution, the impact factor includes a first impact parameter. Each business change sequence includes M business features, where M is an integer greater than or equal to 1. The first impact parameter includes M business parameters, which represent the degree of influence of the corresponding business feature on the object change data. Step S410 above generates the impact factor based on P business change sequences and P object change data, including the following steps:
[0108] Step S510: Determine the M business parameters based on the weighted sum of the M business parameters and the corresponding M business features, as well as the object change data corresponding to the business change sequence.
[0109] Step S520: Merge the business parameters corresponding to each business change sequence into a first influence parameter. The first influence parameter is a P×M matrix, where each element represents the degree of influence of the business features in the corresponding business change sequence on the object change data.
[0110] Specifically, each business change sequence includes M business features, which are characteristics obtained from business data that reflect changes in business conditions. In acquiring the business change sequence, data is extracted for each business feature to form the business change sequence. Therefore, the P business change sequences include the same business features, but the specific feature values depend on the specific data. Specifically, referring to equation (1) above, for each business change sequence, there is an equation relationship between the weighted sum of the M business features and the corresponding M business parameters and the corresponding object change data. For example, for the (P-3)th business change sequence, without considering the influence of the other (P-4) business change sequences preceding it in time sequence, the following equation should exist:
[0111]
[0112] Among them, Y M-3 This is the changed data for the (P-3)th object, X. M-3-i,j Let represent the j-th business feature in the P-3-i-th business change sequence. The other parameters have the same meaning as the parameters in equation (1). Substituting the M business features and object change data in the P-3-th business change sequence into equation (2) above for calculation, we can obtain the various business parameters b in the first influence parameter. ij For P business change sequences, iterative calculations are performed on each sequence to obtain P×M business parameters, which are then combined to form the first influencing parameter B = {b}. ij|i=0,1,...,p; j=1,...,m}.
[0113] In the embodiments of this application, corresponding business parameters are calculated for each business feature in the business change data, thereby enabling a detailed breakdown of the impact of each specific business feature on the object change data to be analyzed, improving the granularity of data analysis and enhancing the accuracy of the analysis results.
[0114] In one embodiment of this application, please refer to Figure 6 , Figure 6 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 6 As shown, based on the above technical solution, step S420, which weights the P business change sequences according to the influencing factors to generate object change estimation data, includes the following steps:
[0115] Step S610: Weight the corresponding business features in the corresponding business change sequence according to each element in the first influence parameter to obtain P×M weighted business features;
[0116] Step S620: Sum the P×M weighted business features to obtain P estimated data.
[0117] The data analysis device weights the corresponding business features in the business change sequence according to each element in the first influence parameter, thereby obtaining P×M weighted business features. Specifically, the first influence parameter B = {b ij The sequence {i = 0, 1, ..., p; j = 1, ..., m} contains P × M business parameters. Each business parameter is then compared with the business change sequence {X}. t-i,j By weighting the business features corresponding to the subsets |i=0,1,...,p;j=1,...,m}, we can obtain P×M weighted business features. Summing these P×M weighted business features yields the estimated business object change data, which can then be used as the estimated object change data. Specifically, the estimated object change data can be calculated as follows:
[0118]
[0119] in, This is data for estimating changes in business objects. Each business feature sequence can be used to calculate a corresponding estimate of object changes.
[0120] In the embodiments of this application, the object change estimation data is calculated by weighted summation, which can fully take into account the impact of each business feature in the business feature sequence on the object change data to be analyzed, and is conducive to improving the completeness of the analysis process.
[0121] In one embodiment of this application, please refer to Figure 7 , Figure 7 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 7 As shown, based on the above technical solution, step S340, which involves performing model statistics based on the object change estimation data and the object change sequence to obtain statistical data, includes the following steps:
[0122] Step S710: Based on the mapping relationship between the estimated object change data and the P object change data, determine the discrete relationship value. The discrete relationship value indicates the degree of dispersion between the estimated object change data and the object change data.
[0123] Step S720: Statistically sum the discrete relation values to obtain statistical data.
[0124] Specifically, discrete relationship values can be determined using forms such as variance, and the mapping relationship between the estimated object change data and the change data of P objects is determined by the variance calculation method. The data analysis device will calculate the statistical data of the P business change sequences according to a predetermined statistical calculation model. Specifically, the statistical data can be calculated according to the following equation:
[0125]
[0126] in, These are the calculated estimates of changes in P business objects. Let be the mean of the changing data of P objects, where After calculating the statistical data F corresponding to P business change sequences. X Then, the statistical data can be compared with the statistical thresholds to determine the sequence of business changes that affect the object's changing data. The statistical thresholds can be predetermined and are usually related to the number of periods for which the data is acquired, i.e., the amount of object's changing data. For example, the statistical thresholds can be obtained by looking up tables based on distributions such as chi-square distribution, normal distribution, or F-distribution.
[0127] In the embodiments of this application, statistical values are calculated by comparing the variance of the estimated object change data with that of P object change data, and then the data analysis results are determined based on the statistical values and statistical thresholds. This enables accurate analysis of the correlation between the business change sequence and the object change data, thereby ensuring the accuracy of the analysis results.
[0128] In one embodiment of this application, please refer to Figure 8 , Figure 8 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 8As shown, based on the above technical solution, step S350, which determines the correlation between the changes in the feature change sequence and the changes in the object change sequence based on statistical data, and obtains the data analysis results, includes the following steps:
[0129] Step S810: Input the quantity of object change data into the statistical distribution function to obtain the statistical threshold;
[0130] Step S820: If the statistical data is greater than the statistical threshold, then it is determined that the changes in the P business change sequences are related to the changes in the Pth object data.
[0131] Step S830: If the statistical data is less than or equal to the statistical threshold, then it is determined that there is no correlation between the P business change sequences and the change data of the Pth object.
[0132] The data analysis device inputs the quantity of change data of an object into a statistical distribution function for calculation, obtaining a statistical threshold. The statistical analysis function is determined based on the statistical test method used. For example, using F-statistics, the statistical threshold could be F... 0.95 (2, p-2). Then, the statistical values of the business change sequence are compared with the statistical threshold. If the statistical value of the business change sequence is greater than the statistical threshold F... X >F 0.95 If (2, p-2), then it is determined that P business change sequences affect the change data of the Pth object. Otherwise, if the statistical value of the business change sequence is less than or equal to the statistical threshold F... X ≤F 0.95 If (2, p-2), then it is determined that the P business change sequences have no impact on the change data of the Pth object. It is understandable that the calculation method for the statistical threshold can vary depending on the statistical test method used.
[0133] In the embodiments of this application, a specific implementation method is provided to determine the business change sequence that affects the change data of the object by comparing statistical values and statistical thresholds, thereby improving the feasibility of the solution.
[0134] In one embodiment of this application, please refer to Figure 9 , Figure 9 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 9 As shown, based on the above technical solution, after step S350, which determines the correlation between the changes in the feature change sequence and the changes in the object change sequence according to statistical data and obtains the data analysis results, the method further includes the following steps:
[0135] Step S910: Determine the object standard deviation of the P object change data;
[0136] Step S920: For the determined business change sequence, determine the ratio of the element in the first influence parameter corresponding to each business feature in the business change sequence to the object standard deviation, and obtain the feature statistics value of each business feature;
[0137] Step S930: Based on the comparison results of the feature statistics value and the threshold of the feature statistics value for each business feature, determine the correlation between the change of each business feature and the change of object data.
[0138] Specifically, the data analysis device calculates the standard deviation of the changes in P objects. Specifically, the object standard deviation can be calculated in the following way:
[0139]
[0140] Subsequently, for the business change sequence that is determined to have an impact on the object's changed data, the ratio of the element in the first influence parameter corresponding to each business feature in the business change sequence to the object's standard deviation is calculated to obtain the feature statistics of each business feature. Specifically, the feature statistics of each business feature are calculated in the following manner:
[0141]
[0142] Among them, B j ={b ij |i=1,...,p} represents the j-th business parameter of the business feature. According to equation (6), for the j-th business feature, the business parameters corresponding to the j-th feature of each of the P business change sequences are summed and then divided by the object standard deviation to form the feature statistics of the j-th business feature.
[0143] After calculating the feature statistics for each business feature, the data analysis device compares each feature statistics value with a feature statistics threshold to determine the correlation between changes in each business feature and changes in object data. The feature statistics threshold is determined based on the calculation method of the feature statistics and the number of business change sequences. Specifically, the feature statistics can be expressed as t... 0.95 (p), the specific value is determined by looking up a table. For the j-th business feature, if t j >t 0.95 If (p), then the business feature has an impact on the changed data of the Pth object; otherwise, the business feature has no impact on the changed data of the Pth object.
[0144] In the embodiments of this application, statistical calculations are performed on the business parameters of each business feature to determine the impact of each business feature on the object change data, thereby improving the granularity of the data analysis results and helping to improve the accuracy of data analysis.
[0145] In one embodiment of this application, please refer to Figure 10 , Figure 10 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 10 As shown, based on the above technical solution, the impact factor also includes a second impact parameter, which includes P-1 object parameters. Each object parameter represents the degree of influence of the corresponding object change data on the Pth object change data. Step S510, which determines the M business parameters based on the weighted sum of the M business parameters and the corresponding M business features, as well as the object change data corresponding to the business change sequence, includes the following steps:
[0146] Step S1010: Determine the M business parameters and P-1 object parameters based on the weighted sum of the M business parameters and their corresponding M business features, the weighted sum of the P-1 object parameters and their corresponding P-1 object change data, and the object change data corresponding to the business change sequence.
[0147] Step S510: After determining the M business parameters based on the weighted sum of the M business parameters and the corresponding M business features, as well as the object change data corresponding to the business change sequence, the method further includes the following steps:
[0148] Step S1020: Determine the second influence parameter based on the preset object parameters corresponding to the P-1 object parameters and the Pth object change data.
[0149] Among them, the P object change data are sorted in chronological order. The object change data that is later in the time series may be affected by the object change data that came before. The P-1 object parameters included in the second influence parameter correspond to the P-1 object change data that came before the P-th object change data. It can be understood that the P-1 object change data that came earlier in the time series can be considered historical data relative to the P-th object change data. The object parameter represents the degree of influence of the corresponding object change data on the P-th object change data. The data analysis device calculates the M business parameters and the P-1 object parameters based on the weighted sum of the M business parameters and the corresponding M business features, the weighted sum of the P-1 object parameters and the corresponding P-1 object change data, and the object change data corresponding to the business change sequence. Specifically, the calculation method of the M business parameters and the P-1 object parameters can be referred to the equation (1) introduced above for calculation, and the obtained object change data rate {Y} t-i |i=0,1,...,p} and business change sequence {X} t-i,j The values of |i=0,1,...,p;j=1,...,m} are input into equation (1) to solve for P-1 object parameters A={ai |i=1,...,p} and the first influence parameter B={b ij |i=0,1,...,p;j=1,...,m}. Then, the data analysis device calculates P-1 object parameters A={a i The second influencing parameter is determined by the preset object parameter a0 corresponding to the changed data of the P-th object |i=1,...,p}. Specifically, the P-1 object parameters A={a i By combining |i=1,...,p} with the preset object parameter a0, we can obtain the second influence parameter A={a i |i=0,1,...,p}.
[0150] In the embodiments of this application, during the calculation of the impact factor, the time series of object change data and the business change series are calculated together to obtain the first impact parameter corresponding to the business change series and the second impact parameter corresponding to the object change data. Thus, when determining the impact relationship of each business feature on the object change data, the impact from the time series is taken into account, which helps to improve the completeness of the data analysis results.
[0151] In one embodiment of this application, please refer to Figure 11 , Figure 11 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 11 As shown, based on the above technical solution, after estimating the object change data according to the feature change sequence in step S330, the method described in this application further includes the following steps:
[0152] Step S1110: Based on the object parameters in the second influence parameters, the corresponding object change data are weighted to obtain P weighted change data.
[0153] Step S1120: Sum the P weighted change data to obtain the change prediction data, which represents the result predicted based on the change data of P-1 objects.
[0154] After calculating the second influence parameter, the data analysis device uses the second influence parameter and the corresponding P object change data to estimate P weighted change data. Specifically, the data analysis device first calculates the second influence parameter A = {a...} i The parameters of each object in |i=0,1,...,p} and the corresponding object change data {Y} t-i We perform weighted calculations on the data points |i=0,1,...,2p} to obtain P weighted change data points. Then, we sum these P weighted change data points to obtain the estimated change data. The estimated change data can be calculated using the following equation:
[0155]
[0156] in, For estimating object data.
[0157] In the embodiments of this application, a weighted summation method is used to calculate the estimated object change data, which can fully take into account the impact of each historical object change data based on time series on the object change data to be analyzed, and is conducive to improving the completeness of the analysis process.
[0158] In one embodiment of this application, please refer to Figure 12 , Figure 12 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 12 As shown, based on the above technical solution, step S340, which involves performing model statistics based on the object change estimation data and the object change sequence to obtain statistical data, includes the following steps:
[0159] Step S1210: Based on the mapping relationship between the estimated object change data and the change data of P objects, determine the discrete relationship value. The discrete relationship value indicates the degree of dispersion between the estimated object change data and the object change data.
[0160] Step S1220: Statistically sum the discrete relation values to obtain statistical data.
[0161] Specifically, discrete relationship values can be determined using forms such as variance, and the mapping relationship between the estimated object change data and the change data of P objects is determined by the variance calculation method. Specifically, the data analysis device calculates the change data {Y} of the P objects based on a predetermined statistical calculation model. t-i The statistical data for |i=0,1,...,p}. Specifically, the statistical data can be calculated according to the following equation:
[0162]
[0163] in, These are the calculated data for P estimated objects. Let be the mean of the changing data of P objects, where The statistical data F corresponding to the changes in P objects is calculated. Y Then, the statistical data can be compared with the statistical threshold to determine the impact of changes in P-1 objects on changes in the Pth object. The statistical threshold can be predetermined and is usually related to the number of periods from which the data is obtained, i.e., to the number of object changes. For example, the statistical threshold can be obtained by looking up a table based on a chi-square distribution, a normal distribution, or an F-distribution.
[0164] In the embodiments of this application, statistical values are calculated by estimating the variance of the object data and the change data of P objects, and then the data analysis results are determined based on the statistical values and statistical thresholds. This enables accurate analysis of the correlation between historical object change data and the object change data to be analyzed, thereby ensuring the scope of influencing factors covered by the cause analysis process is improved, and thus improving the completeness of the attribution analysis results.
[0165] In one embodiment of this application, please refer to Figure 13 , Figure 13 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 13 As shown, based on the above technical solution, step S350, which determines the correlation between the changes in the feature change sequence and the changes in the object change sequence based on statistical data, includes the following steps:
[0166] Step S1310: Input the quantity of object change data into the statistical distribution function to obtain the statistical threshold;
[0167] Step S1320: If the statistical value of the object change data is greater than the statistical threshold, then it is determined that the changes of the P-1 object change data are related to the changes of the Pth object change data.
[0168] Step S1330: If the statistical value of the object change data is less than or equal to the statistical threshold, then it is determined that the changes of the P-1 object change data are not related to the changes of the Pth object change data.
[0169] The data analysis device inputs the quantity of change data of an object into a statistical distribution function for calculation, obtaining a statistical threshold. The statistical analysis function is determined based on the statistical test method used. For example, using F-statistics, the statistical threshold could be F... 0.95 (2, p-2). Then, the statistical values of the object change data are compared with the statistical threshold. If the statistical value of the object change data is greater than the statistical threshold F... Y >F 0.95 If (2, p-2), then the changes in P-1 objects affect the changes in the Pth object. Otherwise, if the statistical value of the object change data is less than or equal to the statistical threshold F... Y ≤F 0.95 If (2, p-2), then the changes in the P-1 objects have no effect on the changes in the P-th object. There is a correlation between the changes in the P-1 objects and the changes in the P-th object, which can typically be positive, negative, or linear. It is understandable that the calculation method for the statistical threshold can vary depending on the statistical test method used.
[0170] In the embodiments of this application, a specific implementation method is provided to determine the object change data that affects the object change data by comparing statistical values and statistical thresholds, thereby improving the feasibility of the solution.
[0171] In one embodiment of this application, please refer to Figure 14 , Figure 14 This is a schematic flowchart illustrating a data analysis method according to an embodiment of this application. Figure 14 As shown, based on the above technical solution, after step S350, which determines the correlation between the changes in the feature change sequence and the changes in the object change sequence according to the statistical data, and obtains the data analysis results, the method further includes the following steps:
[0172] Step S1410: Determine the object standard deviation of the P-1 object change data;
[0173] Step S1420: For P-1 object change data, determine the ratio of the object parameter to the object standard deviation corresponding to each object change data, and obtain the object statistics of P-1 object change data.
[0174] Step S1430: Based on the comparison results of the characteristic statistical values of the P-1 object change data and the object statistical value threshold, determine the impact of the P-1 object change data on the Pth object change data.
[0175] Specifically, the data analysis device calculates the object standard deviation of P-1 object variation data. Specifically, the object standard deviation can be calculated as follows:
[0176]
[0177] Subsequently, for the object change data that is determined to have an impact on the P-th object change data, the ratio of the element in the second influence parameter corresponding to each object change data in the P-1 object change data to the object standard deviation is calculated to obtain the feature statistics of each business feature. Specifically, the feature statistics of each business feature are calculated in the following manner:
[0178]
[0179] Among them, A={a l |l=1,...,p} represents the changed data of the l-th object.
[0180] After calculating the characteristic statistics of each object's change data, the data analysis device compares the characteristic statistics of P-1 object change data with a characteristic statistics threshold to determine the impact of each object change data on the overall object change data. The characteristic statistics threshold is determined based on the calculation method of the characteristic statistics and the number of business change sequences. Specifically, the characteristic statistics can be expressed as t... 0.95 (p) determines the specific value by looking up a table. For the l-th object's changed data, if t j >t 0.95 If (p), it means that the change in the l-th object has an impact on the change in the p-th object; otherwise, the change in the l-th object has no impact on the change in the p-th object.
[0181] In the embodiments of this application, statistical calculations are performed on P-1 object change data and corresponding object parameters to determine the impact of each business feature on the object change data. This takes into account the time series impact of the object change data, which helps to improve the completeness of data analysis and the coverage of factors.
[0182] The following uses gas station business data as an example to describe the complete process of the solution in this application. Specifically, assuming that the data source already contains t periods of business data, the data analysis device obtains p periods of data from the data source and extracts the rate of change in the number of retained customers among the participating customers from period tp to period t as the target data to be analyzed {Y}. t-i |i=0,1,...,p}, and obtain business characteristics such as customer changes during holidays, changes in customer clicks / favorites / comments on discounted refueling behavior, changes in data collection links, and price changes during the tp, ..., t periods as input feature data {X}. t-i,j |i=0,1,...,p;j=1,...,m}。 Then, based on the type and quantity of the acquired target data and input feature data, a causal model based on time series and feature data is constructed. The specific causal model can be implemented using equation (1) from the upper part of the equation. The target data {Y} t-i |i=0,1,...,p} and input feature data {X t-i,j Substituting |i=0,1,...,p;j=1,...,m} into equation (1) for regression calculation, after calculating for each target data, the model parameters A={a} of the causal model are obtained. i |i=0,1,...,p} and B={b ij|i=0,1,...,p;j=1,...,m}。 After obtaining the model parameters, the parameters can be used to estimate the target data, and then the estimated target data and the real target data can be used to perform causal analysis. The estimation calculation of the feature data is performed using equation (3) above, and the model parameters B={b ij |i=0,1,...,p;j=1,...,m} and feature data {X t-i,j The values |i=0,1,...,p;j=1,...,m} are input into equation (3) for calculation, thereby obtaining the target data sequence based on feature data estimation. The estimation of the target data is performed using equation (7) from the above text, with the model parameters A = {a}. i |i=0,1,...,p} and time series data {Y t-i The values of |i=0,1,...,2p} are input into equation (7) for calculation, thereby obtaining the target data predicted based on time series data. Subsequently, the data analysis device employs the F-statistical model to determine whether the feature data and time series data have an impact on the actual target sequence. The calculation method of the F-statistical model is based on equations (4) and (8) described above. Specifically, the data analysis device estimates the target data sequence based on the feature data. Compared with the real target data sequence {Y t-i Substituting |i=0,1,...,p} into equation (4) allows us to calculate the F-statistic F of the feature data. X The target data will be estimated based on time series data. Compared with the real target data sequence {Y t-i The F-statistic F is calculated using equation (8) on the expression |i=0,1,...,p}, thus obtaining the F-statistic F based on historical data from time series data. Y The statistical threshold F can be obtained by looking up the table based on the period p of the acquired data. 0.95 (2,p-2). F X and F Y respectively with F 0.95 (2, p-2) comparison, if F X >F 0.95 (2, p-2) indicates that feature data X has an impact on target data Y; otherwise, feature data X is not an influencing factor on target data Y. The same applies to historical data based on time series data. If F... Y >F 0.95 (2, p-2) represents the historical time series {Y} t-i For target data Y, |i=1,...,p} tThere is an impact; otherwise, the historical time series {Y} t-i |i=1,...,p} is not the target data Y t Influencing factors.
[0183] Based on the identified influencing factors, it is further possible to determine which specific characteristics among these factors affect the target data. Specifically, the data analysis device uses a T-statistical model to make this determination. The calculation formula for the T-statistical model can be obtained using equations (6) and (10) described above. For the feature data, the model parameter B = {b} ij |i=0,1,...,p;j=1,...,m} and the actual target data sequence {Y} t-i Substituting |i=0,1,...,p} into equation (6) allows us to calculate the various business features in the feature data. Statistical value t j The threshold t is obtained by querying a statistics table. 0.95 (p), if the statistical value is greater than the threshold t j >t 0.95 (p) indicates that the corresponding business feature is the target data Y. t If the influencing factors are not specified, then it indicates that the business characteristic is not the target data Y. t Influencing factors. The process for historical target data is similar, where the model parameters A = {a} are... i |i=1,...,p} and the actual target data sequence {Y} t-i Substituting |i=0,1,...,p} into equation (10) yields the time series-based statistic t. l The threshold t is obtained by querying the statistics table. 0.95 (p), if the statistical value is greater than the threshold t l >t 0.95 (p) indicates that the l-th historical target data is the target data Y. t The influencing factors must be identified; otherwise, the historical target data is not the target data Y. t Influencing factors.
[0184] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0185] The following describes the implementation of the apparatus of this application, which can be used to perform the data analysis method in the above embodiments of this application. Figure 15A schematic block diagram illustrating the composition of the data analysis apparatus in an embodiment of this application is shown. For example... Figure 15 As shown, the data analysis device 1500 mainly includes:
[0186] The feature change acquisition module 1510 is used to acquire the feature change sequence of the service, the feature change sequence including the service data collected within a preset time period;
[0187] The object change acquisition module 1520 is used to acquire the object change sequence of the business, wherein the object change sequence represents the result obtained based on the business data statistics.
[0188] Data estimation module 1530 is used to estimate based on the feature change sequence and generate object change estimation data, wherein the object change estimation data represents the result predicted based on the business data;
[0189] The model statistics module 1540 is used to perform model statistics based on the object change estimation data and the object change sequence to obtain statistical data, wherein the statistical data represents the significant relationship between the feature change sequence and the object change sequence;
[0190] The results analysis module 1550 is used to determine the correlation between the changes in the feature change sequence and the changes in the object change sequence based on the statistical data, and to obtain data analysis results.
[0191] In some embodiments of this application, based on the above technical solutions, the feature change sequence includes P business change sequences, and the object change sequence includes P object change data; the data estimation module 1530 includes:
[0192] The factor generation submodule is used to generate an influence factor based on the P business change sequences and the P object change data. The influence factor represents the degree to which the P-th object change data is affected by the P business change sequences or by the P-1 object change data.
[0193] The estimated data generation submodule is used to weight the P business change sequences according to the influencing factors to generate object change estimated data.
[0194] In some embodiments of this application, based on the above technical solutions, the influence factor includes a first influence parameter, each business change sequence includes M business features, where M is an integer greater than or equal to 1, the first influence parameter includes M business parameters, and the business parameters represent the degree of influence of the corresponding business features on the object change data; the factor generation submodule includes:
[0195] The business parameter determination unit is used to determine the M business parameters based on the weighted sum of the M business parameters and the corresponding M business features, as well as the object change data corresponding to the business change sequence.
[0196] The parameter generation unit is used to merge the business parameters corresponding to each business change sequence into the first influence parameter. The first influence parameter is a P×M matrix, and each element represents the degree of influence of the business features in the corresponding business change sequence on the object change data.
[0197] In some embodiments of this application, based on the above technical solutions, the estimated data generation submodule includes:
[0198] The business weighting unit is used to weight the corresponding business features in the corresponding business change sequence according to each element in the first influence parameter, so as to obtain P×M weighted business features.
[0199] The feature summation unit is used to sum the P×M weighted service features to obtain P estimated data.
[0200] In some embodiments of this application, based on the above technical solutions, the model statistics module 1540 includes:
[0201] The first discrete relation value determination submodule is used to determine discrete relation values based on the mapping relationship between the estimated object change data and the P object change data. The discrete relation values identify the degree of dispersion between the estimated object change data and the object change data.
[0202] The first statistical summation submodule is used to perform statistical summation on the discrete relation values to obtain the statistical data.
[0203] In some embodiments of this application, based on the above technical solutions, the result analysis module 1550 includes:
[0204] The first statistical distribution submodule is used to input the quantity of the object's change data into a statistical distribution function to obtain the statistical threshold.
[0205] The first association determination submodule is used to determine, if the statistical data is greater than the statistical threshold, that the changes in the P business change sequences are associated with the changes in the Pth object change data.
[0206] The first association determination submodule is further configured to determine that if the statistical data is less than or equal to the statistical threshold, then the P business change sequences are not associated with the change data of the Pth object.
[0207] In some embodiments of this application, based on the above technical solutions, the data analysis device further includes:
[0208] The first standard deviation determination module is used to determine the object standard deviation of P object change data;
[0209] The first statistical value determination module is used to determine, for the determined business change sequence, the ratio of the element in the first influence parameter corresponding to each business feature in the business change sequence to the standard deviation of the object, so as to obtain the feature statistical value of each business feature;
[0210] The first correlation determination module is used to determine the correlation between changes in each business feature and changes in object data based on the comparison results of the feature statistics value and the threshold of the feature statistics value for each business feature.
[0211] In some embodiments of this application, based on the above technical solutions, the influence factor further includes a second influence parameter, which includes P-1 object parameters, wherein the object parameters represent the degree of influence of the corresponding object change data on the Pth object change data; the business parameter determination unit includes:
[0212] The object parameter determination subunit is used to determine the M business parameters and the P-1 object parameters based on the weighted sum of the M business parameters and the corresponding M business features, the weighted sum of the P-1 object parameters and the corresponding P-1 object change data, and the object change data corresponding to the business change sequence.
[0213] The data analysis device also includes:
[0214] The second influence parameter determination subunit is used to determine the second influence parameter based on the calculated P-1 object parameters and the preset object parameters corresponding to the Pth object change data.
[0215] In some embodiments of this application, based on the above technical solutions, the data analysis device further includes:
[0216] The weighted change data determination module is used to weight the corresponding object change data according to each object parameter in the second influence parameter to obtain P weighted change data;
[0217] The change prediction data determination module is used to sum the P weighted change data to obtain change prediction data, which represents the result predicted based on the P-1 object change data.
[0218] In some embodiments of this application, based on the above technical solutions, the model statistics module 1540 includes:
[0219] The second discrete relation value determination submodule is used to determine discrete relation values based on the mapping relationship between the estimated object change data and the P object change data. The discrete relation values identify the degree of dispersion between the estimated object change data and the object change data.
[0220] The second statistical summation submodule is used to perform statistical summation on the discrete relation values to obtain the statistical data.
[0221] In some embodiments of this application, based on the above technical solutions, the result analysis module 1550 includes:
[0222] The second statistical distribution submodule inputs the quantity of the object's changing data into a statistical distribution function for calculation to obtain the statistical threshold.
[0223] The second association determination submodule is used to determine that if the statistical data is greater than the statistical threshold, there is an association between the changes in the P-1 object change data and the changes in the Pth object change data.
[0224] The second association determination submodule is further configured to determine that if the statistical data is less than or equal to the statistical threshold, the changes in the P-1 object change data are not associated with the changes in the Pth object change data.
[0225] In some embodiments of this application, based on the above technical solutions, the data analysis device further includes:
[0226] The second standard deviation determination module is used to determine the object standard deviation of the P-1 object change data;
[0227] The second statistical value determination module is used to calculate the ratio of the object parameter corresponding to each object change data to the object standard deviation for the P-1 object change data, and obtain the object statistical value of the P-1 object change data.
[0228] The second correlation determination module is used to determine the impact of the changes in the P-1 objects on the changes in the Pth object based on the comparison results of the characteristic statistical values of the P-1 object change data and the object statistical value threshold.
[0229] It should be noted that the apparatus provided in the above embodiments and the method provided in the above embodiments belong to the same concept, and the specific way in which each module performs the operation has been described in detail in the method embodiments, and will not be repeated here.
[0230] Figure 16 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.
[0231] It should be noted that, Figure 16 The computer system 1600 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0232] like Figure 16 As shown, the computer system 1600 includes a Central Processing Unit (CPU) 1601, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1602 or programs loaded from Storage Unit 1608 into Random Access Memory (RAM) 1603. The RAM 1603 also stores various programs and data required for system operation. The CPU 1601, ROM 1602, and RAM 1603 are interconnected via a bus 1604. An Input / Output (I / O) interface 1605 is also connected to the bus 1604.
[0233] The following components are connected to I / O interface 1605: an input section 1606 including a keyboard, mouse, etc.; an output section 1607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to I / O interface 1605 as needed. Removable media 1611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1610 as needed so that computer programs read from them can be installed into storage section 1608 as needed.
[0234] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1609, and / or installed from removable medium 1611. When the computer program is executed by central processing unit (CPU) 1601, it performs various functions defined in the system of this application.
[0235] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0236] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0237] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0238] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.
[0239] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0240] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A data analysis method, characterized in that, include: Acquire a feature change sequence of the service, the feature change sequence including service data collected within a preset time period, the feature change sequence including P service change sequences; Obtain the object change sequence of the business, the object change sequence represents the result obtained based on the business data statistics, and the object change sequence includes P object change data; Based on the P business change sequences and the P object change data, an impact factor is generated, wherein the impact factor represents the degree to which the P-th object change data is affected by the P business change sequences or by the P-1 object change data. The P business change sequences are weighted according to the influencing factors to generate object change estimation data, which represents the result predicted based on the business data. Based on the estimated object change data and the object change sequence, model statistics are performed to obtain statistical data, which represents the significant relationship between the feature change sequence and the object change sequence; Based on the statistical data, the correlation between the changes in the feature change sequence and the changes in the object change sequence is determined, and the data analysis results are obtained.
2. The method according to claim 1, characterized in that, The influencing factor includes a first influencing parameter. Each business change sequence includes M business features, where M is an integer greater than or equal to 1. The first influencing parameter includes M business parameters, where each business parameter represents the degree of influence of the corresponding business feature on the object change data. The step of generating influencing factors based on the P business change sequences and the P object change data includes: The M business parameters are determined based on the weighted sum of the M business parameters and the corresponding M business features, as well as the object change data corresponding to the business change sequence. The business parameters corresponding to each business change sequence are merged into the first influence parameter, which is a P×M matrix, where each element represents the degree of influence of the business characteristics in the corresponding business change sequence on the object change data.
3. The method according to claim 2, characterized in that, The step of weighting the P business change sequences according to the influencing factors to generate object change estimation data includes: Based on each element in the first influence parameter, the corresponding business features in the corresponding business change sequence are weighted to obtain P×M weighted business features; The summation of the P×M weighted business features yields P estimated data points.
4. The method according to claim 2, characterized in that, The statistical data obtained by performing model statistics based on the estimated object change data and the object change sequence includes: Based on the mapping relationship between the estimated object change data and the P object change data, a discrete relationship value is determined, and the discrete relationship value identifies the degree of dispersion between the estimated object change data and the object change data. The discrete relation values are statistically summed to obtain the statistical data.
5. The method according to claim 4, characterized in that, The step of determining the correlation between the changes in the feature change sequence and the changes in the object change sequence based on the statistical data, and obtaining data analysis results, includes: Input the quantity of the object's change data into a statistical distribution function to obtain a statistical threshold; If the statistical data is greater than the statistical threshold, then it is determined that the changes in the P business change sequences are related to the changes in the Pth object change data; If the statistical data is less than or equal to the statistical threshold, then it is determined that the P business change sequences are not related to the change data of the Pth object.
6. The method according to claim 4, characterized in that, After determining the correlation between the changes in the feature change sequence and the changes in the object change sequence based on the statistical data, and obtaining the data analysis results, the method further includes: Determine the standard deviation of the change data for P objects; For the determined business change sequence, determine the ratio of the element in the first influence parameter corresponding to each business feature in the business change sequence to the standard deviation of the object, and obtain the feature statistics of each business feature; Based on the comparison results of the feature statistics of each business feature and the threshold of the feature statistics, the correlation between the changes in each business feature and the changes in object data is determined.
7. The method according to claim 2, characterized in that, The impact factor further includes a second impact parameter, which includes P-1 object parameters, each object parameter representing the degree of influence of the corresponding object change data on the Pth object change data; determining the M business parameters based on the weighted sum of the M business parameters and the corresponding M business features, and the object change data corresponding to the business change sequence, includes: The M business parameters and the P-1 object parameters are determined based on the weighted sum of the M business parameters and the corresponding M business features, the weighted sum of the P-1 object parameters and the corresponding P-1 object change data, and the object change data corresponding to the business change sequence. After determining the M business parameters based on the weighted sum of the M business parameters and the corresponding M business features, and the object change data corresponding to the business change sequence, the method further includes: The second influence parameter is determined based on the P-1 object parameters obtained and the preset object parameters corresponding to the Pth object change data.
8. The method according to claim 7, characterized in that, After weighting the P business change sequences according to the influencing factors to generate object change estimation data, the method further includes: Based on the object parameters in the second influence parameter, the corresponding object change data are weighted to obtain P weighted change data; The P weighted change data are summed to obtain the change prediction data, which represents the result predicted based on the P-1 object change data.
9. The method according to claim 7, characterized in that, The statistical data obtained by performing model statistics based on the estimated object change data and the object change sequence includes: Based on the mapping relationship between the estimated object change data and the P object change data, a discrete relationship value is determined, and the discrete relationship value identifies the degree of dispersion between the estimated object change data and the object change data. The discrete relation values are statistically summed to obtain the statistical data.
10. The method according to claim 8, characterized in that, The step of determining the correlation between the changes in the feature change sequence and the changes in the object change sequence based on the statistical data, and obtaining data analysis results, includes: Input the quantity of the object's change data into a statistical distribution function to obtain a statistical threshold; If the statistical data is greater than the statistical threshold, then it is determined that the changes in the P-1 object change data are related to the changes in the Pth object change data; If the statistical data is less than or equal to the statistical threshold, then it is determined that the changes in the P-1 object change data are not related to the changes in the Pth object change data.
11. The method according to claim 8, characterized in that, After determining the correlation between the changes in the feature change sequence and the changes in the object change sequence based on the statistical data, and obtaining the data analysis results, the method further includes: Determine the object standard deviation of the P-1 object change data; For the P-1 object change data, determine the ratio of the object parameter corresponding to each object change data to the object standard deviation, and obtain the object statistics of the P-1 object change data; Based on the comparison results of the characteristic statistical values of the P-1 object change data and the object statistical value threshold, the influence of the P-1 object change data on the Pth object change data is determined.
12. A data analysis device, characterized in that, include: The feature change acquisition module is used to acquire the feature change sequence of the business, the feature change sequence includes business data collected within a preset time period, and the feature change sequence includes P business change sequences; The object change acquisition module is used to acquire the object change sequence of the business, the object change sequence represents the result obtained based on the business data statistics, and the object change sequence includes P object change data; The factor generation submodule is used to generate an influence factor based on the P business change sequences and the P object change data. The influence factor represents the degree to which the P-th object change data is affected by the P business change sequences or by the P-1 object change data. The estimation data generation submodule is used to weight the P business change sequences according to the influencing factors to generate object change estimation data, wherein the object change estimation data represents the result predicted based on the business data; The model statistics module is used to perform model statistics based on the object change estimation data and the object change sequence to obtain statistical data, which represents the significant relationship between the feature change sequence and the object change sequence. The results analysis module is used to determine the correlation between the changes in the feature change sequence and the changes in the object change sequence based on the statistical data, and to obtain data analysis results.
13. An electronic device, characterized in that, include: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the data analysis method of any one of claims 1 to 11 by executing the executable instructions.
14. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data analysis method as described in any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, a processor of a computer device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions to cause the computer device to perform the data analysis method as described in any one of claims 1 to 11.