Data prediction method, device and equipment and storage medium thereof

By using principal component analysis to screen important features of financial and medical business data, constructing a data matrix and filling missing values, and using a continuous estimation model to set dynamic goals, we solved the problem of insufficient flexibility in traditional methods and achieved efficient business goal setting.

CN120687812APending Publication Date: 2025-09-23CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510585321.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies lack flexibility and foresight in setting business goals, rely on historical data, lead to deviations, cannot adjust strategies in a timely manner, are difficult to process data, and traditional methods ignore changes in the market environment.

Method used

Use principal component analysis to screen important feature data, construct a data matrix, identify and fill missing values, use a continuous estimation model to set dynamically changing task goals, and conduct analysis in combination with real-time business data.

Benefits of technology

It enables data analysis on important feature dimensions, reduces the difficulty of data processing, improves analysis efficiency, and provides scientific business goal setting, which is suitable for financial and medical business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687812A_ABST
    Figure CN120687812A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of big data processing, and relates to a data prediction method, device and equipment and a storage medium thereof.The method includes the steps that firstly, a principal component analysis method is adopted to analyze a small amount of current analysis data, important feature dimensions are determined, feature extraction is conducted on historical data on the important feature dimensions, and a trend line is generated; according to the method, data matrix missing value supplementation is performed in combination with a trend line, a continuous prediction model is constructed by using a data matrix without missing values, and a service prediction model is constructed only in combination with important dimension features, so that processing of unimportant dimension data is avoided, and specific services are more efficiently assisted to perform data analysis and service target setting. The data prediction method is specifically applied to a financial business data analysis scene, and a business target can be set for business personnel more efficiently and quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data processing technology, and is applied to business or market data forecasting and analysis scenarios when used in financial or medical business services, and involves a data forecasting method, device, equipment and storage medium thereof. Background Art

[0002] With the development of big data technology, business forecasting often requires integrating big data, for example, forecasting business goals based on historical batches of business data. During data analysis, setting forecast targets is crucial for ensuring healthy business development. Currently, monthly and daily targets are typically set by fine-tuning them based on historical performance over the same period. This approach lacks rationality and is cumbersome and requires significant data processing.

[0003] Furthermore, with the influence of cyclical and uncertain factors such as market fluctuations, competitive landscape, and the diversity of customer behavior, traditional goal-setting methods lack flexibility and adjustment mechanisms. Goal-setting methods are too rigid and ignore changes in the market environment. This may result in the agent team being unable to adjust their strategies in a timely manner when faced with emergencies or market changes, affecting goal achievement. Furthermore, over-reliance on historical data can lead to biases, lack of foresight, and inability to integrate more specific real-time business data, which also increases the difficulty of data processing. Therefore, there is an urgent need for a data prediction method that can achieve scientific goal setting while paying more attention to specific real-time business data and reducing the difficulty of data processing. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to propose a data prediction method, device, equipment and storage medium thereof, so as to achieve scientific setting of the goal while paying more attention to specific real-time business data and reducing the difficulty of data processing.

[0005] In a first aspect, the present invention provides a data prediction method, which adopts the following technical solution:

[0006] A data prediction method comprises the following steps:

[0007] Real-time collection of data to be analyzed within the current preset time period;

[0008] The principal component analysis method is used to perform principal component analysis on the data to be analyzed, and the characteristic data of the target quantity dimension is screened out in combination with a preset component influence threshold;

[0009] Constructing a data matrix based on the characteristic data of the target quantity dimension;

[0010] Extracting feature data from the reference data within the historical reference time period according to the target quantity dimension to obtain historical feature data trend lines corresponding to the target quantity dimension respectively;

[0011] Identifying whether there are missing values ​​in the data matrix;

[0012] If there are missing values ​​in the data matrix, then according to the data dimension corresponding to the missing value, a trend line change similarity algorithm is used to filter out filling values ​​from the corresponding historical feature data trend line, and fill them into the data matrix to obtain the target data matrix;

[0013] A continuous prediction model is constructed using the target data matrix, and dynamically changing task objectives are set according to continuous output results of the continuous prediction model.

[0014] In a second aspect, the present application also provides a data prediction device, which adopts the following technical solution:

[0015] A data prediction device, comprising:

[0016] The data collection module for analysis is used to collect the data to be analyzed within the current preset time period in real time;

[0017] A principal component analysis module is used to perform principal component analysis on the data to be analyzed using a principal component analysis method, and to filter out characteristic data of a target quantity dimension in combination with a preset component influence threshold;

[0018] A data matrix construction module, configured to construct a data matrix based on the characteristic data of the target quantity dimension;

[0019] A historical characteristic data trend line acquisition module is used to extract characteristic data from the reference data within the historical reference time period according to the target quantity dimension, and obtain the historical characteristic data trend lines corresponding to the target quantity dimension respectively;

[0020] A missing value identification module is used to identify whether there are missing values ​​in the data matrix;

[0021] A missing value filling module is used to, if there are missing values ​​in the data matrix, select filling values ​​from the corresponding historical feature data trend lines using a trend line change similarity algorithm based on the data dimensions corresponding to the missing values, and fill them into the data matrix to obtain a target data matrix;

[0022] The continuous task target estimation module is used to construct a continuous estimation model using the target data matrix and set dynamically changing task targets according to the continuous output results of the continuous estimation model.

[0023] In a third aspect, an embodiment of the present application further provides a computer device that adopts the following technical solution:

[0024] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the data prediction method described above when executing the computer-readable instructions.

[0025] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0026] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data prediction method described above.

[0027] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0028] The data prediction method described in the embodiment of the present application collects the data to be analyzed within the current preset time period in real time; uses the principal component analysis method to perform principal component analysis on the data to be analyzed, and combines the preset component influence threshold to screen out the characteristic data of the target quantity dimension; constructs a data matrix based on the characteristic data of the target quantity dimension; extracts characteristic data of the reference data within the historical reference time period according to the target quantity dimension to obtain the historical characteristic data trend line corresponding to the target quantity dimension; identifies whether there are missing values ​​in the data matrix; if there are missing values ​​in the data matrix, fills them to obtain the target data matrix; uses the target data matrix to construct a continuous estimation model, and sets the dynamically changing task target according to the continuous output results of the continuous estimation model. It realizes that the principal component analysis method is first used to analyze a small amount of current analysis data to determine the important characteristic dimensions, then extracts the characteristics of the historical data on the important characteristic dimensions to generate trend lines, and supplements the missing values ​​of the data matrix based on the trend lines, and constructs a continuous estimation model based on the data matrix without missing values. The business estimation model is constructed based on only the important dimension features, avoiding the processing of non-important dimension data, and more efficiently assisting specific businesses in data analysis and business goal setting. The data prediction method is specifically applied to financial business data analysis scenarios, such as insurance renewal rate prediction scenarios and stock market data fluctuation prediction scenarios in the financial business field, which can more efficiently and quickly set business goals for business personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0030] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0031] Figure 2 is a flow chart of an embodiment of a data prediction method according to the present application;

[0032] Figure 3 yes Figure 2 A flowchart of a specific embodiment of step 203 is shown;

[0033] Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 204 is shown;

[0034] Figure 5 yes Figure 4 A flowchart of a specific embodiment of step 404 is shown;

[0035] Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 205 is shown;

[0036] Figure 7 yes Figure 2 A flowchart of a specific embodiment of step 206 is shown;

[0037] Figure 8 is a structural diagram of an embodiment of a data prediction device according to the present application;

[0038] Figure 9 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0040] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0041] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0042] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0043] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0044] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.

[0045] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .

[0046] It should be noted that the data prediction method provided in the embodiment of the present application is generally executed by a server, and accordingly, a data prediction device is generally set in the server.

[0047] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0048] Continue to refer Figure 2 , shows a flow chart of an embodiment of a data prediction method according to the present application. The data prediction method comprises the following steps:

[0049] Step 201: collect data to be analyzed within a current preset time period in real time.

[0050] In this embodiment, the data prediction method is specifically applied to financial business data analysis scenarios, for example, insurance renewal rate prediction scenarios and stock market data fluctuation prediction scenarios in the financial business field.

[0051] Specifically, taking the insurance renewal rate prediction scenario as an example, step 201 is to collect the renewal rate representation data within the current preset time period in real time as the data to be analyzed; similarly, if taking the stock market data fluctuation prediction scenario as an example, step 201 is to collect the target stock change representation data within the current preset time period in real time as the data to be analyzed, and no further examples will be given here.

[0052] In this embodiment, the current preset time period is specifically a time period formed from a preset historical time point to the current time point, for example, a time period formed from a date half a year ago to a date corresponding to the current time.

[0053] By collecting the data to be analyzed within the current preset time period in real time, it is convenient to analyze the latest data, thereby providing the latest business decision-making basis for related businesses.

[0054] Step 202: principal component analysis is performed on the data to be analyzed using a principal component analysis method, and characteristic data of a target quantity dimension is screened out in combination with a preset component influence threshold.

[0055] Specifically, the principal component analysis method includes the PCA principal component analysis method, which performs principal component analysis on the data to be analyzed, analyzes several influencing dimensions that are more important for the changes in the data to be analyzed, and screens out characteristic data of the target quantity dimension based on a preset component influence threshold.

[0056] In this embodiment, it is assumed that the principal component analysis method is used to analyze 10 influencing dimensions that are more important in affecting the changes in the data to be analyzed. There are certain differences in the influence thresholds corresponding to each influencing dimension, that is, there are certain differences in the influence of different influencing dimensions on data changes. For example, the influence values ​​corresponding to the above 10 influencing dimensions range from 95% to 50%, among which the influence values ​​corresponding to the three influencing dimensions A, B and C are above 90%. Assuming that the preset component influence threshold is 90%, the feature data of the three dimensions A, B and C will be screened out based on the influence value and the preset component influence threshold.

[0057] By using principal component analysis (PCA) to analyze the data to be analyzed and, combined with a preset component influence threshold, filtering out feature data from a target number of dimensions, the analysis focuses solely on the key dimensional feature data, eliminating the need for excessive processing of unimportant dimensions and feature data. This reduces the complexity and difficulty of data analysis or prediction, particularly in large and complex financial business data analysis scenarios.

[0058] Step 203: construct a data matrix based on the feature data of the target quantity dimension.

[0059] Specifically, a data matrix is ​​constructed for the characteristic data of several important dimensions. For example, in an insurance renewal rate prediction scenario, if four dimensions of characteristic data are selected, a data matrix is ​​constructed only for these four dimensions. When constructing the data matrix, the time series is taken into account, with accuracy down to each day of the month, and the rows of the data matrix are used as the rows, and the columns of the data matrix are used as the columns. The time series state of the data matrix construction can be customized according to business needs.

[0060] By constructing a data matrix based on the characteristic data of the target number dimensions, it is convenient to subsequently perform data analysis on the characteristic data of relevant important dimensions under the representation of the data matrix, thereby achieving an overall analysis of the data to be analyzed.

[0061] Step 204 : extracting characteristic data from the reference data within the historical reference time period according to the target quantity dimension, and obtaining historical characteristic data trend lines corresponding to the target quantity dimension.

[0062] Compared to the previous approach of collecting all reference data within a historical reference time period, this embodiment extracts feature data from the reference data within the historical reference time period based solely on the target quantity dimension. This significantly reduces the amount of reference data collected, and the resulting historical feature data trend lines are only those corresponding to relatively important dimensions. This avoids blindly collecting reference data. In particular, in large data collection and analysis scenarios, this can significantly reduce the amount of non-critical data collected, improving the efficiency of data collection, processing, and analysis.

[0063] Step 205: Identify whether there are missing values ​​in the data matrix.

[0064] Step 206: If there are missing values ​​in the data matrix, fill values ​​are selected from the corresponding historical feature data trend lines using a trend line change similarity algorithm based on the data dimensions corresponding to the missing values, and the fill values ​​are filled into the data matrix to obtain a target data matrix.

[0065] Specifically, a trend line change similarity algorithm is used to screen out filling values, thereby realizing the filling of missing data in the form of trend line comparison, which is more reasonable and scientific, avoids the use of the existing commonly used average method or interpolation method to directly fill in missing values, and is more reliable.

[0066] Step 207 : constructing a continuous estimation model using the target data matrix, and setting dynamically changing task objectives according to continuous output results of the continuous estimation model.

[0067] Of course, if there are no missing values ​​in the data matrix, the data matrix is ​​directly set as the target data matrix and step 207 is executed.

[0068] Specifically, a continuous prediction model is constructed using a data matrix that has undergone missing value processing, and dynamically changing task objectives are set based on the continuous output results of the continuous prediction model. This allows subsequent business data forecasts to be made based on the combination of important dimensions and historical data, providing customers with a more efficient basis for business decision-making.

[0069] In this embodiment, data to be analyzed within a preset time period is collected in real time; principal component analysis is performed on the data to be analyzed, and characteristic data of the target quantity dimension is screened out based on a preset component impact threshold; a data matrix is ​​constructed based on the characteristic data of the target quantity dimension; characteristic data of reference data within a historical reference time period is extracted based on the target quantity dimension to obtain historical characteristic data trend lines corresponding to the target quantity dimension; missing values ​​are identified in the data matrix; if missing values ​​exist in the data matrix, missing values ​​are filled to obtain a target data matrix; a continuous prediction model is constructed using the target data matrix, and dynamically changing task objectives are set based on the continuous output results of the continuous prediction model. The embodiment first uses principal component analysis to analyze a small amount of current analysis data to determine important characteristic dimensions, then extracts features from historical data on the important characteristic dimensions to generate trend lines, and uses the trend lines to supplement missing values ​​in the data matrix. A continuous prediction model is constructed using a data matrix without missing values, and the business prediction model is constructed based only on the characteristics of important dimensions, avoiding the processing of non-important dimension data and more efficiently assisting specific businesses in data analysis and business goal setting. The data prediction method is specifically applied to financial business data analysis scenarios, such as insurance renewal rate prediction scenarios and stock market data fluctuation prediction scenarios in the financial business field, which can more efficiently and quickly set business goals for business personnel.

[0070] In this embodiment, before executing the step of constructing a data matrix based on the feature data of the target quantity dimension, that is, before executing step 203, the method also includes: using the Z-score normalization algorithm to normalize the feature data of the target quantity dimension according to the dimension category, to obtain the feature data corresponding to the target quantity dimension after eliminating the dimensional influence.

[0071] By using the Z-score standardization algorithm to perform standardization processing on the characteristic data of the target quantity dimension, the dimension effect of the characteristic data is eliminated, which facilitates the subsequent use of the characteristic data to construct a data matrix.

[0072] Continue to refer Figure 3 , Figure 3 yes Figure 2 The flowchart of a specific embodiment of step 203 shown includes the following steps:

[0073] Step 301: updating the characteristic data corresponding to the target quantity dimensions after eliminating the dimension effect to the characteristic data for constructing the data matrix;

[0074] Step 302: construct the feature data of the data matrix for each dimension and generate a feature vector according to a preset time unit;

[0075] Specifically, the feature vector is generated according to the preset time unit. For example, the feature vector is generated according to the characterization value corresponding to the feature data at the specific time of each day of each month, and a time-series data matrix is ​​constructed. If there are 31 days in January, a 1x31 data matrix containing 31 feature vectors in one row is constructed, and each element in the data matrix represents the characterization value on different dates in January.

[0076] Step 303: Fill the feature vectors into the corresponding data matrix in chronological order to obtain the data matrix corresponding to the feature data of each dimension.

[0077] By constructing a data matrix based on the characteristic data of the target number dimensions, it is convenient to subsequently perform data analysis on the characteristic data of relevant important dimensions under the representation of the data matrix, thereby achieving an overall analysis of the data to be analyzed.

[0078] Continue to refer Figure 4 , Figure 4 yes Figure 2 The flowchart of a specific embodiment of step 204 shown includes the following steps:

[0079] Step 401: Determine, through the data matrix, the target number of data dimensions as the main dimension information of the data to be analyzed;

[0080] Step 402: Extract feature data from the reference data within the historical reference time period using the main dimension information of the data to be analyzed as an extraction dimension to obtain feature data corresponding to all main dimension information of the reference data within the historical reference time period.

[0081] Step 403: Using a Z-score normalization algorithm, the feature data corresponding to all major dimensional information of the reference data within the historical reference time period are normalized according to the dimensional categories to obtain the feature data corresponding to all major dimensional information of the reference data within the historical reference time period after eliminating the dimensional effect.

[0082] Step 404 : constructing the historical characteristic data trend line based on the characteristic data corresponding to the reference data in the historical reference time period in all main dimensional information after eliminating the dimensional influence.

[0083] Specifically, based on the data matrix, the primary dimensional information for data analysis and prediction is determined. Based on this primary dimensional information, corresponding feature data is filtered from the historical data. The filtered feature data is then normalized using the same standardization method to eliminate dimensionality effects and construct a corresponding historical feature data trend line. This facilitates subsequent data prediction and analysis in conjunction with the corresponding historical feature data trend line.

[0084] Continue to refer Figure 5 , Figure 5 yes Figure 4 The flowchart of a specific embodiment of step 404 includes the following steps:

[0085] Step 501: Inputting the characteristic data corresponding to the reference data in the historical reference time period in all main dimensional information after eliminating the dimension effect as graphical input parameters into a preset data graphical processing component;

[0086] Step 502: Using the preset time unit as a timing control parameter of the preset data graphics processing component;

[0087] Step 503 : Generate a historical characteristic data trend line corresponding to each main data dimension according to the timing control parameter and the graphical input parameter.

[0088] Specifically, when generating the historical characteristic data trend line corresponding to each main data dimension, a preset data graphical processing component can be used for graphical processing, thereby achieving a more intuitive acquisition of the historical characteristic data trend line.

[0089] Continue to refer Figure 6 , Figure 6 yes Figure 2 The flowchart of a specific embodiment of step 205 shown includes the following steps:

[0090] Step 601, identifying each element in the data matrix using a traversal method;

[0091] Step 602: Based on the recognition result, determine whether the element value corresponding to the current row and column intersection is a NULL value or an empty value;

[0092] Step 603: If the element values ​​corresponding to all intersections of rows and columns are neither NULL nor empty, then there are no missing values ​​in the data matrix.

[0093] Step 604: If the element value corresponding to the intersection of a row and a column is a NULL value or an empty value, then there is a missing value in the data matrix.

[0094] By identifying whether there are missing values ​​in the data matrix, it is possible to decide whether to perform filling processing on the data matrix in the future.

[0095] Continue to refer Figure 7 , Figure 7 yes Figure 2 The flowchart of a specific embodiment of step 206 includes the following steps:

[0096] Step 701: Determine the data dimension information and time series information corresponding to the missing value based on the row and column position information of the missing value in the data matrix;

[0097] Step 702: constructing a data trend line with breakpoints corresponding to the data to be analyzed on the data dimension information according to the data dimension information and time series information corresponding to the missing value;

[0098] Step 703: Filter out the historical characteristic data trend line corresponding to the data dimension information from the historical characteristic data trend lines corresponding to all main data dimensions based on the data dimension information;

[0099] Step 704: Using a trend line change similarity algorithm, perform regression comparison on the data trend line with the breakpoint and the historical characteristic data trend line, and select the trend line segment that is most similar to the historical characteristic data trend line corresponding to the data trend line with the breakpoint;

[0100] Step 705: supplementing the feature data of all breakpoints in the data trend line having breakpoints according to the feature data corresponding to each point in the most similar trend line segment, and generating corresponding fill values ​​according to the supplemented feature data;

[0101] Step 706 , fill the corresponding filling values ​​into the data according to the row and column position information corresponding to all breakpoints to obtain the target data matrix.

[0102] In this embodiment, missing data is filled in the form of trend line comparison, which is more reasonable and scientific, avoids the use of the existing commonly used average method or interpolation method to directly fill in missing values, and is more reliable.

[0103] In this embodiment, the step of constructing a continuous estimation model using the target data matrix and setting dynamically changing task objectives based on the continuous output results of the continuous estimation model specifically includes: using the target data matrix as a parameter matrix for constructing the continuous estimation model, and constructing the continuous estimation model using a regularized Ridge algorithm, wherein, when constructing the continuous estimation model using the regularized Ridge algorithm, a cross-validation method is used to screen out the maximum regularization parameter, and model construction control settings are performed to reduce model complexity.

[0104] Specifically, when using the regularized Ridge algorithm to build a model, the larger the regularization parameter, the lower the model complexity. In this embodiment, the regularization parameter is used as the model building control parameter when it is at its maximum value to maximize the reduction of model complexity.

[0105] In this embodiment, after executing the steps of constructing a continuous estimation model using the target data matrix and setting dynamically changing task goals based on the continuous output results of the continuous estimation model, the method also includes: judging whether the actual task result corresponding to the current time point exceeds the estimated task goal; if the actual task result corresponding to the current time point exceeds the estimated task goal, continuing to monitor the completion of the task goal; if the actual task result corresponding to the current time point does not exceed the estimated task goal, judging whether the actual task result reaches a preset warning threshold compared with the estimated task goal; if the actual task result does not reach the preset warning threshold compared with the estimated task goal, sending an email reminder to the target business processing end; if the actual task result reaches the preset warning threshold compared with the estimated task goal, continuing to monitor the completion of the task goal.

[0106] By setting warning thresholds and task targets, the actual task results can be monitored. When the actual task results are far below or fail to meet the requirements, an early warning will be sent to the monitoring end, so that the background can make corresponding adjustments to the business processing end in a timely manner.

[0107] In this embodiment, data to be analyzed within a preset time period is collected in real time; principal component analysis is performed on the data to be analyzed, and characteristic data of the target quantity dimension is screened out based on a preset component impact threshold; a data matrix is ​​constructed based on the characteristic data of the target quantity dimension; characteristic data of reference data within a historical reference time period is extracted based on the target quantity dimension to obtain historical characteristic data trend lines corresponding to the target quantity dimension; missing values ​​are identified in the data matrix; if missing values ​​exist in the data matrix, missing values ​​are filled to obtain a target data matrix; a continuous prediction model is constructed using the target data matrix, and dynamically changing task objectives are set based on the continuous output results of the continuous prediction model. The embodiment first uses principal component analysis to analyze a small amount of current analysis data to determine important characteristic dimensions, then extracts features from historical data on the important characteristic dimensions to generate trend lines, and uses the trend lines to supplement missing values ​​in the data matrix. A continuous prediction model is constructed using a data matrix without missing values, and the business prediction model is constructed based only on the characteristics of important dimensions, avoiding the processing of non-important dimension data and more efficiently assisting specific businesses in data analysis and business goal setting. The data prediction method is specifically applied to financial business data analysis scenarios, such as insurance renewal rate prediction scenarios and stock market data fluctuation prediction scenarios in the financial business field, which can more efficiently and quickly set business goals for business personnel.

[0108] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0109] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0110] In this embodiment, data to be analyzed within a preset time period is collected in real time; principal component analysis is performed on the data to be analyzed, and characteristic data of the target quantity dimension is screened out based on a preset component impact threshold; a data matrix is ​​constructed based on the characteristic data of the target quantity dimension; characteristic data of reference data within a historical reference time period is extracted based on the target quantity dimension to obtain historical characteristic data trend lines corresponding to the target quantity dimension; missing values ​​are identified in the data matrix; if missing values ​​exist in the data matrix, missing values ​​are filled to obtain a target data matrix; a continuous prediction model is constructed using the target data matrix, and dynamically changing task objectives are set based on the continuous output results of the continuous prediction model. The embodiment first uses principal component analysis to analyze a small amount of current analysis data to determine important characteristic dimensions, then extracts features from historical data on the important characteristic dimensions to generate trend lines, and uses the trend lines to supplement missing values ​​in the data matrix. A continuous prediction model is constructed using a data matrix without missing values, and the business prediction model is constructed based only on the characteristics of important dimensions, avoiding the processing of non-important dimension data and more efficiently assisting specific businesses in data analysis and business goal setting. The data prediction method is specifically applied to financial business data analysis scenarios, such as insurance renewal rate prediction scenarios and stock market data fluctuation prediction scenarios in the financial business field, which can more efficiently and quickly set business goals for business personnel.

[0111] Further references Figure 8 , as a response to the above Figure 2 In order to realize the method shown in FIG, the present application provides an embodiment of a data prediction device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0112] like Figure 8 As shown, the data prediction device 800 of this embodiment includes: a data collection module 801 for analysis, a principal component analysis module 802, a data matrix construction module 803, a historical feature data trend line acquisition module 804, a missing value identification module 805, a missing value filling module 806, and a continuous task target estimation module 807.

[0113] The data collection module 801 is used to collect the data to be analyzed in real time within the current preset time period;

[0114] The principal component analysis module 802 is configured to perform principal component analysis on the data to be analyzed using a principal component analysis method, and filter out characteristic data of a target quantity dimension in combination with a preset component influence threshold;

[0115] A data matrix construction module 803 is used to construct a data matrix according to the characteristic data of the target number dimension;

[0116] A historical characteristic data trend line obtaining module 804 is configured to extract characteristic data from the reference data within the historical reference time period according to the target quantity dimension, and obtain historical characteristic data trend lines corresponding to the target quantity dimension respectively;

[0117] A missing value identification module 805 is used to identify whether there are missing values ​​in the data matrix;

[0118] A missing value filling module 806 is configured to, if there are missing values ​​in the data matrix, select filling values ​​from the corresponding historical feature data trend lines using a trend line change similarity algorithm based on the data dimension corresponding to the missing values, and fill the missing values ​​into the data matrix to obtain a target data matrix;

[0119] The continuous task target estimation module 807 is used to construct a continuous estimation model using the target data matrix and set dynamically changing task targets according to the continuous output results of the continuous estimation model.

[0120] The present application collects the data to be analyzed within the current preset time period in real time; uses the principal component analysis method to perform principal component analysis on the data to be analyzed, and combines the preset component influence threshold to screen out the characteristic data of the target quantity dimension; constructs a data matrix based on the characteristic data of the target quantity dimension; extracts characteristic data of the reference data within the historical reference time period according to the target quantity dimension, and obtains the historical characteristic data trend line corresponding to the target quantity dimension; identifies whether there are missing values ​​in the data matrix; if there are missing values ​​in the data matrix, fills them to obtain the target data matrix; uses the target data matrix to construct a continuous estimation model, and sets dynamically changing task goals according to the continuous output results of the continuous estimation model. It realizes that a small amount of current analysis data is first analyzed by the principal component analysis method to determine the important characteristic dimensions, and then the historical data is extracted on the important characteristic dimensions to generate trend lines, and the missing values ​​of the data matrix are supplemented in combination with the trend lines, and the continuous estimation model is constructed with the data matrix without missing values. The business estimation model is constructed only in combination with the important dimension features, avoiding the processing of non-important dimension data, and more efficiently assisting specific businesses in data analysis and business goal setting. The data prediction method is specifically applied to financial business data analysis scenarios, such as insurance renewal rate prediction scenarios and stock market data fluctuation prediction scenarios in the financial business field, which can more efficiently and quickly set business goals for business personnel.

[0121] In this embodiment, the data prediction device 800 also includes a data standardization processing module, which is used to use the Z-score standardization algorithm to standardize the feature data of the target quantity dimension according to the dimension category, and obtain the feature data corresponding to the target quantity dimension after eliminating the dimensional effect; and is also used to use the Z-score standardization algorithm to standardize the feature data corresponding to the reference data in the historical reference time period on all major dimensional information according to the dimension category, and obtain the feature data corresponding to the reference data in the historical reference time period on all major dimensional information after eliminating the dimensional effect.

[0122] In this embodiment, the data matrix construction module 803 includes a feature data updating unit, a feature vector generating unit and a data matrix obtaining unit.

[0123] A feature data updating unit, configured to update the feature data corresponding to the target quantity dimensions after eliminating the dimension effect to feature data for constructing a data matrix;

[0124] A feature vector generating unit is used to construct feature data of a data matrix for each dimension and generate a feature vector according to a preset time unit;

[0125] The data matrix obtaining unit is used to fill the feature vectors into the corresponding data matrix in chronological order to obtain the data matrix corresponding to the feature data of each dimension.

[0126] In this embodiment, the historical feature data trend line acquisition module 804 includes a main dimension information determination unit, a reference data feature extraction unit, a reference data annotation processing unit, and a historical feature data trend line construction unit.

[0127] a main dimension information determining unit, configured to determine, through the data matrix, the target number of data dimensions as the main dimension information of the data to be analyzed;

[0128] A reference data feature extraction unit is configured to extract feature data from the reference data within the historical reference time period using the main dimension information of the data to be analyzed as an extraction dimension, and obtain feature data corresponding to all main dimension information of the reference data within the historical reference time period;

[0129] A reference data annotation processing unit is configured to use a Z-score standardization algorithm to perform standardization processing on the feature data corresponding to all main dimensional information of the reference data within the historical reference time period according to the dimensional categories, thereby obtaining the feature data corresponding to all main dimensional information of the reference data within the historical reference time period after eliminating the dimensional influence;

[0130] The historical characteristic data trend line construction unit is used to construct the historical characteristic data trend line according to the characteristic data corresponding to the reference data in the historical reference time period in all main dimensional information after eliminating the dimension effect.

[0131] In this embodiment, the data prediction device 800 further includes a graphical processing input module, a timing control parameter setting module, and a historical characteristic data trend line generation module.

[0132] A graphical processing input module is used to input the characteristic data corresponding to the reference data within the historical reference time period on all main dimensional information after eliminating the dimensional influence as graphical input parameters into a preset data graphical processing component;

[0133] A timing control parameter setting module, configured to use the preset time unit as a timing control parameter of the preset data graphics processing component;

[0134] The specific generation module of the historical characteristic data trend line is used to generate the historical characteristic data trend line corresponding to each main data dimension according to the timing control parameter and the graphical input parameter.

[0135] In this embodiment, the missing value identification module 805 includes a traversal unit, a judgment unit, a first recognition result unit and a second recognition result unit.

[0136] A traversal unit, configured to identify each element in the data matrix in a traversal manner;

[0137] A judgment unit, configured to judge whether the element value corresponding to the current row and column intersection is a NULL value or an empty value based on the recognition result;

[0138] A first recognition result unit is used to determine that if the element values ​​corresponding to the intersections of all rows and columns are neither NULL values ​​nor empty values, then there are no missing values ​​in the data matrix;

[0139] The second recognition result unit is used to determine that there are missing values ​​in the data matrix if the element value corresponding to the intersection of rows and columns is a NULL value or an empty value.

[0140] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0141] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0142] To solve the above technical problems, the present application also provides a computer device. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.

[0143] The computer device 9 includes a memory 9a, a processor 9b, and a network interface 9c that are interconnected via a system bus. Figure 9 Only a computer device 9 having components such as a memory 9a, a processor 9b, and a network interface 9c is shown. However, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead. It should be understood by those skilled in the art that a computer device herein is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, and the like.

[0144] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0145] The memory 9a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 9a can be an internal storage unit of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 9a can also be an external storage device of the computer device 9, such as a plug-in hard disk equipped on the computer device 9, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory 9a can also include both the internal storage unit of the computer device 9 and its external storage device. In this embodiment, the memory 9a is generally used to store the operating system and various application software installed on the computer device 9, such as computer-readable instructions of a data prediction method. In addition, the memory 9a can also be used to temporarily store various types of data that have been output or are to be output.

[0146] The processor 9b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 9b is generally used to control the overall operation of the computer device 9. In this embodiment, the processor 9b is used to execute computer-readable instructions stored in the memory 9a or process data, such as computer-readable instructions for executing the data prediction method.

[0147] The network interface 9c may include a wireless network interface or a wired network interface. The network interface 9c is generally used to establish a communication connection between the computer device 9 and other electronic devices.

[0148] The computer device proposed in this embodiment belongs to the field of big data processing technology and is used in business or market data forecasting and analysis scenarios when it is applied to financial or medical business services. This application collects the data to be analyzed in real time within the current preset time period; uses the principal component analysis method to perform principal component analysis on the data to be analyzed, and combines the preset component influence threshold to screen out the characteristic data of the target quantity dimension; constructs a data matrix based on the characteristic data of the target quantity dimension; extracts characteristic data from the reference data within the historical reference time period based on the target quantity dimension, and obtains the historical characteristic data trend lines corresponding to the target quantity dimension; identifies whether there are missing values ​​in the data matrix; if there are missing values ​​in the data matrix, fills them in to obtain the target data matrix; uses the target data matrix to construct a continuous estimation model, and sets dynamically changing task goals based on the continuous output results of the continuous estimation model. The method first uses principal component analysis to analyze a small amount of current analysis data to determine important feature dimensions. It then extracts features from historical data along these important feature dimensions to generate trend lines. The trend lines are then used to supplement missing values ​​in the data matrix. A continuous forecasting model is constructed using the data matrix without missing values. This business forecasting model is constructed based solely on important dimension features, avoiding the processing of non-important dimension data and more efficiently assisting specific businesses in data analysis and business goal setting. The data forecasting method is specifically applied to financial business data analysis scenarios, such as insurance renewal rate forecasting and stock market data fluctuation forecasting, enabling more efficient and rapid business goal setting for business personnel.

[0149] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by a processor to enable the processor to perform the steps of a data prediction method as described above.

[0150] The computer-readable storage medium proposed in this embodiment belongs to the field of big data processing technology and is applied to business or market data forecasting and analysis scenarios when used in financial or medical business services. This application collects the data to be analyzed within the current preset time period in real time; uses the principal component analysis method to perform principal component analysis on the data to be analyzed, and combines the preset component influence threshold to screen out the characteristic data of the target quantity dimension; constructs a data matrix based on the characteristic data of the target quantity dimension; extracts characteristic data from the reference data within the historical reference time period based on the target quantity dimension, and obtains the historical characteristic data trend lines corresponding to the target quantity dimension; identifies whether there are missing values ​​in the data matrix; if there are missing values ​​in the data matrix, fills them in to obtain the target data matrix; uses the target data matrix to construct a continuous estimation model, and sets dynamically changing task goals based on the continuous output results of the continuous estimation model. The method first uses principal component analysis to analyze a small amount of current analysis data to determine important feature dimensions. It then extracts features from historical data along these important feature dimensions to generate trend lines. The trend lines are then used to supplement missing values ​​in the data matrix. A continuous forecasting model is constructed using the data matrix without missing values. This business forecasting model is constructed based solely on important dimension features, avoiding the processing of non-important dimension data and more efficiently assisting specific businesses in data analysis and business goal setting. The data forecasting method is specifically applied to financial business data analysis scenarios, such as insurance renewal rate forecasting and stock market data fluctuation forecasting, enabling more efficient and rapid business goal setting for business personnel.

[0151] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0152] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is also within the scope of patent protection of this application. The non-company software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.

Claims

1. A data prediction method, characterized in that: The steps include: Real-time collection of data to be analyzed within the current preset time period; The principal component analysis method is used to perform principal component analysis on the data to be analyzed, and the characteristic data of the target quantity dimension is screened out in combination with a preset component influence threshold; Constructing a data matrix based on the characteristic data of the target quantity dimension; Extracting feature data from the reference data within the historical reference time period according to the target quantity dimension to obtain historical feature data trend lines corresponding to the target quantity dimension respectively; Identifying whether there are missing values ​​in the data matrix; If there are missing values ​​in the data matrix, then according to the data dimension corresponding to the missing value, a trend line change similarity algorithm is used to filter out filling values ​​from the corresponding historical feature data trend line, and fill them into the data matrix to obtain the target data matrix; A continuous prediction model is constructed using the target data matrix, and dynamically changing task objectives are set according to continuous output results of the continuous prediction model.

2. The data prediction method according to claim 1, characterized in that Before executing the step of constructing a data matrix according to the feature data of the target number dimension, the method further includes: The Z-score standardization algorithm is used to standardize the feature data of the target quantity dimensions according to the dimension categories, so as to obtain the feature data corresponding to the target quantity dimensions after eliminating the dimension effect.

3. The data prediction method according to claim 2, characterized in that: The step of constructing a data matrix based on the feature data of the target number dimension specifically includes: Updating the characteristic data corresponding to the target quantity dimensions after eliminating the dimension effect to the characteristic data for constructing the data matrix; The feature data of each dimension is constructed by the data matrix, and the feature vector is generated according to the preset time unit; The feature vectors are filled into the corresponding data matrix in chronological order to obtain the data matrix corresponding to the feature data of each dimension.

4. The data prediction method according to claim 1, wherein: The step of extracting feature data from the reference data within the historical reference time period according to the target quantity dimension to obtain the historical feature data trend lines corresponding to the target quantity dimension specifically includes: Determining, through the data matrix, the target number of data dimensions as the main dimension information of the data to be analyzed; Extracting feature data from the reference data within the historical reference time period using the main dimension information of the data to be analyzed as an extraction dimension to obtain feature data corresponding to all main dimension information of the reference data within the historical reference time period; The Z-score normalization algorithm is used to normalize the feature data corresponding to all major dimensional information of the reference data within the historical reference time period according to the dimensional categories, thereby obtaining the feature data corresponding to all major dimensional information of the reference data within the historical reference time period after eliminating the dimension effect; The historical characteristic data trend line is constructed based on the characteristic data after eliminating the dimension effect corresponding to the reference data in the historical reference time period on all main dimensional information.

5. The data prediction method according to claim 1, characterized in that: The step of identifying whether there are missing values ​​in the data matrix specifically includes: Using a traversal method, each element in the data matrix is ​​identified; According to the recognition result, determine whether the element value corresponding to the current row and column intersection is NULL or empty; If the element values ​​corresponding to all row and column intersections are neither NULL nor empty, then there are no missing values ​​in the data matrix; If the element value corresponding to the intersection of a row and a column is a NULL value or an empty value, there is a missing value in the data matrix.

6. The data prediction method according to claim 1, characterized in that: The step of selecting fill values ​​from the corresponding historical feature data trend lines using a trend line change similarity algorithm based on the data dimensions corresponding to the missing values, and filling the fill values ​​into the data matrix to obtain the target data matrix specifically includes: Determine the data dimension information and time series information corresponding to the missing value based on the row and column position information of the missing value in the data matrix; Constructing a data trend line with breakpoints corresponding to the data to be analyzed on the data dimension information according to the data dimension information and time series information corresponding to the missing value; Based on the data dimension information, filter out the historical characteristic data trend line corresponding to the data dimension information from the historical characteristic data trend lines corresponding to all main data dimensions; Using a trend line change similarity algorithm, the data trend line with breakpoints is regressed and compared with the historical feature data trend line, and the most similar trend line segment corresponding to the data trend line with breakpoints in the historical feature data trend line is screened out; Supplementing feature data of all breakpoints in the data trend line having breakpoints according to feature data corresponding to each point in the most similar trend line segment, and generating corresponding filling values ​​according to the supplemented feature data; The corresponding filling values ​​are filled into the data according to the row and column position information corresponding to all breakpoints to obtain the target data matrix.

7. The data prediction method according to claim 1, characterized in that: The step of using the target data matrix to construct a continuous estimation model and setting a dynamically changing task goal according to the continuous output results of the continuous estimation model specifically includes: The target data matrix is ​​used as a parameter matrix for constructing the continuous prediction model, and the continuous prediction model is constructed using a regularized Ridge algorithm. When constructing the continuous prediction model using the regularized Ridge algorithm, a maximum regularization parameter is screened using a cross-validation method, and model construction control settings are performed to reduce model complexity. After executing the steps of constructing a continuous estimation model using the target data matrix and setting a dynamically changing task goal based on a continuous output result of the continuous estimation model, the method further includes: Determine whether the actual task result at the current time point exceeds the estimated task goal; If the actual task result at the current time point exceeds the estimated task goal, then continue to monitor the completion of the task goal; If the actual task result corresponding to the current time point does not exceed the estimated task target, then determine whether the actual task result has reached a preset warning threshold compared to the estimated task target; If the actual task result does not reach the preset warning threshold compared to the estimated task target, an email reminder is sent to the target business processing terminal; If the actual task result reaches the preset warning threshold compared with the estimated task goal, the task goal completion status monitoring will continue.

8. A data prediction device, characterized in that: include: The data collection module for analysis is used to collect the data to be analyzed within the current preset time period in real time; A principal component analysis module is used to perform principal component analysis on the data to be analyzed using a principal component analysis method, and to filter out characteristic data of a target quantity dimension in combination with a preset component influence threshold; A data matrix construction module, configured to construct a data matrix based on the characteristic data of the target quantity dimension; A historical characteristic data trend line acquisition module is used to extract characteristic data from the reference data within the historical reference time period according to the target quantity dimension, and obtain the historical characteristic data trend lines corresponding to the target quantity dimension respectively; A missing value identification module is used to identify whether there are missing values ​​in the data matrix; A missing value filling module is used to, if there are missing values ​​in the data matrix, select filling values ​​from the corresponding historical feature data trend lines using a trend line change similarity algorithm based on the data dimensions corresponding to the missing values, and fill them into the data matrix to obtain a target data matrix; The continuous task target estimation module is used to construct a continuous estimation model using the target data matrix and set dynamically changing task targets according to the continuous output results of the continuous estimation model.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the data prediction method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the data prediction method according to any one of claims 1 to 7.