A life cycle-based investment management method and system

By analyzing the correlation and constraint factors of investment management data and adopting the isolation forest algorithm with adaptive segmentation points, the problems of low accuracy and efficiency in abnormal data detection in traditional methods are solved, achieving higher investment management data accuracy and decision reliability.

CN120163650BActive Publication Date: 2025-09-23SHANGRAO HIGH-SPEED RAILWAY ECONOMIC PILOT ZONE INVESTMENT & CONSTRUCTION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510236891.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-09-23
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The traditional isolation forest algorithm has problems with low accuracy and efficiency in detecting abnormal data in investment project survey data and previous project information data, which affects the reliability of investment decisions.

Method used

By analyzing the correlation and constraint factors between investment management data, an isolation forest algorithm with adaptive segmentation points and segmentation features is used to construct an isolation tree to eliminate abnormal data and improve detection accuracy and efficiency.

Benefits of technology

It improves the accuracy and robustness of abnormal data detection, enhances the ability to capture complex dynamic patterns, and improves the accuracy of investment management data and the reliability of investment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163650B_ABST
    Figure CN120163650B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of investment anomaly data detection, and specifically to a lifecycle-based investment management method and system, the method comprising: collecting various types of investment management data within a preset time period; determining the degree of correlation between any type of investment management data and the remaining types of investment management data; determining the constraint factor between the investment management data of any type and the remaining types of investment management data; obtaining the constraint correlation between the investment management data of any type and the remaining types of investment management data; obtaining various types of investment management data that have a strong correlation with the investment management data of any type; evaluating the preference of each data point in the investment management data of any type as a segmentation point; obtaining each abnormal data in the investment management data of any type based on the preference and eliminating the abnormal data, and performing investment management on the various types of investment management data, thereby improving the security and reliability of investment decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of investment abnormal data detection, and in particular to a life cycle-based investment management method and system. Background Art

[0002] Lifecycle investment management (LIM) is a systematic approach that uses time as a framework to dynamically adjust investment strategies, risk appetite, and resource allocation based on the characteristics of an entity's lifecycle. Its core is to view investment as a dynamic process rather than a static decision, emphasizing the value of time, risk-return matching, and full-lifecycle cost-benefit optimization. This approach not only improves the financial health of individuals and institutions but also promotes a shift in resource utilization from short-term consumption to long-term value creation.

[0003] When making investment project decisions at various stages of the life cycle, it is often necessary to integrate project research data and previous project information data in order to comprehensively evaluate the investment project and improve the scientific nature of project investment through data analysis. However, in the process of collecting project research data and previous project information data, abnormal data are usually inevitable. Direct analysis of the collected data will greatly affect the reliability of project decisions.

[0004] When using the isolation forest algorithm to detect anomaly data in project survey data and previous project information data, the traditional algorithm greatly affects the efficiency and accuracy of isolation tree construction by randomly selecting segmentation features and segmentation points, which in turn affects the accuracy and efficiency of anomaly data detection, leading to investment decision-making bias. Summary of the Invention

[0005] In order to solve the above technical problems, the purpose of this application is to provide a lifecycle-based investment management method and system. The technical solutions adopted are as follows:

[0006] In a first aspect, an embodiment of the present application provides a lifecycle-based investment management method, the method comprising the following steps:

[0007] Collect various investment management data within a preset time period;

[0008] Analyze the differences between adjacent data points in any category of investment management data, as well as the data distribution of the remaining categories of investment management data within the collection time interval corresponding to the adjacent data points, to determine the degree of correlation between the investment management data in any category and the remaining categories of investment management data;

[0009] Taking the data of any one type of investment management data and the data of any other type of investment management data collected at the same time as a feature point pair; determining a constraint factor between the any one type of investment management data and any other type of investment management data based on the frequency of occurrence of the feature point pairs and the time difference between the feature point pairs;

[0010] Combining the correlation degree with the constraint factor, obtaining the constraint correlation degree between the any one type of investment management data and the remaining types of investment management data; classifying the remaining types of investment management data using the correlation between the constraint correlation degrees; obtaining the types of investment management data that have a strong correlation with the any one type of investment management data;

[0011] Obtaining relevant data points of each data point in any one type of investment management data from each type of investment management data having strong correlation, using each data point and the relevant data point as a segmentation point, segmenting the investment management data to which the segmentation point belongs, analyzing the difference in the number of data between the two categories of each type of investment management data after segmentation, and the difference in the average level, and evaluating the preference of each data point in any one type of investment management data as a segmentation point in combination with the constraint correlation degree;

[0012] An isolation forest algorithm is used to construct an isolation tree for any type of investment management data, and a split point in the isolation tree construction process is selected based on the preference degree. Abnormal data in any type of investment management data is obtained and eliminated, and investment management is performed based on each type of investment management data after the abnormal data is eliminated.

[0013] In one embodiment, determining the degree of association includes:

[0014] Calculate the ratio of the absolute value of the difference between adjacent data points in any one type of investment management data to the maximum value among the adjacent data points, and record it as a first ratio; count the number of data points of the remaining types of investment management data that are not within the collection time interval of any one type of investment management data, and calculate the correlation degree by combining the first ratio and the number of data points, as well as the data distribution of the remaining types of investment management data within the collection time interval corresponding to the adjacent data points in any one type of investment management data.

[0015] In one embodiment, the degree of association is calculated as follows:

[0016] Where, GL QW is the correlation between Q-type investment management data and W-type investment management data, exp[] is an exponential function with a natural constant as the base, N Q is the number of data points of Q-type investment management data, For the first ratio of the nth data point in the Q-type investment management data, obtain all data points of the W-type investment management data in the collection time interval corresponding to the nth data point and the (n+1)th data point in the Q-type investment management data, calculate the ratio of the extreme value to the maximum value of all the data points of the W-type investment management data, record it as the second ratio, and calculate the degree of dispersion of all the data points of the W-type investment management data. is the product of the second ratio and the discrete degree, β is a value preset to be greater than 0, and n W is the number of data points in the W-type investment management data that are not within the collection time interval corresponding to the Q-type investment management data.

[0017] In one embodiment, determining the constraint factor includes:

[0018] Among all the feature point pairs of any one type of investment management data and any remaining type of investment management data, calculate the ratio of the number of occurrences of each feature point pair to the number of all feature point pairs, record it as the third ratio, calculate the cumulative sum of the time intervals between all occurrences of each feature point pair, and the constraint factor is the fusion result of the third ratio and the cumulative sum.

[0019] In one embodiment, the constraint correlation degree is positively correlated with the correlation degree and negatively correlated with the constraint factor.

[0020] In one embodiment, the acquiring of various types of investment management data having a strong correlation with any type of investment management data includes:

[0021] The constraint correlation between any one type of investment management data and all remaining types of investment management data is divided into two categories using a clustering algorithm, and the types of investment management data in the category with the largest mean constraint correlation are regarded as the types of investment management data with strong correlation with any one type of investment management data.

[0022] In one embodiment, the relevant data point is a data point in the various types of investment management data with strong correlation that is closest in collection time to each data point in any type of investment management data.

[0023] In one embodiment, the preference is calculated as follows:

[0024] Where Y Qq is the preference of data point q as the split point in Q-type investment management data; n q1 is the number of data points in one category after the Q-category investment management data is segmented with data point q as the segmentation point, n q2 is the number of the second-class data points, is the mean of one type of data points, is the mean of the two types of data points, R u is the constraint correlation degree between the Q-type investment management data and the u-th type of investment management data with strong correlation; U is the number of types of investment management data with strong correlation with the Q-type investment management data, is the number of data points in one category after the u-th category investment management data is segmented by the related data points of data point q, is the number of the second-class data points, is the mean of one type of data points, is the mean of the two types of data points.

[0025] In one embodiment, when the Q-type investment management data is segmented using data point q as a segmentation point, data points in the Q-type investment management data that are greater than or equal to data point q are treated as first-type data points, and data points that are less than data point q are treated as second-type data points; when the u-th type of investment management data is segmented using data points related to data point q, the same segmentation method as that used in the Q-type investment management data using data point q as a segmentation point is used for segmentation.

[0026] The method of selecting a segmentation point for the isolated tree construction process based on the preference is: selecting a segmentation point with the largest preference as a segmentation point for the isolated tree construction process.

[0027] In a second aspect, an embodiment of the present application also provides a lifecycle-based investment management system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of any one of the above-mentioned methods when executing the computer program.

[0028] This application has at least the following beneficial effects:

[0029] This application collects various types of investment management data within a preset time period; analyzes the differences between adjacent data points in any type of investment management data, and the data distribution of the remaining types of investment management data within the collection time interval corresponding to the adjacent data points, and determines the correlation between the aforementioned any type of investment management data and the remaining types of investment management data; the calculation of the correlation can improve the accuracy and robustness of anomaly detection when performing anomaly detection on various types of investment management data, avoid the limitations of single-dimensional detection, and enhance the ability to capture complex dynamic patterns; the aforementioned any type of investment management data and the data of the same collection time in the remaining any type of investment management data are used as special feature point pairs; based on the frequency of occurrence of feature point pairs and the time difference between feature point pairs, determine the constraint factor between any one type of investment management data and any other type of investment management data; combine the correlation and the constraint factor to obtain the constraint correlation between any one type of investment management data and the remaining types of investment management data; the determination of the constraint factor avoids the deviation in the correlation calculation process, introduces the frequency and time difference constraints of feature point pairs, eliminates false information that is statistically relevant but has no causal or business logic association, and improves the accuracy of correlation analysis; uses the correlation between the constraint correlations to separate the remaining types of investment management data Classification is performed; various types of investment management data with strong correlation with any of the types of investment management data are obtained; relevant data points of each data point in the various types of investment management data with strong correlation are respectively obtained, each data point and the relevant data point are used as a segmentation point, the investment management data to which the segmentation point belongs are segmented, the difference in the number of data between the two categories of each type of investment management data after segmentation and the difference in the average level are analyzed, and the preference of each data point in the any of the types of investment management data as a segmentation point is evaluated in combination with the constraint correlation degree; the preference reflects the suitability of each data point in each type of investment management data as a segmentation point, avoids the problem of low accuracy and efficiency of abnormal data detection caused by random selection of segmentation points by the traditional isolation forest algorithm, and improves the accuracy of segmentation point determination; an isolation tree is constructed for the any of the types of investment management data using the isolation forest algorithm, segmentation points in the isolation tree construction process are selected based on the preference, each abnormal data in the any of the types of investment management data is obtained and eliminated, and investment management is performed based on the various types of investment management data after the abnormal data is eliminated, thereby improving the accuracy and efficiency of abnormal data detection in the various types of investment management data, thereby improving the accuracy of investment management data and enhancing the reliability and security of investment decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0031] Figure 1 A flowchart of a life cycle-based investment management method provided in accordance with one embodiment of the present application;

[0032] Figure 2 This is the flow chart of abnormal data detection. DETAILED DESCRIPTION

[0033] To further illustrate the technical means and effectiveness of this application to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of a lifecycle-based investment management method and system proposed in this application. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0034] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0035] The following describes in detail a specific scheme of a life cycle-based investment management method and system provided by this application with reference to the accompanying drawings.

[0036] See also Figure 1 , which shows a flowchart of a life cycle-based investment management method provided by one embodiment of the present application, the method comprising the following steps:

[0037] S1, collects various types of investment management data within a preset time period.

[0038] This embodiment acquires various investment research data for the investment project through market research and field visits at the initial stage of the investment project. This investment research data in this embodiment includes market size data, growth rate data, investment cost data, and revenue forecast data. Next, actual execution data for similar historical projects of the investment project is obtained based on previous project data. This includes various historical data, including actual project revenue data, cost overrun data, construction delay data, risk loss data, and material resource data. These various investment research data and historical data are collectively referred to as various types of investment management data.

[0039] It should be noted that the collection period of various types of investment management data in this embodiment is one year, and the analysis is conducted taking one stage of the life cycle as an example. The implementers can set the types of investment research data and historical data according to actual conditions, and this embodiment does not impose any restrictions on this.

[0040] Since different types of investment management data have different characteristics and formats, this embodiment digitizes all types of investment management data, that is, each type of investment management data is a numerical sequence. Moreover, since the types of investment management data are different, the collection frequency of different types of investment management data may also be different, that is, the lengths of the numerical sequences of various types of investment management data are not exactly the same.

[0041] S2, analyzing the differences between adjacent data points in any type of investment management data, and the data distribution of the remaining types of investment management data within the collection time interval corresponding to the adjacent data points, to determine the correlation between the said any type of investment management data and the remaining types of investment management data.

[0042] When using the Isolation Forest algorithm to detect anomalies in various investment management data, the traditional Isolation Forest algorithm randomly selects split points and features when constructing isolation trees. The choice of split points and features directly affects the accuracy and efficiency of the isolation tree in detecting anomalies. Therefore, this embodiment improves the accuracy and efficiency of isolation tree construction by adaptively selecting split points and features, thereby improving the accuracy and efficiency of isolation forest detection of anomalies.

[0043] When performing anomaly detection on various types of investment management data, this example sets the number of isolation trees to 200. Each time, 256 data points are sampled with replacement from each type of investment management data. For any type of investment management data, using Q-type investment management data as an example, the associated data category of the Q-type investment management data is first obtained. This associated data category is then used as a feature selected for each segmentation to assist in the construction of an isolation tree. This maximizes the accuracy and efficiency of the segmentation performed on each node of the isolation tree.

[0044] For all types of investment management data, obtain any type of investment management data except Q type investment management data. Take W type investment management data as an example to analyze the correlation between Q type investment management data and W type investment management data. The specific analysis process is as follows:

[0045] Calculate the absolute value of the difference between each data point in the Q-type investment management data and its adjacent subsequent data point, and calculate the ratio of the absolute value of the difference to the maximum value of the two data points corresponding to the absolute value of the difference, which is recorded as a first ratio; then, count the number of data points in the W-type investment management data that are not located in the collection time zone corresponding to the Q-type investment management data, and combine the first ratio with the number of data points in the W-type investment management data that are not located in the collection time zone corresponding to the Q-type investment management data, and the data distribution of the W-type investment management data in the collection time interval corresponding to adjacent data points in the Q-type investment management data to calculate the correlation between the Q-type investment management data and the W-type investment management data. The specific calculation method is:

[0046] Where, GL QW is the correlation between Q-type investment management data and W-type investment management data, exp[] is an exponential function with a natural constant as the base, N Q is the number of data points of Q-type investment management data, For the first ratio of the nth data point in the Q-type investment management data, obtain all data points of the W-type investment management data in the collection time interval corresponding to the nth data point and the (n+1)th data point in the Q-type investment management data, calculate the ratio of the extreme value to the maximum value of all the data points of the W-type investment management data, record it as the second ratio, and calculate the degree of dispersion of all the data points of the W-type investment management data. is the product of the second ratio and the discrete degree, β is a value preset to be greater than 0, and n W is the number of data points in the W-type investment management data that are not within the collection time interval corresponding to the Q-type investment management data.

[0047] It should be noted that, in this embodiment, β=0.01 is used to avoid the denominator being 0. In this embodiment, the degree of dispersion is calculated using the standard deviation. The implementer may choose other feasible existing methods for calculating the degree of dispersion, such as variance, coefficient of variation, etc.

[0048] It should be understood that, when the number of data points in the W-type investment management data that are not within the collection time interval corresponding to the Q-type investment management data is greater, it means that the intersection of the collection time intervals corresponding to the W-type investment management data and the Q-type investment management data is smaller, that is, more data points in the W-type investment management data are collected before or after the collection time interval of the Q-type investment management data, resulting in a smaller correlation between the W-type investment management data and the Q-type investment management data, that is, the smaller the correlation between the Q-type investment management data and the W-type investment management data. Reflects the numerical distribution of the nth data point and its adjacent data points in the Q-type investment management data. It reflects the numerical distribution of the nth data point in the Q-type investment management data to the corresponding data point in the W-type investment management data. If The greater the difference between 1 and Q, the greater the difference in the numerical distribution of data points at the same time in the Q-type investment management data and the W-type investment management data, and the smaller the correlation between the Q-type investment management data and the W-type investment management data, that is, the smaller the correlation between the Q-type investment management data and the W-type investment management data, and the less suitable the W-type investment management data is as the segmentation feature for the Q-type investment management data when using the isolation forest algorithm for abnormal data detection.

[0049] By using the same calculation method for the correlation between the Q-type investment management data and the W-type investment management data, the correlation between any two types of investment management data in all types of investment management data can be obtained.

[0050] S3, taking the data of any one type of investment management data and the data of any other type of investment management data at the same collection time as a feature point pair; based on the frequency of occurrence of the feature point pairs and the time difference between the feature point pairs, determining the constraint factor between the any one type of investment management data and any other type of investment management data.

[0051] Taking Q-type investment management data and W-type investment management data as an example, when analyzing the correlation between Q-type investment management data and W-type investment management data, if both Q-type investment management data and W-type investment management data are data with relatively gentle fluctuations, but the correlation between Q-type investment management data and W-type investment management data in time is poor, it will cause a phenomenon of being too high when calculating the correlation between Q-type investment management data and W-type investment management data. Therefore, in order to further accurately reflect the correlation between Q-type investment management data and W-type investment management data, this embodiment records the data points at the same time in Q-type investment management data and W-type investment management data as feature point pairs, and performs statistical analysis on the co-occurrence frequency of the obtained feature point pairs. The co-occurrence frequency is the ratio of the number of occurrences of each feature point pair in Q-type investment management data and W-type investment management data to the number of all feature point pairs, which is recorded as the third ratio, that is, the ratio of the number of identical feature point pairs to the number of all feature point pairs.

[0052] This embodiment constructs a constraint factor through the distribution of the co-occurrence frequency of each feature point pair. When the Q-type investment management data and the W-type investment management data both fluctuate relatively slowly or even remain stable, resulting in a high correlation between the calculated Q-type investment management data and the W-type investment management data, the number of feature point pairs obtained by statistics is relatively small, and the co-occurrence frequency of each feature point pair is often large. For the two types of investment management data with a higher actual correlation degree, their data values ​​are changing, and the number of feature point pairs obtained is relatively large, and the co-occurrence frequency of each feature point pair is often small.

[0053] Based on this, this embodiment constructs a constraint factor between the Q-type investment management data and the W-type investment management data. The specific calculation method is: Where y QW is the constraint factor between Q-type investment management data and W-type investment management data, M is the number of feature point pairs in Q-type investment management data and W-type investment management data, f m is the co-occurrence frequency of the mth feature point pair in the Q-type investment management data and the W-type investment management data, LS m It is the cumulative sum of the time intervals between all occurrences of the mth feature point pair in the Q-type investment management data and the W-type investment management data.

[0054] It should be noted that a pair of feature points with the same value at different times is considered a feature point pair. For example, the set of feature point pairs between Q-type investment management data and W-type investment management data is [(1,2)(2,3)(4,6)(2,3)(5,7)(1,2)(2,3)], then the number of feature point pairs is 4, and the co-occurrence frequency of the feature point pair (1,2) is 2 / 4=0.5. If the moments of each feature point pair in the feature point pair set [(1,2)(2,3)(4,6)(2,3)(5,7)(1,2)(2,3)] are 1, 2, 3, 4, 5, 6, 7 respectively, then the cumulative sum of the time intervals between all occurrences of the feature point pair (2,3) is (4-2)+(7-4)=5.

[0055] y QW This is the fusion result. Fusion means combining multiple variables. Specifically, it can be calculated by addition, multiplication, addition and multiplication mixture, and averaging.

[0056] It should be understood that the greater the co-occurrence frequency of each pair of feature points, the more likely it is that the calculated correlation between the Q-type investment management data and the W-type investment management data is to be biased high. Therefore, the larger the constraint factor, the larger the cumulative sum of the time intervals between all occurrences of each pair of feature points, the more discrete the co-occurrence frequency distribution of the feature point pairs is. It is also more likely that both the Q-type investment management data and the W-type investment management data fluctuate relatively slowly or remain stable, resulting in a higher correlation between the Q-type investment management data and the W-type investment management data, and the larger the constraint factor.

[0057] S4. Combining the correlation degree with the constraint factor, obtain the constraint correlation degree between any one type of investment management data and the remaining types of investment management data; classify the remaining types of investment management data using the correlation between the constraint correlation degrees; and obtain various types of investment management data that have a strong correlation with any one type of investment management data.

[0058] This embodiment constrains the correlation between the Q-type investment management data and the W-type investment management data through a constraint factor. The specific calculation method is: Where R QW is the constraint correlation between Q-type investment management data and W-type investment management data, GL QW is the correlation between Q-type investment management data and W-type investment management data, y QW is the constraint factor between Q-type investment management data and W-type investment management data, μ is a preset value greater than 0 to avoid the denominator being 0. In this embodiment, μ=0.01. The implementer can set it according to actual conditions, and this embodiment does not impose any restrictions here.

[0059] The same calculation method as that for the constraint correlation between the Q-type investment management data and the W-type investment management data is used to obtain the constraint correlation between the Q-type investment management data and the remaining types of investment management data. In this embodiment, the constraint correlation between the Q-type investment management data and all other types of investment management data is clustered using the k-means clustering algorithm, with k set to 2. The implementer can set it according to the actual situation, and this embodiment does not impose any restrictions. The k-means clustering algorithm is used to divide the constraint correlation between the Q-type investment management data and all other types of investment management data into two categories, and the mean of all constraint correlations in each category is calculated. The various types of investment management data in the category corresponding to the maximum mean value are regarded as the various types of investment management data that have a strong correlation with the Q-type investment management data.

[0060] The k-means clustering algorithm is a well-known technology. The implementer can choose other feasible clustering algorithms at will, and this embodiment does not limit it.

[0061] S5, respectively obtaining relevant data points of each data point in any one type of investment management data from the various types of investment management data with strong correlation, using each data point and the relevant data points as segmentation points, segmenting the investment management data to which the segmentation points belong, analyzing the difference in the number of data between the two categories of each type of investment management data after segmentation, as well as the difference in the average level, and evaluating the preference of each data point in any one type of investment management data as a segmentation point in combination with the constraint correlation degree.

[0062] When using the isolation forest algorithm to detect anomaly data on Q-type investment management data, various types of investment management data that have strong correlation with it are used as segmentation features to assist in the selection of segmentation points for Q-type investment management data, so as to improve the segmentation efficiency and accuracy when constructing each node segmentation of the isolation tree. At the same time, it can also greatly reduce the depth of the isolation tree, thereby improving the efficiency and accuracy of anomaly data detection.

[0063] This embodiment calculates the preference of each data point in the Q-type investment management data as a splitting point. Taking the data point q in the Q-type investment management data as an example, in each type of investment management data that has a strong correlation with the Q-type investment management data, the data point closest to the data point q in the collection time is respectively obtained as the relevant data point of the data point q in the various types of investment management data that have a strong correlation with the Q-type investment management data. For each type of investment management data that has a strong correlation with the Q-type investment management data, the relevant data points of each type of investment management data are used as splitting points to split each type of investment management data into two parts. Specifically, in each type of investment management data, data points that are greater than or equal to the relevant data points are regarded as the first type of data points of each type of investment management data, and data points that are less than the relevant data points are regarded as the second type of data points of each type of investment management data. Similarly, in the Q-type investment management data, the Q-type investment management data is split with the data point q as the splitting point, and the data points in the Q-type investment management data that are greater than or equal to the data point q are regarded as the first type of data points, and the data points that are less than the data point q are regarded as the second type of data points.

[0064] Based on this, the preference of data point q as the split point in the Q-type investment management data is calculated. The specific calculation method is: Where Y Qq is the preference of data point q as the split point in Q-type investment management data; n q1 is the number of data points in one category after the Q-category investment management data is segmented with data point q as the segmentation point, n q2 is the number of the second-class data points, is the mean of one type of data points, is the mean of the two types of data points, R u is the constraint correlation degree between the Q-type investment management data and the u-th type of investment management data with strong correlation; U is the number of types of investment management data with strong correlation with the Q-type investment management data, is the number of data points in one category after the u-th category investment management data is segmented by the related data points of data point q, is the number of the second-class data points, is the mean of one type of data points, is the mean of the two types of data points.

[0065] It should be understood that after the investment type management data is segmented by data point q and data points related to data point q, the greater the difference in data volume and average level between the two categories, the easier it is to find abnormal data in the investment type management data when data point q is used as the segmentation point. Therefore, the greater the preference of data point q as the segmentation point, the more data point q should be selected as the segmentation point.

[0066] S6, using the isolation forest algorithm to construct an isolation tree for any type of investment management data, selecting the split point of the isolation tree construction process based on the preference, obtaining and eliminating each abnormal data in any type of investment management data, and performing investment management based on each type of investment management data after eliminating the abnormal data.

[0067] When using the isolation forest algorithm to detect abnormal data in Q-type investment management data, the first segmentation point selected is the data point corresponding to the maximum preference value in the Q-type investment management data. After segmenting the Q-type investment management data, the preference of each data point in the segmented data sequence as the segmentation point is repeatedly calculated, and the data sequence is continued to be segmented using the recalculated maximum preference value until there is only one data point in the current subset of the isolation tree, or all data points in the subset are the same, or the isolation tree reaches the maximum depth. In this embodiment, the maximum depth is 8, and the implementer can set it according to the actual situation. This embodiment does not impose any restrictions here. At this point, the construction of the isolation tree in the isolation forest algorithm can be completed, and the path length of the data point can be calculated through the isolation tree to obtain the abnormal score of the data point, completing the detection of abnormal data in the Q-type investment management data. Among them, the isolation forest algorithm is an existing well-known technology, and the specific process will not be described in detail. The abnormal data detection flow chart is as follows Figure 2 shown.

[0068] Using the same outlier data detection method as for Q-type investment management data, we identify and remove outliers from various investment management data types. We then apply the C4.5 decision tree construction algorithm to the removed data types, constructing a decision tree. This decision tree is then used to assist project investment decisions, thereby enabling lifecycle-based investment management. The C4.5 decision tree construction algorithm is a well-known technique.

[0069] Based on the same inventive concept as the above method, an embodiment of the present application also provides a lifecycle-based investment management system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned lifecycle-based investment management methods are implemented.

[0070] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0071] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0072] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A life cycle-based investment management method, characterized in that: The method comprises the following steps: Collect various investment management data within a preset time period; Analyze the differences between adjacent data points in any category of investment management data, as well as the data distribution of the remaining categories of investment management data within the collection time interval corresponding to the adjacent data points, to determine the degree of correlation between the investment management data in any category and the remaining categories of investment management data; Taking the data of any one type of investment management data and the data of any other type of investment management data collected at the same time as a feature point pair; determining a constraint factor between the any one type of investment management data and any other type of investment management data based on the frequency of occurrence of the feature point pairs and the time difference between the feature point pairs; Combining the correlation degree with the constraint factor, obtaining the constraint correlation degree between the any one type of investment management data and the remaining types of investment management data; classifying the remaining types of investment management data using the correlation between the constraint correlation degrees; obtaining the types of investment management data that have a strong correlation with the any one type of investment management data; Obtaining relevant data points of each data point in any one type of investment management data from each type of investment management data having strong correlation, using each data point and the relevant data point as a segmentation point, segmenting the investment management data to which the segmentation point belongs, analyzing the difference in the number of data between the two categories of each type of investment management data after segmentation, and the difference in the average level, and evaluating the preference of each data point in any one type of investment management data as a segmentation point in combination with the constraint correlation degree; An isolation forest algorithm is used to construct an isolation tree for any of the investment management data, a split point in the isolation tree construction process is selected based on the preference, each abnormal data in the any of the investment management data is obtained and removed, and investment management is performed based on the various types of investment management data after the abnormal data are removed; Calculating the ratio of the absolute value of the difference between adjacent data points in any one type of investment management data to the maximum value among the adjacent data points, and recording the ratio as a first ratio; The calculation method of the correlation degree is: Where, GL QW is the correlation between Q-type investment management data and W-type investment management data, exp[] is an exponential function with a natural constant as the base, N Q is the number of data points of Q-type investment management data, For the first ratio of the nth data point in the Q-type investment management data, obtain all data points of the W-type investment management data in the collection time interval corresponding to the nth data point and the (n+1)th data point in the Q-type investment management data, calculate the ratio of the extreme value to the maximum value of all the data points of the W-type investment management data, record it as the second ratio, and calculate the degree of dispersion of all the data points of the W-type investment management data. is the product of the second ratio and the discrete degree, β is a value preset to be greater than 0, and n W is the number of data points in the W-type investment management data that are not within the collection time interval corresponding to the Q-type investment management data; The calculation method of the preference degree is: Where Y Qq is the preference of data point q as the split point in Q-type investment management data; n q1 is the number of data points in one category after the Q-category investment management data is segmented with data point q as the segmentation point, n q2 is the number of the second-class data points, is the mean of one type of data points, is the mean of the two types of data points, R u is the constraint correlation degree between the Q-type investment management data and the u-th type of investment management data with strong correlation; U is the number of types of investment management data with strong correlation with the Q-type investment management data, is the number of data points in one category after the u-th category investment management data is segmented by the related data points of data point q, is the number of the second-class data points, is the mean of one type of data points, is the mean of the two types of data points.

2. The life cycle-based investment management method according to claim 1, characterized in that: Determination of the constraint factor includes: Among all the feature point pairs of any one type of investment management data and any remaining type of investment management data, calculate the ratio of the number of occurrences of each feature point pair to the number of all feature point pairs, record it as the third ratio, calculate the cumulative sum of the time intervals between all occurrences of each feature point pair, and the constraint factor is the fusion result of the third ratio and the cumulative sum.

3. The life cycle-based investment management method according to claim 1, characterized in that: The constraint association degree is positively correlated with the association degree and negatively correlated with the constraint factor.

4. The life cycle-based investment management method according to claim 1, characterized in that: The acquiring of various types of investment management data having a strong correlation with any type of investment management data includes: The constraint correlation between any one type of investment management data and all remaining types of investment management data is divided into two categories using a clustering algorithm, and the types of investment management data in the category with the largest mean constraint correlation are regarded as the types of investment management data with strong correlation with any one type of investment management data.

5. The life cycle-based investment management method according to claim 1, characterized in that: The relevant data points are data points in the various types of investment management data with strong correlation that are closest in collection time to the data points in any type of investment management data.

6. The life cycle-based investment management method according to claim 1, characterized in that: When the Q-type investment management data is segmented using the data point q as the segmentation point, the data points in the Q-type investment management data that are greater than or equal to the data point q are regarded as the first-class data points, and the data points that are less than the data point q are regarded as the second-class data points; When segmenting the u-th investment management data using the relevant data points of data point q, the same segmentation method as that used in the Q-th investment management data using data point q as the segmentation point is used for segmentation; The method of selecting a segmentation point for the isolated tree construction process based on the preference is: selecting a segmentation point with the largest preference as a segmentation point for the isolated tree construction process.

7. A lifecycle-based investment management system comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • A screening method and a device for a shale gas development scheme

    CN109102182A

  • Temperature change test box temperature intelligent regulation and control method based on data processing

    CN116661522A