Data processing method, device, computer equipment and storage medium

By obtaining the total amount of data on the data platform and the remaining amount of data in the target data account, the method of calculating the deposit sedimentation rate is simplified, which solves the problem of the difficulty of traditional calculation and achieves more efficient data sedimentation rate calculation.

CN113901112BActive Publication Date: 2025-09-16PING AN BANK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111367131.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2025-09-16
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

Traditional data processing methods are difficult to calculate when calculating deposit sedimentation rates and require high performance from data processing equipment.

Method used

By obtaining the total amount of data on the data platform and the remaining amount of data in the target data account, the data sedimentation rate is calculated, avoiding the processing of complex account transaction details and data preprocessing, and adopting a simplified position data calculation method.

Benefits of technology

It reduces the complexity and difficulty of computer equipment in calculating data sedimentation rate, improves calculation speed, and reduces the demand for equipment performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901112B_ABST
    Figure CN113901112B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a data processing method, apparatus, computer equipment and storage medium, wherein the method includes: obtaining the total amount of data of a data platform in a first time unit; obtaining the first data remaining amount of a target data account on the data platform in the first time unit, and the second data remaining amount of the target data account on the data platform in a second time unit; wherein the second time unit is a time unit before the first time unit; determining the data sedimentation rate of the target data account in the first time unit based on the total amount of data, the first data remaining amount and the second data remaining amount, which can reduce the difficulty of calculating the data sedimentation rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, computer equipment, and storage medium. Background Art

[0002] To better manage data, various data platforms typically develop corresponding data management strategies based on a certain reference indicator. For example, financial institutions typically specify fund management strategies based on the deposit sedimentation rate. The so-called deposit sedimentation rate refers to the ratio of the average deposit balance to the total deposit amount, which can effectively reflect the amount of funds in the financial institution. For example, a bank's deposit sedimentation rate can be used to reflect the amount of funds deposited in the bank, or in other words, the bank's deposit sedimentation rate can reflect the availability of bank funds. Traditional data processing methods for calculating the deposit sedimentation rate typically obtain account consumption details and further calculate the deposit sedimentation rate based on the account consumption details. However, this method is difficult to calculate and has high requirements for the performance of data processing equipment. Therefore, how to reduce the difficulty of calculating the deposit sedimentation rate has become a current research hotspot. Summary of the Invention

[0003] The embodiments of the present application provide a data processing method, apparatus, computer equipment, and storage medium, which can reduce the difficulty of calculating the data sedimentation rate.

[0004] In one aspect, an embodiment of the present application provides a data processing method, comprising:

[0005] Obtain the total amount of data on the data platform in the first time unit;

[0006] Obtaining a first remaining amount of data in the target data account on the data platform at the first time unit, and a second remaining amount of data in the target data account on the data platform at the second time unit; wherein the second time unit is a time unit before the first time unit;

[0007] A data sedimentation rate of the target data account in the first time unit is determined according to the total amount of data, the first remaining amount of data, and the second remaining amount of data.

[0008] In another aspect, an embodiment of the present application provides a data processing device, comprising:

[0009] An acquisition unit, used to acquire the total amount of data from the data platform in the first time unit;

[0010] The acquisition unit is further configured to acquire a first remaining amount of data of the target data account on the data platform at the first time unit, and a second remaining amount of data of the target data account on the data platform at a second time unit; wherein the second time unit is a time unit before the first time unit;

[0011] A determination unit is used to determine the data sedimentation rate of the target data account in the first time unit based on the total amount of data, the first remaining amount of data and the second remaining amount of data.

[0012] In another aspect, an embodiment of the present application provides a computer device, including:

[0013] a processor adapted to execute one or more computer instructions;

[0014] A storage medium storing one or more computer instructions, wherein the one or more computer instructions are suitable for being loaded by the processor and executing the above-mentioned data processing method.

[0015] On the other hand, an embodiment of the present application provides a storage medium, which stores one or more computer instructions, and the one or more computer instructions are suitable for being loaded by a processor and executing the above-mentioned data processing method.

[0016] On the other hand, an embodiment of the present application provides a computer program product or computer program, wherein the computer program product includes a computer program, and the computer program is stored in a storage medium; a processor reads the computer program from the storage medium, and the processor executes the computer program, so that the computer device performs the above-mentioned data processing method.

[0017] In an embodiment of the present application, the computer device can calculate the corresponding data sedimentation rate of the target data account through the total amount of data on the data platform and the corresponding remaining amount of data of the target data account in the data platform. It can be seen that the amount of data that the computer device needs to process when calculating the data sedimentation rate is relatively small, which reduces the computational complexity of the computer device when calculating the data sedimentation rate. In addition, since the computer device can more easily obtain the total amount of data and the remaining amount of data, the computational difficulty of the computer device when calculating the data sedimentation rate is effectively reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 is a schematic diagram of a data processing system provided in an embodiment of the present application;

[0020] Figure 2 is a schematic flow chart of a data processing method provided in an embodiment of the present application;

[0021] Figure 3 is a schematic flow chart of another data processing method provided in an embodiment of the present application;

[0022] Figure 4 This is a classification diagram of a target data account provided in an embodiment of the present application;

[0023] Figure 5 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0024] Figure 6 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The embodiment of the present application provides a data processing solution. When the data processing solution is used to calculate the data sedimentation rate, it can be achieved by obtaining the total amount of data on the data platform and the remaining amount of data corresponding to the target data account in the data platform, which effectively reduces the amount of data to be processed, thereby reducing the complexity of calculating the data sedimentation rate. Then, the difficulty of calculating the data sedimentation rate is further reduced. Among them, the target data account mentioned above can be a separate data account in the data platform, or it can be a data account cluster composed of multiple data accounts in the data platform. This application does not limit this, but unless otherwise specified, the following uses the target data account as an example to explain the embodiment of the present application in detail.

[0026] In one embodiment, the data processing scheme can be executed by a computer device, which can be a terminal device or a server, and this application does not impose any restrictions on this. Based on this, the general principle of the data processing scheme can be as follows: the computer device obtains the total amount of data on the data platform in the target time unit, the first data remaining amount of the target data account in the data platform in the target time unit, and the second data remaining amount of the target data account in the data platform in the previous time unit of the target time unit. Furthermore, the computer device can calculate the data sedimentation rate of the target data account in the target time unit based on the total amount of data, the first data remaining amount and the second data remaining amount. Among them, the target time unit can be understood as any time period, such as: a certain day, a certain week, etc.; then, correspondingly, the previous time unit of the target time unit can be understood as: the day before a certain day, the week before a certain week, etc.

[0027] In another embodiment, the above-mentioned data processing scheme can also be applied to Figure 1 In the data processing system shown in FIG. Figure 1 As shown, the data processing system includes a business server 10 and a target server 11, wherein the business server 10 is a server corresponding to the data platform, which is used to provide data management services; the target server 11 is used to calculate the data sedimentation rate corresponding to the target data account in the data platform, and the business server 10 and the target server 11 can establish a communication connection (or data connection). Based on this, the above-mentioned data processing scheme can be collaboratively executed by the business server 10 and the target server 11, and the specific implementation can be as follows: the target server 11 initiates a data connection request to the business server 10, and after the business server 10 responds to the data connection request and establishes a data connection with the target server 11, the target server 11 obtains from the business server 10 the total amount of data on the data platform in the target time unit, the first remaining amount of data of the target data account in the data platform in the target time unit, and the second remaining amount of data of the target data account in the data platform in the previous time unit of the target time unit; further, the target server 11 can calculate the data sedimentation rate of the target data account in the target time unit based on the acquired data.

[0028] It should be noted that the terminal devices mentioned above may include but are not limited to: smartphones, tablets, laptops, desktop computers, in-vehicle terminals, smart home appliances, smart watches, smart voice interaction devices, etc. In addition, the servers mentioned above may include but are not limited to: independent physical servers, server clusters or distributed systems composed of multiple physical servers, and cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. This application does not impose any restrictions on this.

[0029] For the sake of convenience, unless otherwise specified, the following description will take the case where the data processing solution is executed by a computer device as an example to describe in detail the relevant implementation methods in the embodiments of the present application.

[0030] Based on the above description, it is not difficult to understand that the embodiment of the present application can be widely used in a variety of data processing scenarios by virtue of its simple data processing method, such as: deposit sedimentation rate calculation scenario, employee retention rate calculation scenario, etc. In the deposit sedimentation rate calculation scenario, the data platform can be the XX Bank Fund Management Platform, the target data account can be the bank account of the user who has handled the deposit business in the bank (such as: account A), and the total amount of data can be the total amount of deposits stored in XX Bank. Then, based on this, the first data remaining amount can be the deposit balance of account A in XX Bank at the target time unit; the second data remaining amount can be: the deposit balance of account A in XX Bank at the previous time unit of the target time unit; the data sedimentation rate is the deposit sedimentation rate of account A in XX Bank. So, suppose that a computer device is now needed to calculate the deposit sedimentation rate of account A on X month 10, that is, the target time unit mentioned above can be understood as X month 10, and the previous time unit of the target time unit can be understood as X month 9. In this case, the computer device can obtain the total amount of deposits in XX Bank on X month 10, as well as the deposit balance a of account A in XX Bank on X month 10, and the deposit balance b of account A in XX Bank on X month 9; further, the computer device can calculate the deposit settlement rate of account A on X month 10 based on the obtained total amount of deposits, deposit balance a, and deposit balance b. It should be noted that the deposit settlement rate mentioned here can be the current deposit settlement rate or the time deposit settlement rate. This application does not impose any restrictions on this. Unless otherwise specified, the deposit settlement rate mentioned subsequently is explained using the current deposit settlement rate as an example.

[0031] Alternatively, in the employee retention rate calculation scenario, the data platform can be the employee management platform of XX Enterprise, the target data account can be a department within XX Enterprise (e.g., Department B), and the total data volume can be the total number of employees on the XX Enterprise employee management platform. Therefore, the first remaining amount of data can be the number of employees in Department B on the XX Enterprise employee management platform during the target time unit; the second remaining amount of data can be the number of employees in Department B on the XX Enterprise employee management platform during the time unit preceding the target time unit. The data accumulation rate is the employee retention rate, which can be used to reflect the retention and departure of employees in Department B during the target time unit. For this reason, suppose a computer device is required to calculate the employee retention rate of Department B in October. That is, the target time unit is October and the time unit preceding the target time unit is September. In this case, the computer device can obtain the total number of employees in XX Enterprise in October, as well as the number of employees c in Department B in October and the number of employees d in Department B in September. Furthermore, the computer device can calculate the employee retention rate of Department B in October based on the total number of employees, the number of employees c, and the number of employees d.

[0032] In order to more clearly understand the embodiments of the present application, unless otherwise specified, the following takes the data platform as a bank fund management platform as an example, and combines the deposit sedimentation rate calculation scenario to explain in detail the relevant implementation methods mentioned in the embodiments of the present application.

[0033] See Figure 2 , Figure 2 This is a schematic flow chart of a data processing method provided by an embodiment of the present application. This data processing method is proposed based on the above data processing solution. Therefore, it can be understood that this method can of course be executed by the above-mentioned computer device. Figure 2 As shown, the method includes steps S201-S203:

[0034] S201, obtaining the total amount of data on the data platform in the first time unit.

[0035] Among them, the first time unit can be understood as the target time unit mentioned above, that is, any time period. For example, the embodiment of the present application is explained by taking one time unit as one day. Based on the above description, it can be seen that the data platform in the embodiment of the present application can be a bank fund management platform. Then, in this case, the total amount of data on the data platform in the first time unit can be: the total amount of demand deposits of the bank fund management platform on a certain day. For example, the total amount of demand deposits on any day can be obtained from the fund record table of the bank fund management platform, and the fund record table can be exemplarily shown in Table 1:

[0036] Table 1

[0037] date Daily interbank balance Savings daily balance (normalized) Savings daily balance 20XX / 1 / 1 11736XXX.52 113067XXX.49 22964227XXX.63 20XX / 1 / 2 13648XXX.63 113683XXX.58 32564186XXX.33 20XX / 1 / 3 12836XXX.32 111137XXX.28 22741263XXX.63 ... ... ... ... XXXX / X / X XXXXX.XX XXXX.XX XXXX.XX

[0038] In Table 1 above, the total amount of demand deposits can be represented by the daily savings balance. Specifically, the daily savings balance corresponding to the date 20XX / 1 / 2 can represent: the total amount of demand deposits of the bank corresponding to the bank's fund management platform on the day of 20XX / 1 / 2.

[0039] S202 , obtaining a first remaining amount of data of a target data account on a data platform in a first time unit, and a second remaining amount of data of a target data account on a data platform in a second time unit.

[0040] In a specific embodiment, it can be seen from the above description that the target data account can be a single data account or a data account cluster consisting of multiple data accounts. Then, when the target data account is a single data account, the target data account can be any one of all the data accounts corresponding to the data platform; optionally, the target data account can also be any one of the data accounts in the sampled data obtained after the computer device samples data from all the data accounts of the data platform. When the target data account is a data account cluster, it can be all the data accounts corresponding to the data platform, or it can be all the data accounts corresponding to the sampled data obtained after the computer device samples data from all the data accounts of the data platform. This application does not impose specific restrictions on this. Among them, the specific way in which the computer device samples data for all the data accounts corresponding to the data platform will be described in detail in subsequent embodiments, and this application will not elaborate on it here.

[0041] In addition, the second time unit mentioned above is the time unit before the first time unit. Based on this, when the time unit is one day, the time period corresponding to the second time unit is the day before the time period corresponding to the first time unit. For example, assume that the first time unit is X month 9, and the second time unit is X month 8. Based on the above description, it is not difficult to understand that the target data account in the embodiment of the present application can be a bank account that stores funds in the bank's fund management platform. Based on this, for example, the first data remaining amount can be: the current deposit balance of the bank account on a certain day; the second data remaining amount is: the current deposit balance of the bank account on the day before that day. In this case, the current deposit balance of the bank account on each day can be called a position, that is, the first data remaining amount and the second data remaining amount obtained in the embodiment of the present application are essentially position data of the target data account.

[0042] S203: Determine a data sedimentation rate of the target data account in the first time unit according to the total amount of data, the first remaining amount of data, and the second remaining amount of data.

[0043] In a specific implementation, the computer device can determine the difference between the remaining amount of the first data and the remaining amount of the second data based on the remaining amount of the first data and the remaining amount of the second data; further, the computer device can perform a division operation on the difference between the remaining amount and the total amount of data, and obtain the data sedimentation rate of the target data account in the first time unit based on the operation result. Taking the deposit sedimentation rate calculation scenario as an example, in this scenario, the computer device can determine the data sedimentation rate of the target data account on a certain day based on the balance difference between the current deposit balance of the target data account on a certain day and the current deposit balance of the target data account on the day before that day. It is not difficult to see that in the deposit sedimentation rate calculation scenario, the computer device can determine the corresponding current deposit sedimentation rate of the target data account based on the position data of the target data account in certain time units.

[0044] In the embodiment of the present application, since the computer device can determine the data sedimentation rate based on the remaining amount of data in the target data account, when the computer device uses the embodiment of the present application to calculate the current deposit sedimentation rate of the target data account, the current deposit sedimentation rate of the target data account can be obtained specifically through the position data. This method avoids the computer device from obtaining the complex account transaction details of the target data account, and also avoids the computer device from performing a series of data preprocessing (such as data cleaning processing, data integration processing, data transformation processing, etc.) on the acquired data, thereby reducing the calculation complexity of the current deposit sedimentation rate and reducing the performance requirements of the computer device to a certain extent. In addition, since the amount of data corresponding to the position data of the target data account is small, when the current deposit sedimentation rate of the target data account is calculated based on the position data, the calculation difficulty of the computer device will be effectively reduced, thereby improving the calculation speed of the current deposit sedimentation rate of the computer device.

[0045] See Figure 3 , Figure 3 This is another data processing method provided by the embodiment of the present application. This method can also be executed by the computer device mentioned above, or by the target server and business server mentioned above. Unless otherwise specified, the following takes the computer device executing the data processing method alone as an example to describe in detail the various steps provided by the embodiment of the present application. Figure 3 As shown, the method includes steps S301-S306:

[0046] S301: Obtain the total amount of data accounts corresponding to each sampling time period in multiple sampling time periods.

[0047] As can be seen from the above, the target data account can be obtained after data sampling of all data accounts corresponding to the data platform, and in one embodiment, the data processing method provided by the embodiment of the present application can be executed solely by a computer device. Then, in this case, the computer device can first determine the sampling time interval, and then divide the sampling time interval into multiple sampling time periods. Further, the computer device can obtain the total amount of data accounts corresponding to each sampling time period to determine the target data account from all data accounts corresponding to the sampling time period based on the total amount of data accounts. Among them, the duration corresponding to each sampling time period in the multiple sampling time periods can be the same; the computer device can specifically determine the sampling time interval based on business needs. For example, if the computer device needs to formulate the fund management strategy that the bank fund management platform needs to adopt in the next three years based on the data sedimentation rate, then the computer device can use the previous three years as the sampling time interval based on this requirement to formulate the fund management strategy based on the data sedimentation rate corresponding to the data accounts in the previous three years.

[0048] As can be seen from the foregoing, in another embodiment, the embodiment of the present application can be executed collaboratively by the target server and the business server. That is to say, the data platform can correspond to multiple servers, including target servers and business servers. Among them, the target server is mainly used to obtain the remaining amount of data corresponding to the target data account to calculate the data sedimentation rate; the business server is mainly used to provide target services (such as: retail current deposit business, public current deposit business, interbank current deposit business, etc. in bank deposit business). Then, in this case, it is not difficult to understand that the target server can also be used to determine the target data account from all data accounts corresponding to the data platform, and the target data account can be a data account in the target business.

[0049] In a specific implementation, when determining the target data account, the target server can first determine a sampling time interval, then divide the sampling time interval into multiple sampling time periods, and further obtain the total number of data accounts corresponding to each sampling time period from the business server, so as to determine the target data account from all data accounts corresponding to the sampling time period on the business server based on the obtained total number of data accounts. It should be noted that in order to ensure data security, before the target server obtains the total number of data accounts corresponding to each sampling time period from the business server, the target server needs to send a data connection request to the business server, and the data connection request carries the identity of the target server. Then, after the business server receives the data connection request, it can parse the data connection request to obtain the identity of the target server, and then authenticate the target server based on the identity. Furthermore, if the business server successfully authenticates the target server, it establishes a data connection with the target server, so that the target server can obtain the total number of data accounts corresponding to each sampling time period from the business server through the data connection between the target server and the business server.

[0050] Optionally, the data platform may have multiple business servers, with different services provided by different business servers. For example, the aforementioned corporate demand deposit service is provided by business server A, the retail demand deposit service is provided by business server B, and the interbank demand deposit service is provided by business server C. In this case, these multiple business servers can form a blockchain network, with each business server serving as a node in the blockchain network. Data such as all data accounts corresponding to the data platform and their corresponding data account totals can be stored in a block within the blockchain network. In this case, when a target server wishes to obtain the total data account amount from any business server, it can first initiate a data sharing request to the business server. This data sharing request carries the target server's identity. Furthermore, the business server can authenticate the target server and, upon successful authentication, join the target server to the blockchain network where it resides, making it a node within the blockchain network. This allows the target server to obtain the total data account amount and identify the target data account from the block.

[0051] S302: Determine the extraction amount of data accounts corresponding to each sampling time period according to a preset sampling ratio and the total number of data accounts corresponding to each sampling time period.

[0052] Among them, in the multiple sampling time periods mentioned above, the preset sampling ratio used in each sampling time period is the same. For example, assuming that the sampling time interval is the previous 3 years, and the sampling time interval corresponds to 3 sampling time periods, namely the first year, the second year and the third year. Then, if the preset sampling ratio is 5%, it means that the computer device should extract 5% of the data accounts from all data accounts opened in the first year, and extract 5% of the data accounts from all data accounts opened in the second year, and extract 5% of the data accounts from all data accounts opened in the third year. Then, based on this, the computer device can determine the number of data accounts that need to be extracted from the sampling time period (i.e., the number of data accounts extracted) by multiplying the total number of data accounts of all data accounts corresponding to each sampling time period by the preset sampling ratio.

[0053] It can be seen from the above description that in the embodiment of the present application, a target server can be used to determine the amount of data account extraction corresponding to each sampling time period, and there can be multiple business servers that establish data connections with the target server. Based on this, it should be noted that in the deposit sedimentation rate calculation scenario, if there are multiple business servers, and different business servers provide different deposit services, then, illustratively, the target server can determine the target data account in the following way: if the business server is used to provide corporate demand deposit services or interbank demand deposit services, the target server can extract all data accounts corresponding to the business server; if the business server is used to provide retail demand deposit services, the target server can extract data accounts corresponding to the business server according to a preset sampling ratio. It is not difficult to understand that, since in actual applications, the data accounts corresponding to corporate demand deposit services and interbank demand deposit services are less than the data accounts corresponding to retail demand deposit services, therefore, when the data accounts corresponding to a certain business are less, full sampling is used, and when the data accounts corresponding to a certain business are more, partial sampling is used to ensure that the target data accounts are representative.

[0054] S303: Based on the amount of data accounts extracted in each sampling time period, determine a target data account from all data accounts corresponding to each sampling time period.

[0055] In a specific implementation, the computer device can randomly extract N data accounts from all data accounts corresponding to the sampling time period based on the amount of data accounts extracted in each sampling time period (assuming it is N, where N is a positive integer), and then determine the target data account based on the extracted N data accounts. Optionally, if the target data account is a data account cluster, the computer device can use all extracted data accounts as the target data account. Optionally, if the target data account is a single data account, the computer device can use any one of all extracted data accounts as the target data account. It is not difficult to understand that such a data sampling method can ensure that the number of data accounts corresponding to each sampling time period is the same in the total number of data accounts corresponding to the sampling time period (both are preset sampling ratios), thereby making the target data account more representative. Then, it is not difficult to understand that when the target data account is representative, the computer device can estimate the data sedimentation rate corresponding to the entire data platform based on the data sedimentation rate of the target data account. It can be seen that the data processing method provided in the embodiment of the present application has strong practicality.

[0056] S304: Obtain the total amount of data on the data platform in the first time unit.

[0057] S305: Obtain a first remaining amount of data of the target data account on the data platform in a first time unit, and a second remaining amount of data of the target data account on the data platform in a second time unit.

[0058] In one embodiment, the specific embodiments of steps S304 to S305 can refer to the relevant embodiments of the above-mentioned steps S201 to S202, and the embodiments of this application will not be repeated here.

[0059] S306: Determine a data settling rate of the target data account in the first time unit according to the total amount of data, the first remaining amount of data, and the second remaining amount of data.

[0060] In a specific embodiment, the computer device can use a trained data loss rate calculation model to determine the data loss rate of the target data account in the first time unit based on the total amount of data, the first data remaining amount and the second data remaining amount, so that the computer device can determine the data sedimentation rate of the target data account in the first time unit based on the data loss rate of the target data account in the first time unit. Exemplarily, data sedimentation rate = 1-data loss rate. Specifically, when determining the data loss rate of the target data account in the first time unit, the computer device can first calculate the difference between the remaining amount of the first data remaining amount and the remaining amount of the second data, and then perform a division operation based on the difference in the remaining amount and the total amount of data, and use the obtained operation result as the data loss rate.

[0061] Among them, before the computer device uses the trained data loss rate calculation model to calculate the data loss rate, it can first train the data loss rate calculation model. The specific training method can be as follows: the computer device first obtains training data, and the training data includes: the total amount of training data in the first historical time unit, the remaining amount of training data and the standard data loss rate of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in the second historical time unit. Furthermore, after the computer device obtains the training data, the data loss rate calculation model can be used to determine the actual data loss rate of the sample account in the first historical time unit based on the total amount of training data in the first historical time unit, the remaining amount of training data of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in the second historical time unit. Based on this, the computer device can train the data loss rate calculation model based on the actual data loss rate and the standard data loss rate to obtain a trained data loss rate calculation model. In a specific implementation, the computer device can run an optimization algorithm during the process of determining the actual data churn rate using the data churn rate calculation model based on the difference between the actual data churn rate and the standard data churn rate, thereby training the data churn rate calculation model once, thereby obtaining a trained data churn rate calculation model. The optimization algorithm optimizes the data churn rate calculation model, and the optimization goal of the optimization object is to ensure that the difference between the actual data churn rate and the standard data churn rate is less than a difference threshold.

[0062] Among them, the sample account mentioned above can be any account different from the target data account, and its specific determination method can refer to the above description of the determination method of the target data account, and the embodiments of this application will not be repeated here. The standard data loss rate refers to: the correct data loss rate of the training account in the first time unit; the second historical time unit is the historical time unit before the first historical time unit, and the historical time unit can be any time unit. Then, it is not difficult to understand that in a specific application, the training data can be: the computer device obtains the first time unit and the second time unit corresponding to the current data sedimentation rate calculation process. That is to say, in this case, the historical time unit can be specifically: the first time unit or the second time unit corresponding to the data sedimentation rate calculated this time. Optionally, the training data can also be: the computer device obtains other time units other than the first time unit and the second time unit corresponding to the current data sedimentation rate calculation process. That is to say, in this case, the historical time unit can be specifically: other time units except the first time unit and the second time unit corresponding to the data sedimentation rate calculated this time.

[0063] In one embodiment, based on the above Figure 2 and Figure 3From the description of the relevant embodiments, it is not difficult to understand that the computer device can also determine the data sedimentation rate of the target data account in the target time interval, wherein the target time interval includes multiple time units, and these multiple time units are continuous in time. In other words, if the duration corresponding to the time unit is one day, the target time interval can be a number of consecutive days. Based on this, the following describes the method for the computer device to determine the data sedimentation rate of the target data account in the target time interval with reference to specific examples.

[0064] Assuming that the target time interval is from X month 5 to X month 14, then, in this case, it is not difficult to understand that the target time interval includes 10 time units, namely: X month 5, X month 6, ..., X month 13, X month 14; it is not difficult to see that these 10 time units are continuous in time. Then, based on the above description, the computer device can obtain the data sedimentation rate of the target time interval in the following manner: the computer device first obtains the data sedimentation rate of the target data account in each time unit of the multiple time units, wherein the method for the computer device to obtain the data sedimentation rate of any time unit can refer to the above description of the computer device obtaining the data sedimentation rate of the first time unit, which will not be repeated here in this application. Then, further, the computer device can determine the data sedimentation rate of the target data account in the target time interval based on the data sedimentation rate of the target data account in each time unit. For example, the computer device can perform weighted averaging based on the data sedimentation rate of each time unit in the multiple time units to obtain the data sedimentation rate of the target time interval. In actual application, the format of the target time interval and the data sedimentation rate corresponding to the target time interval can be exemplarily shown in Table 2.

[0065] Table 2

[0066]

[0067] The following takes the data corresponding to X / 1 / 1 as an example to explain in detail the meaning of the data in Table 2. X / 1 / 1 means that the computer device is the total amount of data on a certain data platform on January 1, X. 1D means that the target time interval is the day after January 1, X, that is, January 2, X. Similarly, 2-7D means that the target time interval is from the 2nd to the 7th day after January 1, X, that is, January 3, X to January 8, X; based on this, it is not difficult to understand that 8-31D means that the target time interval is from January 9, X to February 1, X. Similarly, the meaning of 31-90D, 90-180D, 180-365D and >365D can be understood, and this application will not elaborate on this. In addition, "3254XX" in Table 2 represents the deposit balance of the target data account on January 1, year X; "0.60%" in Table 2 represents the data sedimentation rate of the target data account in the time interval January 2, year X; similarly, "1.88%" represents the data sedimentation rate of the target data account in the time interval January 3, year X to January 8, year X.

[0068] In another embodiment, it is not difficult to understand based on the above description that the computer device can determine the data sedimentation rate of the target data account in each of the multiple time intervals based on the above method of determining the data sedimentation rate of the target time interval, and further determine the target sedimentation rate influencing factor that affects the data sedimentation rate based on the data sedimentation rate of each time interval, so that relevant personnel can specify data management strategies for the data platform based on the target sedimentation rate influencing factor. Among them, the multiple time intervals mentioned above have a chronological order in time, that is, the time corresponding to each time interval is different, and the target sedimentation rate influencing factor can be one or more. The following is an example of the execution process of the data management strategy in conjunction with the current deposit sedimentation rate calculation scenario. After determining the target sedimentation rate influencing factor, the relevant data management personnel can formulate the data management strategy of the data platform based on the target sedimentation rate influencing factor. For example, the relevant data management personnel can configure the current deposit sedimentation rate parameters in the FTP (Funds Transfer Pricing, internal funds transfer pricing) system and asset-liability management system corresponding to the data platform so that the data platform can more accurately measure cash flow. In addition, relevant data management personnel can also specify interest rate risk management strategies for operating units at all levels based on the target sedimentation rate influencing factors to enhance the data platform's ability to manage funds in a refined manner.

[0069] Specifically, the computer device can determine the target sedimentation rate influencing factor in the following manner: the computer device first obtains multiple sedimentation rate influencing factors and the data sedimentation rate of the target data account in each of multiple time intervals. For example, in the current deposit sedimentation rate calculation scenario, the sedimentation rate influencing factor can be obtained by the computer device from market data related to the data platform. The market data includes but is not limited to: economic growth data, capital market data, monetary policy information, price index, etc. Based on this, the sedimentation rate influencing factor can be monetary policy or capital market, etc., which is mainly used to analyze the correlation between changes in the current deposit sedimentation rate and market data.

[0070] Furthermore, after the computer device obtains multiple sedimentation rate influencing factors and the data sedimentation rate of the target data account in each time interval, data analysis can be performed based on each sedimentation rate influencing factor and the data sedimentation rate of the target data account in each time interval to determine the correlation coefficient of each sedimentation rate influencing factor. Among them, the so-called data analysis can specifically refer to linear regression analysis. Based on this, further, the computer device can determine the target sedimentation rate influencing factor from multiple sedimentation rate influencing factors based on the correlation coefficient of each sedimentation rate influencing factor. It should be noted that the correlation coefficient can be used to reflect the strength of the correlation between the corresponding sedimentation rate influencing factor and the data sedimentation rate. The larger the correlation coefficient, the greater the impact of the corresponding sedimentation rate influencing factor on the data sedimentation rate. That is to say, in this case, if the value of the sedimentation rate influencing factor (such as: the specific policy content of monetary policy, the specific market data of the capital market) changes, the data sedimentation rate will also change significantly.

[0071] Then, based on the above description, it is not difficult to understand that in actual applications, the computer device can use the sedimentation rate influencing factor with a correlation coefficient greater than the coefficient threshold as the target sedimentation rate influencing factor, or the sedimentation rate influencing factor with the largest correlation coefficient as the target sedimentation rate influencing factor, or arrange all the determined correlation coefficients in order from large to small, and use the sedimentation rate influencing factors corresponding to the correlation coefficients arranged in the first P (P is a positive integer greater than 2) positions as the target sedimentation rate influencing factors, or arrange all the determined correlation coefficients in order from small to large, and use the sedimentation rate influencing factors corresponding to the correlation coefficients arranged in the last P positions as the target sedimentation rate influencing factors. This application does not impose any restrictions on this.

[0072] It should be noted that, in specific applications, the computer device can also group all the extracted data accounts according to the account categories to which the data accounts belong, and obtain multiple data account groups. For each data account group, the computer device can sequentially obtain one of the data accounts as the target data account, and calculate the data sedimentation rate corresponding to the target data account, thereby determining the target sedimentation rate influencing factor for each category of business, so as to formulate a more refined data management strategy. For example, in the current deposit sedimentation rate calculation scenario, the category of the data account can be as follows: Figure 4 shown.

[0073] In the embodiment of the present application, since the target data account can be determined after data sampling of all data accounts on the data platform, and the number of data accounts extracted by the computer device in each sampling period accounts for the same proportion of the total number of data accounts corresponding to the sampling period, the data sedimentation rate determined based on the extracted target data account can accurately reflect the actual business situation. Based on this, and since the present application can be used to calculate the current deposit sedimentation rate, and the current deposit sedimentation rate can be applied in bank interest rate risk management, bank personnel can use the current deposit sedimentation rate calculated by the present application based on the computer device to more accurately formulate fund management plans, new business development plans, etc., thereby improving the bank's management efficiency and business management capabilities. In addition, in the embodiment of the present application, position data is used when applying to the calculation of the deposit sedimentation rate. Since the amount of data corresponding to the position data is small, the computer device can use the simplest method to realize the calculation of the deposit sedimentation rate and the correlation analysis with the sedimentation rate influencing factors, effectively reducing the difficulty of calculating the deposit sedimentation rate. In addition, since the data processing method provided by the present application can also be implemented in combination with blockchain, the data security of the embodiment of the present application is highly guaranteed, which enhances the practicality of the present application to a certain extent.

[0074] Based on the description of the above data processing method embodiments, the present application also discloses a data processing device, which can be a computer program (including program code) running on the above-mentioned computer device. Figure 2 or Figure 3 See the method shown in Figure 5 , the data processing device may include: an acquisition unit 501 and a determination unit 502.

[0075] An acquisition unit 501 is used to acquire the total amount of data on the data platform in a first time unit;

[0076] The obtaining unit 501 is further configured to obtain a first remaining amount of data of the target data account on the data platform at the first time unit, and a second remaining amount of data of the target data account on the data platform at the second time unit; wherein the second time unit is a time unit before the first time unit;

[0077] The determination unit 502 is configured to determine a data sedimentation rate of the target data account in the first time unit based on the total amount of data, the first remaining amount of data, and the second remaining amount of data.

[0078] In one embodiment, the acquiring unit 501 is further configured to execute:

[0079] Obtaining a data sedimentation rate of the target data account in each of a plurality of time units, where the plurality of time units are continuous in time;

[0080] The data sedimentation rate of the target data account in the target time interval is determined according to the data sedimentation rate of the target data account in each time unit, where the target time interval includes the multiple time units.

[0081] In yet another embodiment, the acquiring unit 501 is further configured to execute:

[0082] Obtaining multiple sedimentation rate influencing factors and the data sedimentation rate of the target data account in each time interval of multiple time intervals, where the multiple time intervals are arranged in a chronological order;

[0083] Performing data analysis based on each of the plurality of sedimentation rate influencing factors and the data sedimentation rate of the target data account in each time interval to determine a correlation coefficient of each sedimentation rate influencing factor;

[0084] Based on the correlation coefficient of each sedimentation rate influencing factor, a target sedimentation rate influencing factor is determined from the multiple sedimentation rate influencing factors, and the target sedimentation rate influencing factor is used to formulate a data management strategy for the data platform.

[0085] In yet another embodiment, the determining unit 502 is further configured to execute:

[0086] Get the total amount of data accounts corresponding to each sampling time period in multiple sampling time periods;

[0087] Determine the amount of data accounts to be extracted for each sampling period based on a preset sampling ratio and the total number of data accounts corresponding to each sampling period;

[0088] Based on the data account extraction amount of each sampling time period, the target data account is determined from all data accounts corresponding to each sampling time period.

[0089] In another embodiment, the data processing device may be run in a computer device. Taking the computer device as a target server as an example, the target data account is a data account in the target business; the determining unit 502 is further configured to execute:

[0090] Initiate a data connection request to the service server corresponding to the target service, wherein the data connection request carries the identity information of the target server, so that the service server authenticates the target server based on the identity information. If the service server successfully authenticates the target server, the service server establishes a data connection with the target server;

[0091] The total amount of data accounts corresponding to each sampling time period in the plurality of sampling time periods is obtained from the business server through a data connection with the business server.

[0092] In yet another embodiment, the acquiring unit 501 is further configured to execute:

[0093] Using a trained data loss rate calculation model, determine the data loss rate of the target data account in the first time unit according to the total amount of data, the first remaining amount of data, and the second remaining amount of data;

[0094] A data sedimentation rate of the target data account in the first time unit is determined based on the data loss rate of the target data account in the first time unit.

[0095] In yet another embodiment, the data processing apparatus further includes a training unit 503, which can be configured to execute:

[0096] Acquire training data, the training data including: the total amount of training data in a first historical time unit, the remaining amount of training data and the standard data churn rate of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in a second historical time unit;

[0097] Determine the actual data churn rate of the sample account in the first historical time unit using a data churn rate calculation model based on the total amount of training data in the first historical time unit, the remaining amount of training data of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in the second historical time unit;

[0098] Based on the actual data loss rate and the standard data loss rate, the data loss rate calculation model is trained to obtain the trained data loss rate calculation model.

[0099] According to another embodiment of the present application, Figure 2 and Figure 3 The various steps involved in the method shown can be performed by Figure 5 The data processing apparatus shown is executed by each unit. For example: Figure 2 The steps S201 and S202 shown can be both Figure 5 The data processing device shown in FIG. 1 is used to obtain the data. The acquisition unit 501 in FIG. 1 is used to perform step S203. Figure 5 The determination unit 502 in the data processing device shown in FIG. Figure 3 Steps S301 to S303 and step S306 shown in FIG. Figure 5 The determination unit 502 in the data processing device shown in FIG. 1 is executed, and steps S304 to S305 can be performed by Figure 5 The acquisition unit 501 in the data processing device shown is executed.

[0100] According to another embodiment of the present application, Figure 5 The various units in the data processing device shown are divided based on logical functions. The above-mentioned various units can be individually or all combined into one or more other units to form a structure, or one (or some) of the units can be further divided into multiple functionally smaller units to form a structure, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. In other embodiments of the embodiments of the present application, the data processing device may also include other units. In actual applications, these functions can also be assisted by other units and can be implemented by the collaboration of multiple units.

[0101] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. Figure 2 or Figure 3 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 5 The computer program can be recorded on a storage medium, for example, and loaded into the computing device through the storage medium and run therein.

[0102] In an embodiment of the present application, the data processing device can calculate the corresponding data sedimentation rate of the target data account through the total amount of data on the data platform and the corresponding remaining amount of data of the target data account in the data platform. It can be seen that the amount of data that the data processing device needs to process when calculating the data sedimentation rate is relatively small, which reduces the computational complexity of the computer device when calculating the data sedimentation rate. In addition, since the data processing device can more easily obtain the total amount of data and the remaining amount of data, the computational difficulty of the data processing device when calculating the data sedimentation rate is effectively reduced.

[0103] Based on the above descriptions of the method embodiment and the apparatus embodiment, the present application also provides a computer device. Figure 6 The computer device includes at least a processor 601 and a storage medium 602, and the processor 601 and the storage medium 602 in the computer device can be connected via a bus or other means.

[0104] The storage medium 602 is a memory device in a computer device, used to store programs and data. It is understandable that the storage medium 602 here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The storage medium 602 provides a storage space, which stores the operating system of the computer device. In addition, one or more computer instructions suitable for being loaded and executed by the processor 601 are also stored in the storage space. These computer instructions can be one or more computer programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one storage medium located away from the aforementioned processor. The processor 601 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0105] In one embodiment, the processor 601 may load and execute one or more computer instructions stored in the storage medium 602 to implement the above-mentioned Figure 2 and Figure 3 The corresponding method steps in the data processing method embodiment shown; in a specific implementation, one or more computer instructions in the computer storage medium 602 are loaded by the processor 601 and execute the following steps:

[0106] Obtain the total amount of data on the data platform in the first time unit;

[0107] Obtaining a first remaining amount of data in the target data account on the data platform at the first time unit, and a second remaining amount of data in the target data account on the data platform at the second time unit; wherein the second time unit is a time unit before the first time unit;

[0108] A data sedimentation rate of the target data account in the first time unit is determined according to the total amount of data, the first remaining amount of data, and the second remaining amount of data.

[0109] In one embodiment, the processor 601 is specifically configured to load and execute:

[0110] Obtaining a data sedimentation rate of the target data account in each of a plurality of time units, where the plurality of time units are continuous in time;

[0111] The data sedimentation rate of the target data account in the target time interval is determined according to the data sedimentation rate of the target data account in each time unit, where the target time interval includes the multiple time units.

[0112] In yet another embodiment, the processor 601 is specifically configured to load and execute:

[0113] Obtaining multiple sedimentation rate influencing factors and the data sedimentation rate of the target data account in each time interval of multiple time intervals, where the multiple time intervals are arranged in a chronological order;

[0114] Performing data analysis based on each of the plurality of sedimentation rate influencing factors and the data sedimentation rate of the target data account in each time interval to determine a correlation coefficient of each sedimentation rate influencing factor;

[0115] Based on the correlation coefficient of each sedimentation rate influencing factor, a target sedimentation rate influencing factor is determined from the multiple sedimentation rate influencing factors, and the target sedimentation rate influencing factor is used to formulate a data management strategy for the data platform.

[0116] In yet another embodiment, the processor 601 is specifically configured to load and execute:

[0117] Get the total amount of data accounts corresponding to each sampling time period in multiple sampling time periods;

[0118] Determine the amount of data accounts to be extracted for each sampling period based on a preset sampling ratio and the total number of data accounts corresponding to each sampling period;

[0119] Based on the data account extraction amount of each sampling time period, the target data account is determined from all data accounts corresponding to each sampling time period.

[0120] In another embodiment, taking the computer device as the target server as an example, the target data account is a data account in the target business; the processor 601 is further configured to execute:

[0121] Initiate a data connection request to the service server corresponding to the target service, wherein the data connection request carries the identity information of the target server, so that the service server authenticates the target server based on the identity information. If the service server successfully authenticates the target server, the service server establishes a data connection with the target server;

[0122] The total amount of data accounts corresponding to each sampling time period in the plurality of sampling time periods is obtained from the business server through a data connection with the business server.

[0123] In yet another embodiment, the processor 601 is specifically configured to load and execute:

[0124] Using a trained data loss rate calculation model, determine the data loss rate of the target data account in the first time unit according to the total amount of data, the first remaining amount of data, and the second remaining amount of data;

[0125] A data sedimentation rate of the target data account in the first time unit is determined based on the data loss rate of the target data account in the first time unit.

[0126] In yet another embodiment, the processor 601 is specifically configured to load and execute:

[0127] Acquire training data, the training data including: the total amount of training data in a first historical time unit, the remaining amount of training data and the standard data churn rate of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in a second historical time unit;

[0128] Determine the actual data churn rate of the sample account in the first historical time unit using a data churn rate calculation model based on the total amount of training data in the first historical time unit, the remaining amount of training data of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in the second historical time unit;

[0129] Based on the actual data loss rate and the standard data loss rate, the data loss rate calculation model is trained to obtain the trained data loss rate calculation model.

[0130] In an embodiment of the present application, the computer device can calculate the corresponding data sedimentation rate of the target data account through the total amount of data on the data platform and the corresponding remaining amount of data of the target data account in the data platform. It can be seen that the amount of data that the computer device needs to process when calculating the data sedimentation rate is relatively small, which reduces the computational complexity of the computer device when calculating the data sedimentation rate. In addition, since the computer device can more easily obtain the total amount of data and the remaining amount of data, the computational difficulty of the computer device when calculating the data sedimentation rate is effectively reduced.

[0131] The present application also provides a storage medium that stores one or more computer instructions of the above-mentioned data processing method. When one or more processors load and execute the computer instructions, the description of the data processing method in the data processing method embodiment can be implemented, which will not be repeated here. The description of the beneficial effects of adopting the same method will not be repeated here. It is understood that the computer instructions can be deployed on one or more devices that can communicate with each other for execution.

[0132] It should be noted that according to one aspect of the embodiments of the present application, a computer program product or computer program is also provided, the computer program product or computer program includes computer instructions, and the computer instructions are stored in a storage medium. The processor of the computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer device performs the above Figure 2 and Figure 3 The data processing method shown is provided in various optional aspects of the embodiments.

[0133] In addition, it should be understood that the various embodiments disclosed above are merely preferred embodiments of the present application and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A data processing method, characterized in that: include: Obtain the total amount of data on the data platform in the first time unit; Obtaining a first remaining amount of data in the target data account on the data platform at the first time unit, and a second remaining amount of data in the target data account on the data platform at the second time unit; wherein the second time unit is a time unit before the first time unit; The data sedimentation rate of the target data account in the first time unit is determined based on the total amount of data, the first remaining amount of data, and the second remaining amount of data; if the data processing scenario corresponding to the data sedimentation rate is an employee retention rate calculation scenario, the data platform is an employee management desk of a target enterprise, the total amount of data is the total number of employees of the employee management desk of the target enterprise, the target data account is a target department in the target enterprise, the first remaining amount of data and the second remaining amount of data are used to indicate the number of employees, and the data sedimentation rate is an employee retention rate; wherein, data sedimentation rate = 1 - data loss rate, and the data loss rate of the target data account in the first time unit is determined by calculating the difference between the first remaining amount of data and the second remaining amount of data, dividing the difference by the total amount of data, and using the result of the calculation as the data loss rate; The method is executed by a target server, and the target data account is a data account in a target business; the method further includes: Initiate a data connection request to the service server corresponding to the target service, wherein the data connection request carries the identity information of the target server, so that the service server authenticates the target server based on the identity information. If the service server successfully authenticates the target server, the service server establishes a data connection with the target server; Obtaining, from the business server through a data connection with the business server, a total amount of data accounts corresponding to each sampling time period in a plurality of sampling time periods; Determine the amount of data accounts to be extracted for each sampling period based on a preset sampling ratio and the total number of data accounts corresponding to each sampling period; Based on the data account extraction amount of each sampling time period, determining the target data account from all data accounts corresponding to each sampling time period; Obtaining multiple sedimentation rate influencing factors and the data sedimentation rate of the target data account in each time interval of multiple time intervals, where the multiple time intervals are arranged in a chronological order; performing data analysis based on each of the plurality of sedimentation rate influencing factors and the data sedimentation rate of the target data account in each time interval to determine a correlation coefficient for each sedimentation rate influencing factor; the correlation coefficient is used to reflect the strength of the correlation between the corresponding sedimentation rate influencing factor and the data sedimentation rate, with a larger correlation coefficient indicating a greater impact of the corresponding sedimentation rate influencing factor on the data sedimentation rate; Based on the correlation coefficient of each sedimentation rate influencing factor, determining a sedimentation rate influencing factor with a large correlation coefficient from the multiple sedimentation rate influencing factors as a target sedimentation rate influencing factor, wherein the target sedimentation rate influencing factor is used to formulate a data management strategy for the data platform; Among them, there are multiple business servers corresponding to the data platform, and different businesses are provided by different business servers; the multiple business servers form a blockchain network, each business server is a node in the blockchain network, and all data accounts corresponding to the data platform and the corresponding total amount of data accounts constitute a block stored in the blockchain network; when the target server needs to obtain the total amount of data accounts from any business server, it initiates a data sharing request to the business server, and after the business server authenticates the target server, it adds the target server to the blockchain network where the business server is located, so that the target server becomes a node in the blockchain network, and obtains the total amount of data accounts and determines the target data account from the corresponding block.

2. The method according to claim 1, characterized in that The method further comprises: Obtaining a data sedimentation rate of the target data account in each of a plurality of time units, where the plurality of time units are continuous in time; The data sedimentation rate of the target data account in the target time interval is determined according to the data sedimentation rate of the target data account in each time unit, where the target time interval includes the multiple time units.

3. The method according to claim 1, characterized in that The determining, based on the total amount of data, the first remaining amount of data, and the second remaining amount of data, a data settling rate of the target data account in the first time unit includes: Using a trained data loss rate calculation model, determine the data loss rate of the target data account in the first time unit according to the total amount of data, the first remaining amount of data, and the second remaining amount of data; A data sedimentation rate of the target data account in the first time unit is determined based on the data loss rate of the target data account in the first time unit.

4. The method according to claim 3, characterized in that The method further comprises: Acquire training data, the training data including: the total amount of training data in a first historical time unit, the remaining amount of training data and the standard data churn rate of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in a second historical time unit; Determine the actual data churn rate of the sample account in the first historical time unit using a data churn rate calculation model based on the total amount of training data in the first historical time unit, the remaining amount of training data of the sample account in the first historical time unit, and the remaining amount of training data of the sample account in the second historical time unit; Based on the actual data loss rate and the standard data loss rate, the data loss rate calculation model is trained to obtain the trained data loss rate calculation model.

5. A data processing device, characterized in that: include: An acquisition unit, used to acquire the total amount of data from the data platform in the first time unit; The acquisition unit is further configured to acquire a first remaining amount of data of the target data account on the data platform at the first time unit, and a second remaining amount of data of the target data account on the data platform at a second time unit; wherein the second time unit is a time unit before the first time unit; a determination unit for determining a data sedimentation rate of the target data account in the first time unit based on the total amount of data, the first remaining amount of data, and the second remaining amount of data; if the data processing scenario corresponding to the data sedimentation rate is an employee retention rate calculation scenario, the data platform is an employee management desk of a target enterprise, the total amount of data is the total number of employees of the employee management desk of the target enterprise, the target data account is a target department in the target enterprise, the first remaining amount of data and the second remaining amount of data are used to indicate the number of employees, and the data sedimentation rate is an employee retention rate; wherein, data sedimentation rate = 1 - data loss rate, and the data loss rate of the target data account in the first time unit is determined by calculating the difference between the remaining amount of the first remaining amount of data and the remaining amount of the second remaining amount of data, dividing the difference between the remaining amount and the total amount of data, and using the obtained calculation result as the data loss rate; The determining unit is further configured to initiate a data connection request to a business server corresponding to the target business, the data connection request carrying identity information of the target server, so that the business server performs authentication processing on the target server based on the identity information; if the business server successfully authenticates the target server, the business server establishes a data connection with the target server; obtain, from the business server through the data connection with the business server, the total number of data accounts corresponding to each sampling time period in a plurality of sampling time periods; determine, based on a preset sampling ratio and the total number of data accounts corresponding to each sampling time period, a data account extraction amount corresponding to each sampling time period; and determine, based on the data account extraction amount for each sampling time period, the target data account from all data accounts corresponding to each sampling time period; The acquisition unit is further configured to acquire a plurality of sedimentation rate influencing factors and the data sedimentation rate of a target data account in each of a plurality of time intervals, wherein the plurality of time intervals are arranged in a temporal order; perform data analysis based on each sedimentation rate influencing factor in the plurality of sedimentation rate influencing factors and the data sedimentation rate of the target data account in each of the time intervals to determine a correlation coefficient of each sedimentation rate influencing factor; the correlation coefficient is configured to reflect the strength of the correlation between the corresponding sedimentation rate influencing factor and the data sedimentation rate, wherein a larger correlation coefficient indicates a greater influence of the corresponding sedimentation rate influencing factor on the data sedimentation rate; based on the correlation coefficient of each sedimentation rate influencing factor, determine a sedimentation rate influencing factor with a larger correlation coefficient from the plurality of sedimentation rate influencing factors as a target sedimentation rate influencing factor, and the target sedimentation rate influencing factor is configured to be used for formulating a data management strategy for the data platform; Among them, there are multiple business servers corresponding to the data platform, and different businesses are provided by different business servers; the multiple business servers form a blockchain network, each business server is a node in the blockchain network, and all data accounts corresponding to the data platform and the corresponding total amount of data accounts constitute a block stored in the blockchain network; when the target server needs to obtain the total amount of data accounts from any business server, it initiates a data sharing request to the business server, and after the business server authenticates the target server, it adds the target server to the blockchain network where the business server is located, so that the target server becomes a node in the blockchain network, and obtains the total amount of data accounts and determines the target data account from the corresponding block.

6. A computer device, characterized in that: include: a processor adapted to execute one or more computer instructions; A storage medium storing one or more computer instructions, wherein the one or more computer instructions are suitable for being loaded by the processor and executing the data processing method according to any one of claims 1 to 4.

7. A storage medium, characterized in that: The storage medium stores one or more computer instructions, and the one or more computer instructions are suitable for being loaded by a processor and executing the data processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Communication method, communication system and device, server and storage medium

    CN112235400A

  • Live broadcast target user estimation method and device, electronic equipment and computer readable storage medium

    CN113256342A