Data processing method and device, equipment, storage medium and program product

CN121187508BActive Publication Date: 2026-09-15CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511327883.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2026-09-15
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

[0004]本发明实施例提供一种数据处理方法、装置、设备、存储介质及程序产品,以解决相关技术中存在数据删除的效果较差的问题

Benefits of technology

[0010] In a sixth aspect, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the data processing method described in the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121187508B_ABST
    Figure CN121187508B_ABST
Patent Text Reader

Abstract

The application provides a data processing method, device, equipment, storage medium and program product, and relates to the technical field of data processing. The method comprises the following steps: obtaining a quality coefficient and a structure parameter of each first service in a plurality of first services from a first storage device, wherein the quality coefficient is used for representing the number of times of updating the first service, and the structure parameter is used for representing the position of data of the first service; calculating a set quality threshold based on the quality coefficient and the structure parameter of the plurality of first services; filtering the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each second service in the plurality of second services is greater than the set quality threshold; retaining service data of the plurality of second services in the first storage device, and deleting service data of other services, wherein the other services are services other than the plurality of second services in the plurality of first services. The application can improve the effect of data deletion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a data processing method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the development and application of big data and big model technologies, the demand for server storage resources is constantly increasing. In related technologies, to meet this growing demand, a combination of building storage devices and deleting unnecessary data is typically used. Deleting unnecessary data usually involves setting a deletion time; data is deleted when its storage time exceeds the set time. However, in these technologies, deleting data by setting a deletion time can lead to the deletion of data that users need even after the set time has expired, resulting in poor data deletion effectiveness.

[0003] It is evident that the relevant technologies suffer from poor data deletion effectiveness. Summary of the Invention

[0004] This invention provides a data processing method, apparatus, device, storage medium, and program product to solve the problem of poor data deletion effect in related technologies.

[0005] To solve the above problems, the present invention is implemented as follows: In a first aspect, embodiments of the present invention provide a data processing method, including: The quality coefficient and structural parameters of each of the multiple first services are obtained from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. Based on the quality coefficients and structural parameters of the multiple first services, a set quality threshold is calculated. Based on the set quality threshold, the plurality of first services are filtered to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The service data of the plurality of second services in the first storage device is retained, and the service data of other services are deleted. The other services are services other than the plurality of second services among the plurality of first services.

[0006] Secondly, embodiments of the present invention also provide a data processing apparatus, comprising: The data mining module is used to acquire business data and historical quality coefficients of multiple first-level services; A data magnetic separation module is communicatively connected to the data mining module. The data magnetic separation module is used to generate structural data corresponding to the plurality of first services and to obtain the quality coefficients corresponding to the plurality of first services. A data extraction module, which is communicatively connected to the data magnetic separation module, is used to calculate a set quality threshold based on the quality coefficients and structural parameters of the plurality of first services; filter the plurality of first services based on the set quality threshold to obtain a plurality of second services; retain the service data of the plurality of second services in the first storage device, and delete the service data of other services; The data flow module is communicatively connected to the data extraction module. The data flow module is used to read business data, quality coefficients, structural parameters, and second identifiers from multiple storage structures; and to send the business data, quality coefficients, structural parameters, and second identifiers to external devices through a preset interface.

[0007] Thirdly, embodiments of the present invention also provide a data processing apparatus, comprising: The first acquisition module is used to acquire the quality coefficient and structural parameters of each of the multiple first services from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. The calculation module is used to calculate a set quality threshold based on the quality coefficients and structural parameters of the multiple first services; A filtering module is used to filter the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The first storage module is used to retain the service data of the plurality of second services in the first storage device and delete the service data of other services, wherein the other services are services other than the plurality of second services among the plurality of first services.

[0008] Fourthly, embodiments of the present invention also provide an electronic device, including a transceiver and a processor. The processor is configured to obtain the quality coefficient and structural parameters of each of the multiple first services from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. The processor is also configured to calculate a set quality threshold based on the quality coefficients and structural parameters of the plurality of first services. The processor is further configured to filter the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The processor is further configured to retain the service data of the plurality of second services in the first storage device and delete the service data of other services, wherein the other services are services other than the plurality of second services among the plurality of first services.

[0009] Fifthly, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the data processing method described in the first aspect.

[0010] In a sixth aspect, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the data processing method described in the first aspect.

[0011] In a seventh aspect, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the data processing method described in the first aspect.

[0012] In this embodiment of the invention, the quality coefficient and structural parameters of each of a plurality of first services are obtained from a first storage device. The quality coefficient characterizes the number of times the first service is updated, and the structural parameters characterize the location of the data in the first service. Based on the quality coefficient and structural parameters of the plurality of first services, a set quality threshold is calculated. Based on the set quality threshold, the plurality of first services are filtered to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The service data of the plurality of second services in the first storage device is retained, while the service data of other services (services other than the plurality of second services in the plurality of first services) is deleted. In this way, by setting a quality threshold to filter and obtain a plurality of second services, the plurality of second services are considered to be service data with high importance or value. By retaining the service data of the plurality of second services in the first storage device and deleting the service data of other services, the use of storage resources is reduced while retaining important or valuable service data, thereby improving the data deletion effect. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of a data processing method provided in an embodiment of the present invention; Figure 2 This is a structural diagram of a data processing device provided in an embodiment of the present invention; Figure 3 This is a structural diagram of the data mining module provided in an embodiment of the present invention; Figure 4 This is a structural diagram of the data magnetic separation module provided in an embodiment of the present invention; Figure 5 This is a structural diagram of the data extraction module provided in an embodiment of the present invention; Figure 6 This is a structural diagram of the data flow turnover module provided in an embodiment of the present invention; Figure 7 This is a structural diagram of a data processing device provided in an embodiment of the present invention; Figure 8 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Please see Figure 1 , Figure 1 This is a flowchart of a data processing method provided in an embodiment of the present invention, such as... Figure 1 As shown, it includes the following steps: Step 101: Obtain the quality coefficient and structural parameters of each of the multiple first services from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service.

[0017] The aforementioned first storage device is used to store the business data of the first service, as well as the quality coefficient and structure coefficient of the first service. Specifically, it can be a high-speed storage module or a database.

[0018] The quality coefficient mentioned above is used to characterize the number of times the first business is updated. The larger the quality coefficient, the more times the business data of the first business is updated, and thus the higher the importance or value of the business can be considered.

[0019] The aforementioned structural parameters are used to characterize the location of the data in the first service. In some embodiments, the structural parameters may represent information about the data table where the data of the first service resides, and / or information about the data warehouse where the data table containing the data of the first service resides.

[0020] Step 102: Calculate the set quality threshold based on the quality coefficients and structural parameters of the multiple first services.

[0021] The aforementioned quality threshold is used to filter the data for the primary business that needs to be retained. It is understandable that as the data for the primary business is continuously updated, the storage resources occupied by multiple primary businesses will increase, necessitating the deletion of some business data to reduce storage resource consumption.

[0022] In this embodiment of the invention, a set quality threshold is calculated based on the quality coefficients and structural parameters of multiple first services, such that the quality coefficients of some service data in the service data of multiple first services are less than the set quality threshold, thereby deleting service data with low importance or value by setting the quality threshold.

[0023] The quality threshold can be dynamically adjusted based on the quality coefficients and structural parameters of multiple primary services, allowing for flexible adjustment according to the current status of these services.

[0024] Step 103: Filter the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold.

[0025] If the quality coefficient of each of the aforementioned secondary services is greater than the set quality threshold, then the secondary services can be considered to be business data with high importance or value. Therefore, the business data of the secondary services needs to be retained while the other business data is deleted, thereby reducing the occupation of storage resources while retaining the business data with high importance or value.

[0026] Step 104: Retain the service data of the plurality of second services in the first storage device, and delete the service data of other services, wherein the other services are services other than the plurality of second services in the plurality of first services.

[0027] In this embodiment of the invention, the quality coefficient and structural parameters of each of a plurality of first services are obtained from a first storage device. The quality coefficient characterizes the number of times the first service is updated, and the structural parameters characterize the location of the data in the first service. Based on the quality coefficient and structural parameters of the plurality of first services, a set quality threshold is calculated. Based on the set quality threshold, the plurality of first services are filtered to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The service data of the plurality of second services in the first storage device is retained, while the service data of other services (services other than the plurality of second services in the plurality of first services) is deleted. In this way, by setting a quality threshold to filter and obtain a plurality of second services, the plurality of second services are considered to be service data with high importance or value. By retaining the service data of the plurality of second services in the first storage device and deleting the service data of other services, the use of storage resources is reduced while retaining important or valuable service data, thereby improving the data deletion effect.

[0028] In one embodiment, before obtaining the quality coefficient and structural parameters of each of the plurality of first services from the first storage device, the method further includes: Obtain the business data and historical quality coefficient of the third service, wherein the third service is one of the multiple first services obtained, and the business data of the third service includes a first identifier of the location of the business data. Generate the structured data corresponding to the third service based on the first identifier; The historical quality coefficients are updated to obtain the quality coefficients corresponding to the third service; The business data, quality coefficient, and structure data of the third service are stored in the first storage device.

[0029] It should be noted that each time the business data is updated, the business's structure data and quality coefficients must also be updated so that the quality threshold can be calculated and set based on the updated structure data and quality coefficients.

[0030] In this embodiment of the invention, structural data corresponding to the third service is generated based on the first identifier, and the historical quality coefficient is updated to obtain the quality coefficient corresponding to the third service, thereby updating the structural data and quality coefficient of the third service; after the update, the service data, quality coefficient and structural data of the third service are stored in the first storage device, so that the quality threshold can be calculated and set according to the updated structural data and quality coefficient of the third service.

[0031] In one embodiment, calculating the set quality threshold based on the quality coefficients and structural parameters of the plurality of first services includes: The m-th quality coefficient of the multiple first services arranged in ascending order is set as the set quality threshold, where m is a positive integer greater than or equal to 1 and m is less than the first product. The first product is the product of the first preset coefficient and the first quotient. The first quotient is the quotient of the first sum and the second preset coefficient. The first sum is the sum of the structural parameters of the multiple first services.

[0032] In this embodiment of the invention, the m-th quality coefficient in the ascending order of the quality coefficients of multiple first services is set as the set quality threshold, so that filtering the multiple first services through the quality threshold can retain some service data (i.e., service data of multiple second services) in the multiple first services.

[0033] Here, 'm' can be adjusted through structural parameters, which determine the amount of data from multiple second-service data sets to be retained. Specifically, 'm' is less than the first product, which is the product of a first preset coefficient and a first quotient. The first quotient is the quotient of a first sum and a second preset coefficient, and the first sum is the sum of the structural parameters of multiple first-service data sets. Thus, the larger the structural parameters, the larger 'm', and the less business data is retained. A set quality threshold is calculated based on the quality coefficients and structural parameters of multiple first-service data sets, allowing for flexible adjustment of the set quality threshold.

[0034] In one embodiment, before filtering the plurality of first services based on the set quality threshold to obtain a plurality of second services, the method further includes: Obtain the first time parameter for each first service, where the first time parameter is used to characterize the time of the last collection of service data for the first service; The filtering of the plurality of first services based on the set quality threshold to obtain a plurality of second services includes: Based on the set quality threshold and the preset time threshold, the plurality of first services are filtered to obtain a plurality of second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service.

[0035] It should be noted that, in addition to filtering multiple primary services by setting quality thresholds, multiple primary services can also be filtered by preset time thresholds. Specifically, if the time elapsed since a service last updated is greater than the preset time threshold, meaning that the service has a low update frequency, it can be considered to have low importance or value, and this part of the service can be deleted.

[0036] In this embodiment of the invention, multiple first services are filtered based on a set quality threshold and a preset time threshold to obtain multiple second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service. Thus, by filtering the multiple first services using a set quality threshold and a preset time threshold to obtain multiple second services, multiple services with high importance or value can be more accurately identified and retained, further improving the effectiveness of data deletion.

[0037] In one embodiment, the method further includes: Obtain a second identifier for each first service, the second identifier being used to characterize the type of data warehouse where the service data collection location of the first service is located; The step of retaining the service data of the multiple second services in the first storage device and deleting the service data of other services includes: Multiple storage structures are created within the first storage device, each corresponding to a data warehouse type; Based on the second identifier of the fourth service, the service data, quality coefficient, structural parameters and second identifier of the fourth service are stored in the first storage structure. The fourth service is one of the plurality of second services, and the first storage structure is one of the plurality of storage structures. The second identifier of the fourth service matches the data warehouse type corresponding to the first storage structure.

[0038] In this embodiment of the invention, multiple storage structures are created within the first storage device, each corresponding to a data warehouse type. Based on the second identifier of a fourth service, the service data, quality coefficient, structural parameters, and second identifier of the fourth service are stored in the first storage structure. The fourth service is one of the multiple second services, and the first storage structure is one of the multiple storage structures. The second identifier of the fourth service matches the data warehouse type corresponding to the first storage structure. Thus, by creating multiple storage structures on the first storage device, different types of service data are stored in storage structures of different data warehouse types, facilitating the management of different service data through different storage structures.

[0039] Specifically, in one embodiment, the method further includes: Obtain the service data, quality coefficient, structural parameters, and second identifier of the fourth service in the first storage structure; The service data, quality coefficient, structural parameters, and second identifier of the fourth service are sent to external devices through a preset interface.

[0040] In this embodiment of the invention, a preset interface is configured in the storage structure to provide external devices with the service data, quality coefficient, structural parameters, and second identifier of the fourth service. The second identifier represents the type of data warehouse where the service data is collected, and its location during transmission can be determined.

[0041] Please see Figure 2 , Figure 2 This is a structural diagram of a data processing device provided in an embodiment of the present invention, comprising: The data mining module is used to acquire business data and historical quality coefficients of multiple first-level services; A data magnetic separation module is communicatively connected to the data mining module. The data magnetic separation module is used to generate structural data corresponding to the plurality of first services and to obtain the quality coefficients corresponding to the plurality of first services. A data extraction module, which is communicatively connected to the data magnetic separation module, is used to calculate a set quality threshold based on the quality coefficients and structural parameters of the plurality of first services; filter the plurality of first services based on the set quality threshold to obtain a plurality of second services; retain the service data of the plurality of second services in the first storage device, and delete the service data of other services; The data flow module is communicatively connected to the data extraction module. The data flow module is used to read business data, quality coefficients, structural parameters, and second identifiers from multiple storage structures; and to send the business data, quality coefficients, structural parameters, and second identifiers to external devices through a preset interface.

[0042] In this embodiment of the invention, the data mining module, data magnetic separation module, data extraction module, and data flow turnover module can be used to collect, filter, and store business data, retain important or valuable business data, delete other business data, and thus improve the effect of data deletion.

[0043] Specifically, the structure of the data mining module is as follows: Figure 3 As shown. It should be noted that the data mining module is deployed on servers or big data platforms, such as service clusters like source data snapshot warehouses, integrated data warehouses, aggregation data warehouses, and application layer data warehouses, and is used to mine business data from various business operations in different data warehouses.

[0044] The source data snapshot warehouse is used to store data generated by business systems and regularly saves business data with timestamps; the integration data warehouse is used to store cleaned data that has been formatted and damaged; the aggregation data warehouse is used to store detailed data that has been classified and sorted according to business type; and the application layer data warehouse is used to store big data processed and provided to upstream business applications, such as reports, statistical or application data for artificial intelligence.

[0045] In the application layer data warehouse service cluster, the front-end agent mainly extracts business data representing the usage in the application and saves it to the cache module. In the source data snapshot warehouse, integrated data warehouse, and aggregated data warehouse service clusters, the front-end agent mainly extracts business data from various data warehouses and provides it to the data selection module for processing.

[0046] For example, the data mining module collects business data from the application-layer data warehouse service cluster. When upstream business applications use the application-layer data warehouse, it mines the accessed data to obtain data such as the database, table, corresponding business entity identifier (Identifier, ID), fields used, and usage time, and saves it as business data with a serialized data structure. i This can be expressed by the following formula: S i ={id i ,t i ,k i D i}; ; id in the formula i This indicates the business identification information in this data entry, such as the customer's ID card number, mobile phone number, etc.; t i This indicates the identification information of the data table where the business is located, such as transaction records and status tables in the financial industry; w i k represents the identification information of the data warehouse where the data table containing the business resides. i f represents the time period during which the data is used. n Other identifying information present in this data, such as bank card number, social security number, or membership card number.

[0047] After acquiring the business data, the front-end agent synchronizes the business data to the data selection module. During the initial synchronization, the front-end agent initiates a full transmission mode, transferring all serialized data structures stored on the server to the data selection module. Subsequently, it monitors the application-layer data warehouse service cluster, and when data access is requested, it transmits the requested incremental serialized data structures to the data selection module.

[0048] For example, when the data magnetic separation module sends a request message, the request message S' i This can be expressed by the following formula: S' i= {id' i ,t' i ,w' i ,k' i}; id' in the formula i This represents the business identification information within the data entry. This identification information is consistent with the application layer data warehouse. However, to ensure information and data security, it is usually recorded as an internal identifier within the system. Therefore, it is necessary to convert the identifier and change the ID. i Convert to id' i The same applies to other parameters; t' i This indicates the identifier information of the data table where the business is located. Typically, in the warehouse, this information is organized and merged according to business type or table type; w' i k' represents the identification information of the data warehouse where the data table containing the business resides. i This indicates the point in time when business data was acquired.

[0049] The data mining module obtains business data from the source data snapshot warehouse, integrated data warehouse, and aggregated data warehouse through the aforementioned request message, and saves it as a serialized data structure S'' i S'' i With S i The format is the same, then add S'' i Transmitted to the data magnetic separation module.

[0050] Furthermore, the structure of the data magnetic separation module is as follows: Figure 4 As shown. It should be noted that the data magnetic selection module is deployed in the backend server cluster and is used to carry business data from the big data platform's source data snapshot warehouse, integrated data warehouse, aggregated data warehouse, and application layer data warehouse. It also receives business data of various serialized data structures mined by the data mining module from the big data platform's source data snapshot warehouse, integrated data warehouse, aggregated data warehouse, and application layer data warehouse service cluster.

[0051] The data magnetic separation module assembles the serialized data structure obtained from the data mining module in memory. The business data generates message S' i= {id' i ,t' i ,w' i ,k' i The cumulative number of messages sent is , where id' i1 =id' i ,id' in+1 =f n , id i =id' i1 By sending messages to the data mining module, business data can be obtained from the service cluster, including the source data snapshot warehouse, integrated data warehouse, and aggregation data warehouse. .

[0052] The data magnetic separation module is used to update the structural and quality parameters of the business data. Specifically, the data magnetic separation module obtains the serialized data structure S from the data mining module. i Data, sorted by ID i Exploration archives P for creating or updating data i =(id i ,t' i ,p i ), where p i This is the quality coefficient for the data, initially set to 1. The quality parameter is updated to p each time the business data is updated. i =p i +1; structural parameters updated to After obtaining the structural and quality parameters, P is stored in key-value pairs. i and T i Data is stored in a high-speed storage module and a database, which is used for data persistence and disaster recovery backup.

[0053] Furthermore, the structure of the data extraction module is as follows: Figure 5As shown. It should be noted that the data extraction module is located in the backend server cluster and is used to load the P generated by the data magnetic separation module. i and T i By using an extraction algorithm to calculate a set quality threshold for the data, high-value data can be recovered and extracted.

[0054] In one embodiment, calculating the set quality threshold specifically involves setting the m-th quality coefficient of a plurality of first services in ascending order as the set quality threshold, where m is a positive integer greater than or equal to 1 and less than a first product. The first product is the product of a first preset coefficient and a first quotient, where the first quotient is the quotient of a first sum and a second preset coefficient, and the first sum is the sum of the structural parameters of the plurality of first services.

[0055] The first quotient q can be calculated using the following formula: ; In the formula, M max The second preset coefficient can be the maximum capacity of the cache and database in the data flow module. max If it does not exist, you can set q=1.

[0056] Furthermore, the business data is established in memory as P' i Business data in a chained storage structure arranged in order of quality parameters , can be represented as: ; ; in the formula Represents business data P' i Identification information, Represents business data P' i Acquisition time, Represents business data P' i The quality parameters. Thus, the quality threshold is determined to be... , where n is the first preset coefficient.

[0057] In some implementations, coefficients can be set for setting quality thresholds and preset time thresholds to adjust the retained business data. For example, the business data is established as T'. i Business data in a chained storage structure arranged in order of quality parameters , can be represented as: ; ; in the formula Represents business data Identification information, Represents business data Acquisition time, Represents business data T' i The quality parameters, t now t represents the current time. max Indicates the maximum duration for which data is retained. A coefficient for setting the quality threshold. This is a coefficient for a preset time threshold. In this way, it can be achieved through... and Adjust the retained business data; for example, increase the coefficient when it is necessary to minimize storage resource consumption. and If it is necessary to retain as much data as possible, reduce the coefficient. and .

[0058] P' can be i and T' i It is stored in high-speed storage in key-value pair format for use by the data flow module; at the same time, it is stored in the database for data persistence and disaster recovery backup.

[0059] Furthermore, the structure of the data flow turnover module is as follows: Figure 6 As shown. It should be noted that the data flow module is deployed in the backend server cluster and the frontend data management server. It is used to integrate and manage the extracted business data and provide access capabilities to other external systems.

[0060] Specifically, it is possible to retrieve P' from the cache. i and T' i Key-value pairs are used to establish business data. The storage structures are L OTS L DWI L DWD L DWA , This indicates the data warehouse type representing business data. Specifically, when... When representing a snapshot of the source data, the business data is stored in L. OTS ;when When integrating a data warehouse, business data is stored in L. DWI ,when When a data warehouse is being aggregated, business data is stored in L. DWD ,when When using data from the data mart, business data is stored in L. DWA .

[0061] Furthermore, LOTS L DWI L DWD L DWA Store the data in the database for data persistence and disaster recovery backup.

[0062] The data flow module loads L from the database via a web service. OTS L DWI L DWD L DWA The system displays data and utilizes web services to assist relevant personnel in managing and statistically analyzing it; it also provides pre-defined interfaces (such as API interfaces) to enable access to the data. Since L... OTS L DWI L DWD L DWA All data is transmitted through the identification information of the business within the data, so the direction of the replay data flow can be confirmed based on this.

[0063] Please see Figure 7 , Figure 7 This is a structural diagram of a data processing device provided in an embodiment of the present invention, such as... Figure 7 As shown, the data processing device 700 includes: The first acquisition module 701 is used to acquire the quality coefficient and structural parameters of each of the multiple first services from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. Calculation module 702 is used to calculate a set quality threshold based on the quality coefficients and structural parameters of the plurality of first services; The filtering module 703 is used to filter the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The first storage module 704 is used to retain the service data of the plurality of second services in the first storage device and delete the service data of other services, wherein the other services are services other than the plurality of second services among the plurality of first services.

[0064] In one embodiment, the data processing apparatus 700 further includes: The second acquisition module is used to acquire the business data and historical quality coefficient of the third service, wherein the third service is one of the multiple first services, and the business data of the third service includes a first identifier of the location of the business data. The generation module is used to generate the structured data corresponding to the third service based on the first identifier; The update module is used to update the historical quality coefficients to obtain the quality coefficients corresponding to the third service. The second storage module is used to store the business data, quality coefficients, and structure data of the third service into the first storage device.

[0065] In one embodiment, the computing module 702 includes: The calculation unit is used to set the m-th quality coefficient of a plurality of first services in ascending order as the set quality threshold, where m is a positive integer greater than or equal to 1 and m is less than the first product, the first product is the product of the first preset coefficient and the first quotient, the first quotient is the quotient of the first sum and the second preset coefficient, and the first sum is the sum of the structural parameters of the plurality of first services.

[0066] In one embodiment, the data processing apparatus 700 further includes: The third acquisition module is used to acquire the first time parameter of each first service, wherein the first time parameter is used to characterize the time of the last collection of service data of the first service. The filtering module 703 includes: A filtering unit is used to filter the plurality of first services based on the set quality threshold and the preset time threshold to obtain a plurality of second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service.

[0067] In one embodiment, the data processing apparatus 700 further includes: The fourth acquisition module is used to acquire the second identifier of each first service, wherein the second identifier is used to characterize the type of data warehouse where the business data of the first service is collected; The first storage module 704 includes: A creation unit is used to create multiple storage structures within the first storage device, each storage structure corresponding to a data warehouse type; A storage unit is used to store the business data, quality coefficient, structural parameters and second identifier of the fourth business into a first storage structure based on the second identifier of the fourth business. The fourth business is one of the plurality of second businesses, and the first storage structure is one of the plurality of storage structures. The second identifier of the fourth business matches the data warehouse type corresponding to the first storage structure.

[0068] In one embodiment, the data processing apparatus 700 further includes: The fifth acquisition module is used to acquire the service data, quality coefficient, structural parameters, and second identifier of the fourth service in the first storage structure; The sending module is used to send the service data, quality coefficient, structural parameters and second identifier of the fourth service to an external device through a preset interface.

[0069] The data processing apparatus provided in this embodiment of the invention can implement each process of each embodiment of the above data processing method, with one-to-one correspondence of technical features and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0070] It should be noted that the data processing device in the embodiments of the present invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0071] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the above-described functionality. Figure 1 The various processes of the data processing method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0072] For details, see Figure 8 As shown, this embodiment of the invention also provides an electronic device, including a bus 801, a transceiver 802, an antenna 803, a bus interface 804, a processor 805, and a memory 806.

[0073] The processor 805 is used to obtain the quality coefficient and structural parameters of each of the multiple first services from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. The processor 805 is also used to calculate a set quality threshold based on the quality coefficients and structural parameters of the plurality of first services. The processor 805 is further configured to filter the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The processor 805 is further configured to retain the service data of the plurality of second services in the first storage device and delete the service data of other services, wherein the other services are services other than the plurality of second services among the plurality of first services.

[0074] In one embodiment, the transceiver 802 is used to acquire service data and historical quality coefficients of a third service, wherein the third service is one of a plurality of first services, and the service data of the third service includes a first identifier of the location of the service data. The processor 805 is further configured to generate structured data corresponding to the third service based on the first identifier; The processor 805 is also used to update the historical quality coefficient to obtain the quality coefficient corresponding to the third service. The processor 805 is also used to store the service data, quality coefficient and structure data of the third service into the first storage device.

[0075] In one embodiment, calculating the set quality threshold based on the quality coefficients and structural parameters of the plurality of first services includes: The m-th quality coefficient of the multiple first services arranged in ascending order is set as the set quality threshold, where m is a positive integer greater than or equal to 1 and m is less than the first product. The first product is the product of the first preset coefficient and the first quotient. The first quotient is the quotient of the first sum and the second preset coefficient. The first sum is the sum of the structural parameters of the multiple first services.

[0076] In one embodiment, the transceiver 802 is further configured to acquire a first time parameter for each first service, wherein the first time parameter is used to characterize the time when the service data of the first service was last collected. The filtering of the plurality of first services based on the set quality threshold to obtain a plurality of second services includes: Based on the set quality threshold and the preset time threshold, the plurality of first services are filtered to obtain a plurality of second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service.

[0077] In one embodiment, the transceiver 802 is further configured to acquire a second identifier for each first service, the second identifier being used to characterize the type of data warehouse where the service data of the first service is collected; The step of retaining the service data of the multiple second services in the first storage device and deleting the service data of other services includes: Multiple storage structures are created within the first storage device, each corresponding to a data warehouse type; Based on the second identifier of the fourth service, the service data, quality coefficient, structural parameters and second identifier of the fourth service are stored in the first storage structure. The fourth service is one of the plurality of second services, and the first storage structure is one of the plurality of storage structures. The second identifier of the fourth service matches the data warehouse type corresponding to the first storage structure.

[0078] In one embodiment, the transceiver 802 is further configured to acquire the service data, quality coefficient, structural parameters, and second identifier of the fourth service in the first storage structure; The transceiver 802 is also used to send the service data, quality coefficient, structural parameters and second identifier of the fourth service to an external device through a preset interface.

[0079] exist Figure 8 In this document, a bus architecture (represented by bus 801) is used. Bus 801 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 805 and memory represented by memory 806. Bus 801 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 804 provides an interface between bus 801 and transceiver 802. Transceiver 802 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 805 is transmitted over a wireless medium via antenna 803, which further receives data and transmits data to processor 805.

[0080] The processor 805 manages the bus 801 and handles general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 806 can be used to store data used by the processor 805 during operation.

[0081] Optionally, the processor 805 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a graphics processing unit (GPU).

[0082] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the above-described functions. Figure 1 The various processes corresponding to the data processing method embodiments achieve the same technical effect, and will not be described again here to avoid repetition. The computer-readable storage medium mentioned includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0083] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The various processes corresponding to the data processing method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.

[0084] In the embodiments of this invention, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing eight possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0085] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0087] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A data processing method, characterized in that, include: The quality coefficient and structural parameters of each of the multiple first services are obtained from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. Based on the quality coefficients and structural parameters of the multiple first services, a set quality threshold is calculated. Based on the set quality threshold, the plurality of first services are filtered to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The service data of the plurality of second services in the first storage device is retained, and the service data of other services are deleted. The other services are services other than the plurality of second services in the plurality of first services. The calculation of the set quality threshold based on the quality coefficients and structural parameters of the multiple first services includes: The m-th quality coefficient in the ascending order of the quality coefficients of multiple first services is set as the set quality threshold, where m is a positive integer greater than or equal to 1 and m is less than the first product. The first product is the product of the first preset coefficient and the first quotient. The first quotient is the quotient of the first sum and the second preset coefficient. The first sum is the sum of the structural parameters of the multiple first services. Before filtering the plurality of first services based on the set quality threshold to obtain a plurality of second services, the method further includes: Obtain the first time parameter for each first service, where the first time parameter is used to characterize the time of the last collection of service data for the first service; The filtering of the plurality of first services based on the set quality threshold to obtain a plurality of second services includes: Based on the set quality threshold and the preset time threshold, the plurality of first services are filtered to obtain a plurality of second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service.

2. The method as described in claim 1, characterized in that, Before obtaining the quality coefficient and structural parameters of each of the plurality of first services from the first storage device, the method further includes: Obtain the business data and historical quality coefficient of the third service, wherein the third service is one of the multiple first services obtained, and the business data of the third service includes a first identifier of the location of the business data. Generate the structured data corresponding to the third service based on the first identifier; The historical quality coefficients are updated to obtain the quality coefficients corresponding to the third service; The business data, quality coefficient, and structure data of the third service are stored in the first storage device.

3. The method as described in claim 1 or 2, characterized in that, The method further includes: Obtain a second identifier for each first service, the second identifier being used to characterize the type of data warehouse where the service data collection location of the first service is located; The step of retaining the service data of the multiple second services in the first storage device and deleting the service data of other services includes: Multiple storage structures are created within the first storage device, each corresponding to a data warehouse type; Based on the second identifier of the fourth service, the service data, quality coefficient, structural parameters and second identifier of the fourth service are stored in the first storage structure. The fourth service is one of the plurality of second services, and the first storage structure is one of the plurality of storage structures. The second identifier of the fourth service matches the data warehouse type corresponding to the first storage structure.

4. The method as described in claim 3, characterized in that, The method further includes: Obtain the service data, quality coefficient, structural parameters, and second identifier of the fourth service in the first storage structure; The service data, quality coefficient, structural parameters, and second identifier of the fourth service are sent to external devices through a preset interface.

5. A data processing apparatus, characterized in that, include: The data mining module is used to acquire business data and historical quality coefficients of multiple first-level services; A data magnetic separation module is communicatively connected to the data mining module. The data magnetic separation module is used to generate structural data corresponding to the plurality of first services and to obtain the quality coefficients corresponding to the plurality of first services. A data extraction module, which is communicatively connected to the data magnetic separation module, is used to calculate a set quality threshold based on the quality coefficients and structural parameters of the multiple first services. Based on the set quality threshold, the multiple first services are filtered to obtain multiple second services; Retain the service data of the multiple second services in the first storage device, and delete the service data of other services; A data flow module is communicatively connected to the data extraction module. The data flow module is used to read business data, quality coefficients, structural parameters, and second identifiers from multiple storage structures; and to send the business data, quality coefficients, structural parameters, and second identifiers to external devices through a preset interface. Specifically, the data extraction module is used for: The m-th quality coefficient in the ascending order of the quality coefficients of multiple first services is set as the set quality threshold, where m is a positive integer greater than or equal to 1 and m is less than the first product. The first product is the product of the first preset coefficient and the first quotient. The first quotient is the quotient of the first sum and the second preset coefficient. The first sum is the sum of the structural parameters of the multiple first services. Obtain the first time parameter for each first service, where the first time parameter is used to characterize the time of the last collection of service data for the first service; Based on the set quality threshold and the preset time threshold, the plurality of first services are filtered to obtain a plurality of second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service.

6. A data processing apparatus, characterized in that, include: The first acquisition module is used to acquire the quality coefficient and structural parameters of each of the multiple first services from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. The calculation module is used to calculate a set quality threshold based on the quality coefficients and structural parameters of the multiple first services; A filtering module is used to filter the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. A first storage module is used to retain the service data of the plurality of second services in the first storage device and delete the service data of other services, wherein the other services are services other than the plurality of second services among the plurality of first services; The computing module includes: The calculation unit is used to set the m-th quality coefficient of a plurality of first services in ascending order as the set quality threshold, where m is a positive integer greater than or equal to 1 and m is less than the first product, the first product is the product of the first preset coefficient and the first quotient, the first quotient is the quotient of the first sum and the second preset coefficient, and the first sum is the sum of the structural parameters of the plurality of first services. The data processing device further includes: The third acquisition module is used to acquire the first time parameter of each first service, wherein the first time parameter is used to characterize the time of the last collection of service data of the first service. The filtering module includes: A filtering unit is used to filter the plurality of first services based on the set quality threshold and the preset time threshold to obtain a plurality of second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service.

7. An electronic device, characterized in that, Including transceivers and processors, The processor is configured to obtain the quality coefficient and structural parameters of each of the multiple first services from the first storage device. The quality coefficient is used to characterize the number of times the first service is updated, and the structural parameters are used to characterize the location of the data in the first service. The processor is also configured to calculate a set quality threshold based on the quality coefficients and structural parameters of the plurality of first services. The processor is further configured to filter the plurality of first services based on the set quality threshold to obtain a plurality of second services, wherein the quality coefficient of each of the plurality of second services is greater than the set quality threshold. The processor is further configured to retain the service data of the plurality of second services in the first storage device and delete the service data of other services, wherein the other services are services other than the plurality of second services among the plurality of first services; The calculation of the set quality threshold based on the quality coefficients and structural parameters of the multiple first services includes: The m-th quality coefficient in the ascending order of the quality coefficients of multiple first services is set as the set quality threshold, where m is a positive integer greater than or equal to 1 and m is less than the first product. The first product is the product of the first preset coefficient and the first quotient. The first quotient is the quotient of the first sum and the second preset coefficient. The first sum is the sum of the structural parameters of the multiple first services. The transceiver is also used to acquire a first time parameter for each first service, wherein the first time parameter is used to characterize the time when the service data of the first service was last collected; The filtering of the plurality of first services based on the set quality threshold to obtain a plurality of second services includes: Based on the set quality threshold and the preset time threshold, the plurality of first services are filtered to obtain a plurality of second services. The quality coefficient of each second service is greater than the set quality threshold, and the first difference of each second service is less than the preset time threshold. Each second service includes a fourth service, and the first difference of the fourth service is the difference between the current time parameter and the first time parameter of the fourth service.

8. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the data processing method as described in any one of claims 1 to 4.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the data processing method as described in any one of claims 1 to 4.

10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the data processing method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Cache data deleting method and server

    CN104715020A

  • Data deletion method, terminal, and computer-readable storage medium

    CN109062964A