Method and device for determining space occupation amount, electronic equipment and storage medium
By downsampling the data in the storage space and determining the space occupancy of the target data type, the problems of traditional methods such as being time-consuming and labor-intensive and requiring high computing resources are solved, and efficient space occupancy determination is achieved.
Patent Information
- Application Number
- CN202510263084.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional full-scan methods are time-consuming and labor-intensive when determining the space occupancy of data in data storage systems, and require high computing resources, making it difficult to meet the needs of real-time and large-scale data processing.
By performing first downsampling and second downsampling on the data in the storage space, third sampling data of the target data type is determined, and the space occupied by the data of the target data type in the storage space is estimated based on the third sampling data.
This significantly reduces the amount of data that needs to be processed, reduces the consumption of computing resources, and improves the efficiency of determining data space occupancy.
Smart Images

Figure CN120704592A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to computer technology, and in particular to a method, device, electronic device, and storage medium for determining space occupancy. Background Art
[0002] With the rapid growth of data volumes in data storage systems, accurately and efficiently determining the space usage of different types of data has become crucial. Traditional full-data scanning methods are not only time-consuming and labor-intensive, but also extremely demanding on computing resources, making them difficult to meet the demands of real-time and large-scale data processing. Summary of the Invention
[0003] The embodiments of the present application provide a method, device, electronic device, and storage medium for determining space occupancy, which can improve the efficiency of determining space occupancy.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides a method for determining a space occupation amount, the method comprising:
[0006] Performing a first downsampling on the data in the storage space to obtain first sampled data;
[0007] Based on the first data type of the first sampled data, performing a second downsampling on the first sampled data to obtain second sampled data;
[0008] determining a target data type from the first data type, and determining third sampled data belonging to the target data type from the second sampled data;
[0009] Based on a first space occupied by the third sampled data in the storage space, a second space occupied by data of the target data type in the storage space is determined.
[0010] An embodiment of the present application provides a device for determining space occupancy, including:
[0011] a sampling module, configured to perform a first downsampling on the data in the storage space to obtain first sampled data;
[0012] The sampling module is configured to perform a second downsampling on the first sampled data based on the first data type of the first sampled data to obtain second sampled data;
[0013] an acquisition module, configured to determine a target data type from the first data type, and determine third sampled data belonging to the target data type from the second sampled data;
[0014] The determining module is configured to determine a second space occupancy of data of the target data type in the storage space based on a first space occupancy of the third sampled data in the storage space.
[0015] An embodiment of the present application provides an electronic device, the electronic device comprising:
[0016] a memory for storing computer-executable instructions or computer programs;
[0017] The processor is used to implement the method for determining the space occupancy provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0018] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the method for determining the space occupancy provided in an embodiment of the present application when executed by a processor.
[0019] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the method for determining the space occupancy provided in the embodiment of the present application is implemented.
[0020] The embodiments of the present application have the following beneficial effects:
[0021] By performing first and second downsampling on the data in the storage space, the amount of data to be processed is significantly reduced, thereby lowering computing resource consumption. Furthermore, the second sampled data is sampled data obtained by downsampling the first data type of the first sampled data. Third sampled data belonging to the target data type is determined from the second sampled data, and the second space occupancy of the target data type in the entire storage space is inferred based on the first space occupancy of the third sampled data. This allows the space occupancy of all data of a certain data type to be determined based on the space occupancy of a small amount of data of that data type, thereby improving the efficiency of determining the space occupancy of data in the storage space. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Schematic diagram of the system for determining space occupancy provided by an embodiment of the present application;
[0023] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0024] Figure 3 This is a flow chart of the method for determining the space occupancy provided in the embodiment of the present application. Figure 1 ;
[0025] Figure 4This is a flow chart of the method for determining the space occupancy provided in the embodiment of the present application. Figure 2 ;
[0026] Figure 5 This is a flow chart of the method for determining the space occupancy provided in the embodiment of the present application. Figure 3 ;
[0027] Figure 6 This is a flowchart of traversing data in a database provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0029] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0030] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0031] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0032] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0033] In the embodiments of this application, when collecting and processing relevant data in practical applications, it should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope authorized by laws, regulations and the personal information subject.
[0034] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0035] 1) Uniform distribution: In probability theory and statistics, the uniform distribution, also known as the rectangular distribution, is a symmetric probability distribution where the distribution probabilities in intervals of the same length are equally likely. The uniform distribution function is a < x < b, where f(x) represents the probability that the variable x appears, and a and b represent the range where the variable x appears.
[0036] 2) Normal distribution: Also known as the "normal distribution", also called the Gaussian distribution, which is symmetric. Data that follows a normal distribution obeys a probability distribution with a location parameter of μ and a scale parameter of σ. The normal distribution function f(x; μ, σ ) is as follows:
[0037] 3) Downsampling: To solve the large amount of calculation problems brought by a large amount of data, a subset of the data set is often obtained according to the distribution characteristics of the data for calculation, and the characteristics of the entire data set are inferred from the calculation results of the subset.
[0038] 4) Data Structure Storage System (Remote Dictionary Server, Redis): An open-source, memory-based, high-performance distributed database. Its read and write performance supports 80,000 - 100,000 queries per second (Queries-per-second, qps) on a single machine, supports master-slave synchronization and failover, and the commands in Redis are executed in a single thread.
[0039] 5) Key: The basic data unit in Redis, used to uniquely identify a value (Value). The key is of string type and can contain any binary data.
[0040] 6) Slot: The unit in Redis for storing cached Keys. For example, there are 16,384 slots built into a Redis cluster. When each data is stored in Redis, it needs to be placed in the specified slot according to the hash value of the Key.
[0041] 7) Data backup (Redis Database Backup, RDB) log: A Redis data persistence technology. RDB log is a snapshot of the data set that Redis generates periodically and saves on disk. It can be used for Redis data backup and recovery.
[0042] 8) Prefix: It is the prefix of the key, which refers to the part of characters at the beginning of the string, and is often used to identify the category, type or range of the string.
[0043] 9) Suffix: It is the suffix of the key, which refers to the part of characters at the end of the string, often used to identify the specific attributes, status or additional information of the string.
[0044] The embodiments of the present application provide a method, device, electronic device, and storage medium for determining space occupancy, which can determine the efficiency of the space occupancy of data in a storage space. In the method for determining space occupancy provided in the embodiments of the present application, first, a first downsampling is performed on the data in the storage space to obtain first sampled data; then, based on the first data type of the first sampled data, a second downsampling is performed on the first sampled data to obtain second sampled data; then, a target data type is determined from the first data type, and third sampled data belonging to the target data type is determined from the second sampled data; finally, based on the first space occupancy of the third sampled data in the storage space, a second space occupancy of the data belonging to the target data type in the storage space is determined.
[0045] The following describes an exemplary application of a device for determining the amount of space occupied provided in an embodiment of the present application. The device for determining the amount of space occupied is an electronic device for implementing a method for determining the amount of space occupied. The electronic device provided in an embodiment of the present application can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and vehicle-mounted terminals, and can also be implemented as a server. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application. Below, an exemplary application of the device for determining the amount of space occupied when it is implemented as a terminal or a server will be described.
[0046] See also Figure 1 , Figure 1 It is an architectural diagram of the space occupancy determination system provided in an embodiment of the present application. In order to perform the space occupancy determination operation, a space occupancy determination application can be provided. For example, the space occupancy determination application can be an application dedicated to the determination of the space occupancy, or it can be a functional module in other applications (such as a space occupancy determination module in a database application, etc.). The space occupancy determination system 100 in an embodiment of the present application includes at least a terminal 400, a network 300 and a server 200, wherein the server 200 is a server of the space occupancy determination application. The server 200 can constitute the space occupancy determination device of the embodiment of the present application, that is, the space occupancy determination method of the embodiment of the present application is implemented through the server 200. The terminal 400 is connected to the server 200 via the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.
[0047] See also Figure 1 , a user can perform interactive operations on the client side of the space occupancy determination application through terminal 400. Such interactive operations may include, for example, starting to traverse the storage space or starting to determine the space occupancy. After receiving the user's interactive operation, the client side sends a space occupancy determination request to server 200 via network 300. After receiving the space occupancy determination request, server 200, in response to the space occupancy determination request sent by the terminal, performs a first downsampling on the data in the storage space to obtain first sampled data. Based on the first data type of the first sampled data, server 200 performs a second downsampling on the first sampled data to obtain second sampled data. Server 200 determines a target data type from the first data type and determines third sampled data belonging to the target data type from the second sampled data. Based on the first space occupancy of the third sampled data in the storage space, server 200 determines a second space occupancy of the data belonging to the target data type in the storage space. After determining the second space occupancy of the data of the target data type, server 200 may also send the second space occupancy to terminal 400. Terminal 400 displays the second space occupancy of the data of the target data type on the current interface.
[0048] In some embodiments, the terminal 400 may also perform the method for determining the space occupancy of the embodiment of the present application. That is, after the user performs an interactive operation on the client of the space occupancy determination application through the terminal 400, the terminal 400, in response to the interactive operation, performs a first downsampling on the data in the storage space to obtain first sampled data; the terminal 400 performs a second downsampling on the first sampled data based on the first data type of the first sampled data to obtain second sampled data; the terminal 400 determines the target data type from the first data type and determines third sampled data belonging to the target data type from the second sampled data; the terminal 400 determines the second space occupancy of the data belonging to the target data type in the storage space based on the first space occupancy of the third sampled data in the storage space. The terminal 400 displays the second space occupancy of the data of the target data type on the current interface.
[0049] In a scenario where a big data analysis platform needs to regularly count the storage occupancy of different types of data in a database (Redis), a first downsampling is performed on part or all of the data in the database to obtain first sampling data. For example, one piece of data is selected as the first sampling data for every 100 pieces of data. Based on the first data type (address data, user behavior data, transaction data, etc.) of the first sampling data, a second downsampling is performed on the first sampling data of each first data type to obtain second sampling data. The second sampling data of the target data type (such as address data) is selected from the second sampling data as the third sampling data. The first space occupancy of the third sampling data in the database is calculated. Based on the first space occupancy of the third sampling data in the storage space, the second space occupancy of all address data in the storage space is determined. Through the above method, the big data analysis platform can efficiently estimate the storage occupancy of different types of data, thereby optimizing data storage strategies and query performance, and improving the operating efficiency of the overall system.
[0050] In the data management scenario of cloud storage services, it is necessary to regularly evaluate the occupancy of different types of files in the storage space in order to optimize storage resources and billing strategies. The storage space contains multiple types of files, such as pictures, videos, documents, etc. A first downsampling is performed on all files in the storage space. For example, a file is selected as the first sampling data with a probability of 10%. Based on the file type (picture, video, document, etc.) of the first sampling data, a second downsampling is performed on the first sampling data of each type to obtain the second sampling data. Files belonging to the target data type (such as pictures) are extracted from the second sampling data as the third sampling data. The first space occupancy of the third sampling data in the storage space is calculated, and based on the first space occupancy, the second space occupancy of all image files in the storage space is determined. Through the above method, the cloud storage service provider can efficiently estimate the space occupancy of different types of files, thereby optimizing storage resource allocation and billing strategies.
[0051] See also Figure 2 , Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 2 The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0052] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0053] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0054] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0055] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0056] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0057] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0058] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);
[0059] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0060] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0061] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 A device 455 for determining the amount of space occupied stored in memory 450 is shown. This device 455 may be software in the form of a program or plug-in, and includes the following software modules: a sampling module 4551, an acquisition module 4552, and a determination module 4553. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0062] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the method for determining the space occupancy provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0063] The following describes the method for determining the space usage provided by an embodiment of the present application. As previously mentioned, the electronic device implementing the method for determining the space usage provided by an embodiment of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.
[0064] It should be noted that, based on their understanding of the following text, those skilled in the art can apply the method for determining the space occupancy provided in the embodiments of the present application to a variety of scenarios, such as: data management scenarios in cloud storage services, data statistics scenarios in big data analysis platforms, resource optimization scenarios in enterprise-level data warehouses, device data management scenarios in Internet of Things platforms, content management scenarios in social media platforms, etc.
[0065] See also Figure 3 , Figure 3 This is a flow chart of the method for determining the space occupancy provided in the embodiment of the present application. Figure 1 , the following will be combined Figure 3 The steps shown are explained as Figure 3 As shown, the method for determining the space occupancy is described by taking the server as an example. The method includes the following steps 101 to 104:
[0066] In step 101, a first downsampling is performed on the data in the storage space to obtain first sampled data.
[0067] Here, storage space refers to a physical or virtual area for storing and managing data, which can be a hard disk, a distributed file system, a cloud storage service, etc. Exemplarily, in an embodiment of the present application, the storage space is a database (Redis). Data refers to information stored in the storage space, which can be any form of content, such as files, records, logs, images, videos, etc. Data can be structured data, that is, data with a fixed format, such as rows and columns in a database table. Data can also be unstructured data, that is, data without a fixed format, such as text files, pictures, videos, etc. Data can also be semi-structured data, that is, data between structured and unstructured, such as JSON, XML files, etc. In an embodiment of the present application, taking the storage space as a database (Redis) as an example, data refers to the keys (Key) stored in each slot in Redis.
[0068] The data in the storage space can be all the data in the storage space, or part of the data in the storage space. If the data in the storage space is part of the data extracted from the storage space through the screening condition, use a pre-set regular expression (regular formula) as the screening condition, traverse the storage space, and obtain data that meets the regular formula. Perform a first downsampling on the data that meets the regular expression to obtain the first sampled data. For example, when it is necessary to evaluate the space occupancy of all address data in Redis, the regular expression can be set to "prefix is address (address)", traverse all keys in Redis, check whether the prefix of the key meets the above regular expression, and extract all keys whose prefix is address as the data in the storage space.
[0069] The first downsampling may refer to the process of randomly or according to rules selecting a portion of data from the data in the storage space as the first sampling data, with the aim of reducing the amount of data that needs to be processed, thereby improving the efficiency of subsequent processing. For example, 1 million keys are randomly selected from 10 million keys as the first sampling data. The embodiment of the present application does not limit the downsampling method of the first downsampling, for example, it may be random downsampling, interval downsampling, or condition-based downsampling (selecting data items based on timestamps, data types, etc.). The following describes the first downsampling process of random downsampling and interval downsampling respectively.
[0070] In some embodiments, performing a first downsampling on the data in the storage space to obtain first sampling data can be achieved in the following manner: first, for each data in the storage space, generating a first random number corresponding to the data; then, determining the data corresponding to the first random number that is less than or equal to the first sampling probability as the first sampling data.
[0071] Here, the first downsampling may be random downsampling. Random downsampling is to sample multiple data in the storage space according to a set first sampling probability to obtain multiple first sampling data. The first sampling probability is a preset value that represents the probability of each data being selected as the first sampling data. For example, the first sampling probability is set to 0.1, indicating that each data has a 10% probability of being sampled as the first sampling data. The probability distribution obeyed by random downsampling satisfies the following formula (1).
[0072]
[0073] in, is the first sampling probability, and f(x)1 is the probability distribution of random downsampling. When a=1, the first sampling probability is 1, and all the data in the storage space will be sampled as the first sampling data. When a>1, the data in the storage space will be sampled according to Sampling is performed with the first sampling probability of .
[0074] For each data, a first random number corresponding to the data can be randomly generated, and the value range of the first random number is 0-1. For example, the first random number can be generated using a random number generation function provided by a programming language. The first random number is compared with the first sampling probability. If the first random number is less than or equal to the set first sampling probability, the data is determined to be the first sampling data. Alternatively, if the first random number is greater than the set first sampling probability, the data is skipped and a first random number is randomly generated for the next data. For example, assuming that there are 1000 data in the storage space and the first sampling probability is set to 10% (i.e. 0.1). For each data, a first random number between 0 and 1 is generated. If the generated first random number is less than or equal to 0.1, the data is selected as the first sampling data; otherwise, the data is skipped. In the end, approximately 100 (1000×10%) first sampling data are obtained.
[0075] The embodiment of the present application processes the data in the storage space by randomly downsampling, so that the first sampled data obtained by sampling is evenly distributed in the storage space, thereby significantly reducing the amount of data to be processed without affecting the overall data characteristics, and improving the efficiency of subsequent analysis and processing.
[0076] In some embodiments, the first downsampling refers to sampling at a preset sampling interval. The first downsampling of the data in the storage space to obtain the first sampled data can also be achieved by: determining, for each piece of data in the storage space, the location information of the data within the plurality of data in the storage space; performing a preset operation on the sampling interval based on the location information; and determining the data as the first sampled data if the result of the preset operation is the same as a set first value.
[0077] Here, the first downsampling process is interval downsampling. Interval downsampling is to sample multiple data in the storage space according to a set sampling interval to obtain multiple first sampled data. That is, the data in the storage space can be sampled as the first sampled data based on the set sampling interval. The sampling interval is a preset value. Every sampling interval of data, one data is selected as the first sampled data. For example, setting the sampling interval to 10 means that every 10 data is selected as the first sampled data. The probability distribution obeyed by interval downsampling satisfies the following formula (2).
[0078] f(x)2=1,if x%b=1 Formula (2);
[0079] Where x is the xth data, b is the sampling interval, and f(x)2 is the probability distribution obeyed by the interval downsampling.
[0080] Position information refers to the unique identifier or sequential number of each data in the multiple data in the storage space, which is used to describe the location of the data for indexing, sorting, remaindering and other operations. The position information can be an index or a number, for example, the position information can be an integer. A unique index is assigned to each of the multiple data in the storage space in sequence as the position information. For example, the position information of the 10 data in the storage space is 0-10 respectively. The position information of the first data is 1. The preset operation on the sampling interval can be achieved in the following way: take the remainder of the sampling interval to obtain the remainder, and use the remainder as the result of the preset operation.
[0081] For example, the number of data in the storage space is 1000, the sampling interval is 10, and the first value is a positive integer 1. For each data item, the sampling interval is modulo the data item based on the data item's location information. When the remainder is 1, the data item is determined as the first sampled data item. Thus, the first data item is the first sampled data item, and the eleventh data item is the first sampled data item. Ultimately, approximately 100 (1000 / 10) first sampled data items are obtained.
[0082] The embodiment of the present application processes the data in the storage space by means of interval downsampling, so that the first sampled data obtained by sampling is evenly distributed in the storage space, thereby significantly reducing the amount of data to be processed without affecting the overall data characteristics, and improving the efficiency of subsequent analysis and processing.
[0083] In step 102 , based on the first data type of the first sampled data, the first sampled data is subjected to a second downsampling to obtain second sampled data.
[0084] Here, data type refers to the category or classification of data. In the embodiments of this application, using data as a key as an example, data type refers to the prefix, suffix, or other identifier of the key. For example, a data key with the prefix "address": address.data_0, has an address data type; or a data key with the suffix "name": data_0.name, has a name data type. Each first sampled data item corresponds to a first data type, and each first data type corresponds to one or more first sampled data items. Based on the first data type of each first sampled data item, a second downsampling is performed on the multiple first sampled data items to obtain multiple second sampled data items. The second downsampling is based on the first data type and further downsampling is performed on the first sampled data items to reduce the data volume while maintaining the representativeness of data of different data types. The second downsampling dynamically adjusts the sampling probability based on the number of sampled data items of the same data type to ensure a balanced distribution of data of different data types. The embodiments of this application do not limit the downsampling method of the second downsampling; for example, it can be piecewise linear probability downsampling or exponential probability downsampling. The second downsampling process of the piecewise linear probability downsampling and exponential probability downsampling methods are described below.
[0085] In some embodiments, the first sampled data are arranged in order based on the location information of the first sampled data in the storage space. In step 102, based on the first data type of the first sampled data, performing a second downsampling on the first sampled data to obtain the second sampled data can be achieved by: first, traversing the first sampled data one by one; for the i-th first sampled data traversed, determining the second data type of the i-th first sampled data, and determining a first number of fourth sampled data belonging to the second data type, where the fourth sampled data is the first sampled data sampled as the second sampled data among the first i-1 first sampled data, and i is a positive integer that increases successively; then, based on the first number, the first number threshold, and the second number threshold, determining a second sampling probability that the i-th first sampled data is sampled as the second sampled data; finally, generating a second random number corresponding to the i-th first sampled data, and determining the i-th first sampled data as the second sampled data if the second random number is less than or equal to the second sampling probability.
[0086] Here, traversal starts from the first first sampling data among multiple first sampling data, and it is determined whether each first sampling data is sampled as the second sampling data. For the traversed i-th first sampling data, the first data type of the i-th first sampling data is determined to be the second data type. The first sampling data belonging to the second data type and sampled as the second sampling data in the first i-1 first sampling data are counted, and these first sampling data are determined as the fourth sampling data, and the first quantity of the fourth sampling data is obtained. Based on the first quantity, the first quantity threshold and the second quantity threshold, the second sampling probability that the i-th first sampling data is sampled as the second sampling data is determined, and the second sampling probability is negatively correlated with the first quantity. The embodiment of the present application does not limit the values of the first quantity threshold and the second quantity threshold, and it is only necessary to satisfy that the second quantity threshold is greater than the first quantity threshold.
[0087] For the i-th first sampled data, a second random number can be randomly generated, with a value range of 0-1. For example, the second random number can be generated using a random number generation function provided by a programming language. The second random number is compared with the second sampling probability. If the second random number is less than or equal to the second sampling probability, the i-th first sampled data is determined as the second sampled data. Alternatively, if the second random number is greater than the second sampling probability, the i-th first sampled data is skipped, and the above steps are repeated for the i+1-th first sampled data.
[0088] For example, i=5. Assume that the first data type of the first first sampled data is an address and has been sampled as the second sampled data. The first data type of the second first sampled data is a name and has been sampled as the second sampled data. The first data type of the third first sampled data is an address and has been sampled as the second sampled data. The first data type of the fourth first sampled data is an address and has not been sampled as the second sampled data. The first data type of the fifth first sampled data is an address, and the second data type is an address. The first number of fourth sampled data of the address type is counted among the first four first sampled data: the first first sampled data and the third first sampled data. The first number of fourth sampled data is 2. A second sampling probability that the fifth first sampled data is sampled as the second sampled data can be determined based on the first number, the first number threshold, and the second number threshold.
[0089] In an embodiment of the present application, the first sampled data is subjected to a second downsampling based on the data type. The second sampling probability in the second downsampling process decreases as the first number increases. This can maintain the representativeness of data of different data types while minimizing the probability of the scanned data participating in statistics when the amount of data of the same data type is large to ensure execution efficiency, further reduce the amount of data that needs to be processed, and improve the efficiency of subsequent determination of space occupancy.
[0090] In some embodiments, determining the second sampling probability that the i-th first sampling data is sampled as the second sampling data based on the first quantity, the first quantity threshold, and the second quantity threshold can be achieved in the following manner: first, when the first quantity is less than or equal to the first quantity threshold, determining a set second value as the second sampling probability that the i-th first sampling data is sampled as the second sampling data; when the first quantity is greater than the first quantity threshold and less than or equal to the second quantity threshold, dividing the data based on the first quantity threshold and the second quantity threshold to obtain multiple first value intervals, wherein the second quantity threshold is greater than the first quantity threshold; then, determining the first value interval including the first quantity as the second value interval, and determining the reciprocal of the maximum endpoint value of the second value interval as the second sampling probability that the i-th first sampling data is sampled as the second sampling data; finally, when the first quantity is greater than the second quantity threshold, determining a set third value as the second sampling probability that the i-th first sampling data is sampled as the second sampling data, wherein the second value is greater than the reciprocal of the maximum endpoint value, and the reciprocal of the maximum endpoint value is greater than the third value.
[0091] Here, a second numerical interval is determined from multiple first numerical intervals, and the inverse of the maximum endpoint value of the second numerical interval is determined as the second sampling probability, wherein the multiple first numerical intervals are obtained by partitioning the numerical interval composed of the first quantity threshold and the second quantity threshold, the second numerical interval is the first numerical interval in which the first quantity is located, and the second quantity threshold is greater than the first quantity threshold.
[0092] In an embodiment of the present application, the second downsampling may be piecewise linear probability downsampling. Piecewise linear probability downsampling is a method for dynamically adjusting the second sampling probability based on the first amount of sampled fourth sampled data. Piecewise linear probability downsampling divides the sampling process into multiple stages, with the second sampling probability in each stage varying linearly to ensure a balanced distribution of second sampled data of different data types.
[0093] The probability distribution of the piecewise linear probability downsampling satisfies the following formula (3).
[0094]
[0095] Wherein, f(x)3 is the probability distribution obeyed by the piecewise linear probability downsampling, x is the first quantity, d is the second quantity threshold, 0 is the first quantity threshold, and c is the endpoint value of the first numerical interval.
[0096] Exemplarily, the first quantity threshold is 0, and the second value is set to 1. When the first quantity is less than or equal to 0, the second sampling probability that the i-th first sampled data is sampled as the second sampled data is 1. That is, for each first data type traversed, the first first sampled data will definitely be sampled as the second sampled data. The second quantity threshold is the quantity threshold of the second sampled data belonging to the second data type, and the third value is set to 0. That is, when the first quantity of the fourth sampled data of the sampled second data type is greater than the second quantity threshold d (d is an integer greater than 0), the second sampling probability is 0, and each subsequent traversed first sampled data of the second data type will not be sampled as the second sampled data. The numerical interval formed by the first quantity threshold and the second quantity threshold is 0-d. The numerical interval 0-d is partitioned to obtain at least one first numerical interval. For example, partitioning obtains a first numerical interval (0, c] and a second numerical interval (c, d]. Assuming c is 5, d is 20, and the first number is 2, then the second numerical interval is the first numerical interval (0, 5] including the first number, and the reciprocal of the maximum endpoint value c of the second numerical interval (0, 5], 1 / c, is the second sampling probability 1 / 5.
[0097] The embodiment of the present application processes the first sampling data through piecewise linear probability downsampling to obtain the second sampling data, thereby ensuring that there is first sampling data that is sampled as the second sampling data for each data type, and when the amount of data of the same data type is large, the probability of the scanned data participating in the statistics is reduced as much as possible to ensure execution efficiency, further reduce the amount of data that needs to be processed, and improve the efficiency of subsequent determination of space occupancy.
[0098] In some embodiments, a second sampling probability that the i-th first sampled data is sampled as the second sampled data may also be determined based on the first quantity and the second quantity threshold. The second sampling probability that the i-th first sampled data is sampled as the second sampled data may be determined based on the first quantity and the second quantity threshold by: when the first quantity is less than or equal to the second quantity threshold, performing a calculation on a set fourth value and the first quantity to obtain a fifth value; performing exponential processing on the fifth value to obtain a sixth value, and determining the ratio of the fourth value to the sixth value as the second sampling probability; and when the first quantity is greater than the second quantity threshold, determining the set third value as the second sampling probability.
[0099] Here, the second downsampling may be exponential probability downsampling. Exponential probability downsampling is a method for dynamically adjusting the second sampling probability based on the first amount of sampled fourth sampled data, wherein the second sampling probability decreases according to an exponential function. The probability distribution obeyed by the exponential probability downsampling satisfies the following formula (4).
[0100]
[0101] Where f(x)4 is the probability distribution obeyed by the exponential probability downsampling, x is the first quantity, d is the second quantity threshold, 0 is the third value, g is the fourth value, gx is the fifth value, e -gx The sixth value.
[0102] When the first quantity x is less than or equal to the second quantity threshold d, performing an operation on the set fourth value and the first quantity means multiplying the fourth value and the first quantity, and determining the resulting product gx as the fifth value. Performing an exponential operation on the fifth value means using the natural constant e as the base and the negative value of the fifth value as the exponent to obtain the value e. -gx As the sixth value. The ratio of the fourth value to the sixth value When the first quantity x is greater than the second quantity threshold d, the second sampling probability is 0.
[0103] The embodiment of the present application processes the first sampling data through exponential probability downsampling to obtain the second sampling data, thereby ensuring that for each data type there must be first sampling data that is sampled as the second sampling data, and when the amount of data of the same data type is large, the probability of the scanned data participating in the statistics is reduced as much as possible to ensure execution efficiency, further reduce the amount of data that needs to be processed, and improve the efficiency of subsequent determination of space occupancy.
[0104] In step 103 , a target data type is determined from the first data type, and third sample data belonging to the target data type is determined from the second sample data.
[0105] Here, the target data type is a data type arbitrarily selected from a plurality of first data types, for example, the target data type may be “address.” At least one second sampled data having the target data type “address” is selected from a plurality of second sampled data as the third sampled data.
[0106] In step 104 , based on the first space occupied by the third sampled data in the storage space, a second space occupied by the data of the target data type in the storage space is determined.
[0107] Here, the first space occupancy refers to the actual amount of space occupied by the third sampled data in the storage space. For example, if there are 200 third sampled data items with the target data type being "address," these 200 third sampled data items occupy a total of 200 MB of the storage space, and the first space occupancy is 200 MB. The second space occupancy refers to the actual amount of space occupied by all data items of the target data type in the storage space. This can be calculated using the first space occupancy and the second number of third sampled data items.
[0108] In some embodiments, see Figure 4 , Figure 4 It is shown that in step 104, determining the second space occupancy of the data of the target data type in the storage space based on the first space occupancy of the third sampled data in the storage space can be achieved by the following steps 1041A to 1044A:
[0109] In step 1041A, an average space occupancy of each third sample data is determined based on the first space occupancy and the second number of the third sample data.
[0110] Here, you can use a built-in tool in the database or a third-party memory detection tool to directly obtain the first storage space usage of the third sampled data. Obtain the second number of third sampled data and divide the first storage space usage by the second number to obtain the average storage space usage of each third sampled data item. For example, if the first storage space usage is 200MB and the second number of third sampled data items is 200, then the average storage space usage of each third sampled data item is 200MB / 200=1MB.
[0111] In step 1042A, a third space occupancy of the first sample data of the target data type in the storage space is determined based on the average space occupancy and the third number of the first sample data of the target data type.
[0112] Here, the third amount of first sampled data belonging to the target data type is obtained, and the third amount of space occupied by the first sampled data belonging to the target data type in the storage space is determined as the product of the average space occupancy and the third amount. For example, assuming the average space occupancy is 10 MB and the third amount of first sampled data belonging to the target data type is 50,000, then the third amount of space occupied = 10 MB × 5000 = 500 GB.
[0113] In step 1043A, the ratio of the third space occupancy to the fourth space occupancy is determined as the space occupancy ratio of data belonging to the target data type in the storage space.
[0114] The fourth space occupancy is the space occupancy of the first sample data of each first data type in the storage space.
[0115] Here, the space occupied by the first sampled data of each first data type in the storage space is obtained respectively. For each first data type, the specific process of obtaining the space occupied by the first sampled data of the first data type in the storage space is similar to the process of determining the third space occupied by the first sampled data of the target data type in the storage space in step 1042A, and will not be described again here. The space occupied by the first sampled data of each first data type in the storage space is added together to obtain a fourth space occupied amount. The ratio of the third space occupied amount to the fourth space occupied amount is determined as the space occupied ratio of the data of the target data type in the storage space. For example, the third space occupied amount is 500GB, the fourth space occupied amount (the total occupied amount of the first sampled data of all first data types) is 10TB, and the space occupied ratio = 500GB / 10000GB = 0.05 = 5%.
[0116] In step 1044A, a second space occupancy is determined based on the total space occupancy and the space occupancy ratio of the storage space.
[0117] Here, the total space occupied by the storage space is obtained, and the product of the total space occupied and the space occupied ratio is determined as the second space occupied by the data of the target data type in the storage space. The total space occupied by the storage space can be directly obtained using tools built into the database or third-party memory detection tools. For example, in the database, the total space occupied by the storage space can be directly queried using Structured Query Language (SQL).
[0118] In the embodiment of the present application, the second space occupancy satisfies the following formula (5).
[0119]
[0120] in, is the second space occupied by the data of the jth first data type in the storage space, θ is the total space occupied by the storage space, is the first space occupied by the third sampled data of the jth first data type in the storage space, is the second quantity of the third sampling data, is the third number of the j-th first sample data of the first data type.
[0121] The embodiment of the present application combines the total space occupancy and space occupancy ratio of the storage space to calculate the second space occupancy of the data belonging to the target data type in the storage space, significantly reducing the amount of data processed, improving computing efficiency, and ensuring the high accuracy and reliability of the results through multi-step refined processing.
[0122] In some embodiments, see Figure 5 , Figure 4 It is shown that determining the second space occupancy of the data of the target data type in the storage space based on the first space occupancy of the third sampled data in the storage space in step 104 can be achieved by the following steps 1041B to 1044B:
[0123] In step 1041B, the ratio of the third quantity to the fourth quantity is determined as the quantity ratio of data belonging to the target data type in the storage space.
[0124] The third number is the number of first sampled data belonging to the target data type, and the fourth number is the number of first sampled data of each first data type.
[0125] Here, the fourth number is the total number of first sampled data obtained after the first downsampling. For example, if the target data type is "address," the third number of address-type first sampled data is 50,000, and the fourth number of all first sampled data is 100,000, then the proportion of address-type data in the storage space = 50,000 / 100,000 = 0.5 = 50%.
[0126] In step 1042B, based on the total amount of data in the storage space and the amount ratio, a fifth amount of data of the target data type in the storage space is determined.
[0127] Here, the product of the total quantity and the quantity ratio may be determined as the fifth quantity of data belonging to the target data type in the storage space.
[0128] In step 1043B, an average space occupancy of each third sample data is determined based on the first space occupancy and the second number of the third sample data.
[0129] Here, the process of determining the average space occupancy of each third sampling data based on the first space occupancy and the second number of the third sampling data is consistent with step 1041A and will not be further described.
[0130] In step 1044B, the fifth number and the average space occupancy are processed to obtain a second space occupancy.
[0131] Here, performing operation processing on the fifth number and the average space occupancy refers to multiplying the fifth number by the average space occupancy, and determining the obtained product as the second space occupancy.
[0132] In the embodiment of the present application, the second space occupancy satisfies the following formula (6).
[0133]
[0134] in, is the second space occupied by the data of the jth first data type in the storage space, ω is the total amount of data in the storage space, is the first space occupied by the third sampled data of the jth first data type in the storage space, is the second quantity of the third sampling data, is the third number of first sample data of the j-th first data type, The fourth quantity.
[0135] The embodiment of the present application significantly reduces the amount of data to be processed and reduces computing resource consumption by performing first and second downsampling on the data in the storage space. Furthermore, the second sampled data is sampled data obtained by downsampling the first data type of the first sampled data, and third sampled data belonging to the target data type is determined from the second sampled data. The second space occupancy of the target data type in the entire storage space is inferred based on the first space occupancy of the third sampled data. This allows the space occupancy of all data of a certain data type to be determined based on the space occupancy of a small amount of data of that data type, thereby improving the efficiency of determining the space occupancy of data in the storage space.
[0136] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0137] In the analysis of the memory usage of distributed cache data and related technologies for memory governance in business scenarios, all RDB log files of the Redis cluster are generally obtained, and the RDB log file analysis tool is used to analyze the RDB log files according to business requirements to obtain the results that need to be counted. Alternatively, by exporting all the keys and values stored in Redis to a relational database or a data warehouse database, a data statistics program is written to perform statistics based on business needs. However, when the amount of data stored in Redis is too large, both of the above methods require a large amount of computing resources and time to complete. When the business scenario that needs to be counted is complex, the way to obtain the ideal results often needs to be constantly adjusted, resulting in uncontrollable cycles for obtaining statistical results. To this end, an embodiment of the present application provides a method for determining space usage, which is a method for dynamically evaluating cache memory consumption. It can significantly reduce the data statistics cycle and computing resource consumption while fitting the memory usage of cache components as much as possible. The details are explained below.
[0138] Figure 6 This is a flowchart of traversing data in a database provided by an embodiment of the present application. Figure 6, step 201, obtaining a database link. The database (corresponding to the storage space in the above embodiment) is a Redis database, and obtaining a database link means establishing a communication connection with the Redis server.
[0139] Step 202: Traverse the keys in the database based on the given prefix. Based on the given prefix, the keys in Redis can be traversed using the Scan command. That is, a loop is started to execute the Scan command to scan all keys in Redis that match the given regular expression to obtain a key list (corresponding to the data in the above embodiment).
[0140] Step 203, process the scan result according to the downsampling method. The key list scanned in step 202 can be processed according to the first downsampling method (corresponding to the first downsampling in the above embodiment) to obtain a downsampling set (corresponding to the first sampling data in the above embodiment). The first downsampling method is used to ensure that all samples can be evenly distributed in all slots in Redis. Commonly used first downsampling methods include random downsampling and interval sampling. Random downsampling samples the keys in the key list according to the probability of the above formula (1). Wherein, x represents the xth key to be placed in the downsampling set, a is used to control the probability of the sample to be collected (corresponding to the first sampling probability in the above embodiment). When a=1, all keys in the key list will be sampled. When a>1, all keys in the key list will be sampled according to the probability of 1 / a. Interval sampling samples the keys in the key list according to the interval of the above formula (2). Wherein, x represents the xth key to be placed in the downsampling set, b is used to control the sampling interval, and sampling is performed every b scans.
[0141] Step 204: traverse the downsampling set. Traverse the downsampling set of each Scan.
[0142] Step 205: Get the prefix / suffix of the current key being traversed. Get the prefix / suffix of the current key being traversed, and define it as a variable PrefixKey (corresponding to the first data type in the above embodiment).
[0143] Step 206, downsample the current key. Add the current key to the statistical list of the PrefixKey corresponding to the current key (corresponding to the second sampling data in the above embodiment) according to the probability of the second downsampling method (corresponding to the second downsampling in the above embodiment). The second downsampling method is used to ensure that there is at least one key that can be counted for each prefix / suffix. Commonly used second downsampling methods include piecewise linear probability downsampling and exponential probability downsampling. Piecewise linear probability downsampling samples the keys in the downsampled set according to the probability of the above formula (3) (corresponding to the second sampling probability in the above embodiment), where x represents the xth key to be included in the statistical list, and c and d are used to control the number of sampled keys and the probability of the key being added to the statistical list. This probability distribution function ensures that the first scanned key under each PrefixKey will be added to the statistical list, ensuring that some prefix / suffix keys that need to be counted will not be missed due to downsampling, resulting in inaccurate calculated values. This probability distribution can include as many scanned keys as possible in the statistics when the amount of prefix / suffix data is small, ensuring that the statistical results of the sampled data fit the statistical results of the full data as much as possible. When the amount of prefix / suffix data is large, the probability of the scanned key participating in the statistics will be reduced as much as possible to ensure execution efficiency.
[0144] Exponential probability downsampling samples the keys in the downsampling set according to the probability of the above formula (4), where g is used to control the distribution of the probability of sampling acquisition during the downsampling process, and d is used to control the number of keys sampled at most. This probability distribution allows keys to participate in the data statistics as much as possible when the amount of key data participating in the statistics is small. As the amount of key data participating in the statistics increases, the probability of the next scanned key participating in the statistics will decrease exponentially, thereby improving sampling efficiency.
[0145] It should be noted that in the embodiment of the present application, all keys are evenly distributed in each slot of Redis. The memory usage of keys with the same prefix or suffix is normally distributed.
[0146] Step 207: Determine whether the downsampling is successful. Determine whether the current key passes the second downsampling. If the current key passes the second downsampling, add the key to the statistics list and jump to step 208. If the current key does not pass the second downsampling, jump to step 204 and determine the next key traversed.
[0147] Step 208: Obtain memory usage through a command. If the current key is included in the statistics, obtain the space usage of the current key through the Redis memory usage command (memory usage key). Record the number of keys under each PrefixKey currently scanned, the number of keys under each PrefixKey participating in the statistics in the statistics list, and the memory usage of each key under each PrefixKey participating in the statistics in the statistics list. For each PrefixKey, if the number of keys participating in the statistics under the current PrefixKey is greater than a given threshold, the key is not included in the statistics.
[0148] Step 209: Determine whether the traversal is complete. If the current key is the last key in the downsampling list, the traversal is complete and the process jumps to step 210. Otherwise, the process jumps to step 204 and traverses the next key.
[0149] Step 210: Get all the keys in the statistical list and obtain statistical data. Traverse the number of keys under each PrefixKey counted in this statistical list. For each PrefixKey, calculate the average memory space occupied by each key (corresponding to the average space occupied in the above embodiment), and then evaluate the total space occupied by all keys under the PrefixKey in Redis (corresponding to the second space occupied in the above embodiment) by the proportion of the number of keys under the PrefixKey counted (corresponding to the number ratio in the above embodiment).
[0150] Step 211. Determine whether the statistics are finished. If the data volume of the currently counted Key meets the stopping condition, the statistics are finished and jump to step 212. Otherwise, jump to step 202 and get the remaining Keys from Redis. The stopping condition is one of the following: Get the number of Keys currently traversed. If the number of Keys currently traversed meets the preset condition, the statistics are finished. Alternatively, for each PrefixKey, get the number of Keys under the PrefixKey in the current statistical list. If the number of Keys under the PrefixKey in the statistical list under each PrefixKey meets the respective preset quantity conditions, the statistics are finished. When the data volume of the KEYs of all prefixes is sampled to the preset data volume, the statistical process will be exited in advance, which can reduce unnecessary data access and further reduce the time consumption of data statistics.
[0151] Step 212: Output statistical results, including the number of keys for each PrefixKey in the statistical list, memory usage, and the number of keys for each PrefixKey in the downsampled set at the end of the statistical process.
[0152] After the statistics are completed, the memory usage in Redis is estimated. Assume that the set of all prefixes is all, and the total memory usage of the current system (corresponding to the total space usage of the storage space in the above embodiment) is θ. Through the above process, the number of keys participating in the statistics for each PrefixKey, n, can be obtained. prefix 、Sample data memory usage total prefix , the total number of keys N for each PrefixKey until the statistics program stops scanning prefix , then the overall space occupied by the specified prefix can be estimated by the above formula (5) or (6).
[0153] If the total memory usage and total number of keys in Redis are not obtained, the memory usage of a single prefix key can be estimated using the following formula (7).
[0154]
[0155] The embodiment of the present application adopts a downsampling method when performing data statistics. During the sampling process, the sampling rate will decay rapidly as the number of samples increases. On the one hand, the sampling efficiency is improved, and on the other hand, the data sampled in the later stage of the scan can be collected as much as possible. Taking into account that the distribution of the data set in the statistical target is normally distributed, when enough sample data is sampled, the statistical process will be stopped in time, and only a limited sample data set will be used to fit the distribution of the overall data set. This statistical method can quickly complete data statistics in the scenario of large data volume and obtain approximate statistical results. Through experiments, in the Redis memory of 500 million data, under the premise that the Redis Key prefix and suffix are limited, according to the expected number of samples, the statistical results can be obtained at the minute level (0-30 minutes), and the error varies with the number of samples. The more samples, the more accurate it is. Compared with the prior art, the time cost and machine cost are greatly reduced.
[0156] The following continues to describe the exemplary structure of the space occupancy determination device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the space occupancy determination device 455 of the memory 450 may include:
[0157] The sampling module 4551 is configured to perform a first downsampling on the data in the storage space to obtain first sampled data;
[0158] A sampling module 4551 is configured to perform a second downsampling on the first sampled data based on the first data type of the first sampled data to obtain second sampled data;
[0159] an acquisition module 4552, configured to determine a target data type from the first data type, and determine third sampled data belonging to the target data type from the second sampled data;
[0160] The determining module 4553 is configured to determine a second space occupancy of data of the target data type in the storage space based on the first space occupancy of the third sampled data in the storage space.
[0161] In some embodiments, the sampling module 4551 is further configured to generate a first random number corresponding to each data in the storage space; and determine the data corresponding to the first random number that is less than or equal to the first sampling probability as the first sampling data.
[0162] In some embodiments, the first downsampling refers to sampling at a preset sampling interval. The sampling module 4551 is further configured to determine, for each data item in the storage space, the location information of the data item among the multiple data items in the storage space; perform a preset operation on the sampling interval based on the location information; and determine the data item as the first sampled data item if the result of the preset operation is the same as the set first value.
[0163] In some embodiments, the first sampled data are arranged in sequence based on the location information of the first sampled data in the storage space. The sampling module 4551 is further configured to traverse the first sampled data one by one; for the ith first sampled data traversed, determine the second data type of the ith first sampled data, and determine a first number of fourth sampled data belonging to the second data type, where the fourth sampled data is the first sampled data among the first i-1 first sampled data that is sampled as the second sampled data, and i is a positive integer that increases sequentially; determine a second sampling probability that the ith first sampled data is sampled as the second sampled data based on the first number, the first number threshold, and the second number threshold; generate a second random number corresponding to the ith first sampled data, and determine the ith first sampled data as the second sampled data if the second random number is less than or equal to the second sampling probability.
[0164] In some embodiments, the sampling module 4551 is further used to determine the set second value as the second sampling probability that the i-th first sampling data is sampled as the second sampling data when the first number is less than or equal to the first number threshold; when the first number is greater than the first number threshold and less than or equal to the second number threshold, divide the data based on the first number threshold and the second number threshold to obtain multiple first value intervals, wherein the second number threshold is greater than the first number threshold; determine the first value interval including the first number as the second value interval, and determine the inverse of the maximum endpoint value of the second value interval as the second sampling probability that the i-th first sampling data is sampled as the second sampling data; when the first number is greater than the second number threshold, determine the set third value as the second sampling probability that the i-th first sampling data is sampled as the second sampling data, wherein the second value is greater than the inverse of the maximum endpoint value, and the inverse of the maximum endpoint value is greater than the third value.
[0165] In some embodiments, the determination module 4553 is further used to determine the average space occupancy of each third sampling data based on the first space occupancy and the second number of the third sampling data; determine the third space occupancy of the first sampling data belonging to the target data type in the storage space based on the average space occupancy and the third number of the first sampling data belonging to the target data type; determine the ratio of the third space occupancy to the fourth space occupancy as the space occupancy ratio of the data belonging to the target data type in the storage space, wherein the fourth space occupancy is the space occupancy of the first sampling data of each first data type in the storage space; and determine the second space occupancy based on the total space occupancy and the space occupancy ratio of the storage space.
[0166] In some embodiments, the determination module 4553 is further used to determine the ratio of the third quantity to the fourth quantity as the ratio of the quantity of data belonging to the target data type in the storage space, wherein the third quantity is the quantity of the first sampled data belonging to the target data type, and the fourth quantity is the quantity of the first sampled data of each first data type; based on the total quantity and quantity ratio of the data in the storage space, determine the fifth quantity of the data belonging to the target data type in the storage space; based on the first space occupancy and the second quantity of the third sampled data, determine the average space occupancy of each third sampled data; and perform calculations on the fifth quantity and the average space occupancy to obtain the second space occupancy.
[0167] An embodiment of the present application provides a computer program product, comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the method for determining the space usage described in the embodiment of the present application.
[0168] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the method for determining the space occupancy provided by the embodiment of the present application, for example, Figure 3 The method for determining the space occupancy is shown.
[0169] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0170] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0171] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0172] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0173] In summary, the embodiments of the present application can significantly reduce the data statistical cycle and computing resource consumption, and improve the efficiency of determining data memory occupancy.
[0174] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A method for determining space occupancy, characterized in that: The method comprises: Performing a first downsampling on the data in the storage space to obtain first sampled data; Based on the first data type of the first sampled data, performing a second downsampling on the first sampled data to obtain second sampled data; determining a target data type from the first data type, and determining third sampled data belonging to the target data type from the second sampled data; Based on a first space occupied by the third sampled data in the storage space, a second space occupied by data of the target data type in the storage space is determined.
2. The method according to claim 1, characterized in that The first downsampling of the data in the storage space to obtain first sampled data includes: For each data in the storage space, generate a first random number corresponding to the data; Data corresponding to a first random number that is less than or equal to the first sampling probability is determined as first sampled data.
3. The method according to claim 1, characterized in that The first downsampling refers to sampling according to a preset sampling interval; the first downsampling of the data in the storage space to obtain the first sampled data includes: For each data in the storage space, determining location information of the data among multiple data in the storage space; Based on the position information, performing a preset operation on the sampling interval; When the result of the preset operation is the same as the set first value, the data is determined as the first sampling data.
4. The method according to claim 1, wherein The first sampled data are arranged in sequence based on position information of the first sampled data in the storage space; and the second downsampling of the first sampled data based on the first data type of the first sampled data to obtain the second sampled data includes: Traversing the first sampling data one by one; For the traversed i-th first sampled data, determine the second data type of the i-th first sampled data, and determine a first number of fourth sampled data belonging to the second data type, wherein the fourth sampled data is the first sampled data sampled as the second sampled data among the first i-1 first sampled data, and i is a positive integer that increases successively; Determine, based on the first quantity, the first quantity threshold, and the second quantity threshold, a second sampling probability that the i-th first sampling data is sampled as the second sampling data; Generate a second random number corresponding to the i-th first sampling data, and determine the i-th first sampling data as the second sampling data when the second random number is less than or equal to the second sampling probability.
5. The method according to claim 4, characterized in that The determining, based on the first quantity, the first quantity threshold, and the second quantity threshold, a second sampling probability that the i-th first sampling data is sampled as the second sampling data includes: When the first number is less than or equal to the first number threshold, determining a set second value as a second sampling probability that the i-th first sampling data is sampled as the second sampling data; When the first number is greater than the first number threshold and less than or equal to the second number threshold, dividing based on the first number threshold and the second number threshold to obtain multiple first numerical intervals, wherein the second number threshold is greater than the first number threshold; Determine a first numerical interval including the first number as a second numerical interval, and determine the reciprocal of a maximum endpoint value of the second numerical interval as a second sampling probability that the i-th first sampling data is sampled as the second sampling data; When the first number is greater than the second number threshold, the set third value is determined as the second sampling probability that the i-th first sampling data is sampled as the second sampling data, wherein the second value is greater than the inverse of the maximum endpoint value, and the inverse of the maximum endpoint value is greater than the third value.
6. The method according to any one of claims 1 to 5, characterized in that The determining, based on the first space occupied by the third sampled data in the storage space, a second space occupied by the data of the target data type in the storage space, includes: determining an average space occupancy of each third sample data based on the first space occupancy and the second number of the third sample data; determining a third space occupancy of the first sampled data of the target data type in the storage space based on the average space occupancy and a third number of the first sampled data of the target data type; determining a ratio of the third space occupancy to the fourth space occupancy as a space occupancy ratio of data of the target data type in the storage space, wherein the fourth space occupancy is the space occupancy of each first sample data of the first data type in the storage space; The second space occupancy is determined based on the total space occupancy of the storage space and the space occupancy ratio.
7. The method according to any one of claims 1 to 5, characterized in that The determining, based on the first space occupied by the third sampled data in the storage space, a second space occupied by the data of the target data type in the storage space, includes: determining a ratio of the third quantity to the fourth quantity as a quantity ratio of data belonging to the target data type in the storage space, wherein the third quantity is the quantity of first sampled data belonging to the target data type, and the fourth quantity is the quantity of first sampled data of each first data type; determining a fifth amount of data of the target data type in the storage space based on the total amount of data in the storage space and the amount ratio; determining an average space occupancy of each third sample data based on the first space occupancy and the second number of the third sample data; The fifth number and the average space occupancy are processed to obtain the second space occupancy.
8. A device for determining space occupancy, characterized in that: The device comprises: a sampling module, configured to perform a first downsampling on the data in the storage space to obtain first sampled data; The sampling module is configured to perform a second downsampling on the first sampled data based on the first data type of the first sampled data to obtain second sampled data; an acquisition module, configured to determine a target data type from the first data type, and determine third sampled data belonging to the target data type from the second sampled data; The determining module is configured to determine a second space occupancy of data of the target data type in the storage space based on a first space occupancy of the third sampled data in the storage space.
9. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; A processor is configured to implement the method for determining the space occupancy as described in any one of claims 1 to 7 when executing the computer-executable instructions or computer program stored in the memory.
10. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method for determining the space occupancy according to any one of claims 1 to 7 is implemented.