A cache management method of a CDN network

By using sequential disk writes and consistent hashing algorithms in the CDN network, combined with forward distance calculation to determine hot files, the problem of hot file management occupying memory and CPU is solved, and the performance and efficiency of the cache server are improved.

CN117857631BActive Publication Date: 2025-10-14CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311704362.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-10-14
Estimated Expiration
2043-12-13

AI Technical Summary

Technical Problem

In existing CDN networks, the management and maintenance of hot files occupy a large amount of system memory and CPU resources, reducing the performance of the cache server.

Method used

The cache is managed by sequentially writing to disk, combining the consistent hashing algorithm and forward distance calculation to determine the probability of file hot spots, optimize the storage location of hot files, and avoid additional memory queries.

Benefits of technology

It reduces the memory and CPU consumption of hot file management and improves the response speed and utilization of the cache server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117857631B_ABST
    Figure CN117857631B_ABST
Patent Text Reader

Abstract

The application discloses a CDN network cache management method, relates to the technical field of content distribution network, cache management, hot spot cache and performance, and effectively manages and maintains the hot spot cache on a cache server to improve the hit rate of the cache, greatly reduces the consumption of CPU and memory of the hot spot cache management and maintenance, improves the response speed of the cache server and other performance indexes, maintains the hot spot file with the latest access, makes the hot spot file save for the longest time on the disk, improves the utilization rate of the cache, simultaneously does not need additional memory information to save and count the hot spot file, avoids the process of searching the memory information table of the hot spot file when judging the hot spot file, and greatly improves the processing performance of the cache server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cache management, and in particular relates to a cache management method for a CDN network. Background Art

[0002] In a CDN network, on-demand videos or website images and other files need to be cached on a cache server. Caching means copying the required static resource files and storing them on the CDN cache server. Since the storage capacity of a single cache server is actually upper bounded, the cached files can be considered unlimited.

[0003] To improve the quality of CDN services, ensure cache hit rates, and reduce the large number of back-to-source requests due to cache misses, it's crucial to effectively store and maintain hotspot files on the cache server. Currently, the industry's general practice is to record the access time of recently visited URLs in memory or to count the number of URL accesses. Regardless of which method is used, since cache servers hold tens of millions or even hundreds of millions of files, a large amount of memory must be allocated to store hotspot file information. Furthermore, when the cache is eliminated, memory must be queried to determine the hotspot file. Maintaining hotspot files consumes a significant amount of system memory and CPU resources for hotspot queries, which undoubtedly reduces the system performance of the cache server. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a cache management method for a CDN network in response to the shortcomings of the background technology, effectively manage and maintain the hotspot cache on the cache server to improve the cache hit rate, greatly reduce the CPU and memory consumption of hotspot cache management and maintenance, and improve the performance indicators such as the response speed of the cache server.

[0005] The present invention adopts the following technical solutions to solve the above technical problems:

[0006] A CDN network cache management method specifically includes the following steps:

[0007] Step 1: Upon receiving a user request, the cache server checks whether the cache file exists on the cache server. If not, it retrieves the file content from the origin server, selects a disk based on the configured HASH algorithm, and writes it to the cache disk in order. It also records the location of the file that has been written to the current disk.

[0008] Step 2: If user requests continue to be received before the disk is full and the cache does not exist in the cache server, the current disk write location is recorded from step 1 and the new cache file is continuously written, while the current disk write location is updated.

[0009] Step 3: After the disk is full, if a user request is received and the requested cache does not exist in the cache server, the server will continue writing from the current write position, directly overwriting the old cache at the current position.

[0010] Step 4: When the disk is full, if the currently requested file exists in the cache server, the file is read from the disk to complete the client response, and the forward distance between the current file position in the cache and the current write position is calculated;

[0011] Step 5: Determine whether the currently accessed cache is a hotspot based on the hotspot probability obtained from the forward distance. If so, read the currently accessed cache and rewrite it to the current write location. Otherwise, keep it in the original storage location.

[0012] As a further preferred solution of the cache management method of a CDN network of the present invention, in step 1, it is determined whether the cache file exists on the cache server, specifically as follows: the cache server will save the index of all cached URLs, and calculate a unique HASH value based on the URL requested by the user to search in the cache index. If found, the file exists on the cache server.

[0013] As a further preferred solution of the cache management method of a CDN network of the present invention, in step 1, the cache server adopts a sequential circular writing method for the cache.

[0014] As a further preferred solution of the cache management method of a CDN network of the present invention, in step 1, the configured HASH algorithm adopts a consistent hashing algorithm, which specifically includes the following steps:

[0015] Step 1.1: Use the disk size as the weight to build a consistent hash ring for all disks.

[0016] In step 1.2, a unique HASH value is calculated based on the URL requested by the user, and a specific disk is selected on the hash ring based on the HASH value.

[0017] As a further preferred solution of the cache management method of a CDN network of the present invention, in step 4, the forward distance is calculated as follows: if the storage position of the current file is after the write position, the current write position is subtracted from the current file storage position.

[0018] As a further preferred embodiment of a cache management method for a CDN network of the present invention, in step 4, the forward distance is calculated as follows: if the current file storage position is before the current write position, the forward distance is: the end position of the disk minus the current write position plus the current file storage position.

[0019] As a further preferred embodiment of the cache management method of a CDN network of the present invention, in step 5, the hotspot probability obtained according to the forward distance is specifically as follows: the relationship between the forward distance and the hotspot probability has a corresponding hotspot curve, and a uniquely determined hotspot probability can be obtained from the hotspot curve according to the forward distance.

[0020] As a further preferred embodiment of the cache management method for a CDN network of the present invention, in step 5, it is determined whether the currently accessed cache is a hotspot, specifically as follows: a probability value is obtained in a probability curve based on the obtained forward distance, and whether the currently accessed cache is a hotspot is determined based on the probability value.

[0021] As a further preferred solution of the cache management method of a CDN network of the present invention, in step 5, a curve of the forward distance of the cache file and the hotspot probability is set to adopt a percentage of the disk.

[0022] As a further preferred solution of the cache management method of a CDN network of the present invention, in step 5, a curve of the forward distance of the cache file and the hotspot probability is set to adopt a specific value.

[0023] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0024] The present invention achieves the purpose of maintaining the hot files that have been accessed recently, so that the hot files can be stored on the disk for the longest time, thereby improving the utilization rate of the cache; at the same time, since the present invention does not require additional memory information to save statistical hot files, it also avoids the need to search the hot file memory information table when judging the hot files, thereby greatly improving the processing performance of the cache server. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 It is a flow chart of the present invention;

[0027] Figure 2 The present invention sets a relationship curve between the forward distance and the hotspot probability;

[0028] Figure 3 is an embodiment of the present invention;

[0029] Figure 4 It is the calculation of the forward distance of the file cache position after the write position of the present invention;

[0030] Figure 5 It is the calculation of the forward distance of the file cache position before the write position of the present invention. DETAILED DESCRIPTION

[0031] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings:

[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The present invention is described in detail below based on the drawings and preferred embodiments. The purpose and effect of the present invention will become more clear. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0033] In view of the background technology of the hot file cache storage and update method on the cache server, the solution proposed in this article can greatly reduce the memory consumption and CPU consumption of hot cache query, and effectively maintain the hot files on the cache server, thereby improving the system performance of the cache server.

[0034] The specific scheme of the present invention is as follows:

[0035] 1) Upon receiving a user request, the cache server determines whether the cache file exists on the cache server. If not, it obtains the file content from the origin server, selects the disk according to the configured HASH algorithm, and writes it to the cache disk in order. It also records the current disk write location. Before the disk is full, if user requests continue to be received and the cache file does not exist on the cache server, it continues to write the new cache file from the location in step 1 and updates the current disk write location.

[0036] The HASH algorithm uses a consistent hashing algorithm: all disks are weighted according to the size of the disk to build a consistent hash ring, a unique HASH value is calculated according to the URL requested by the user, and a certain disk is selected on the hash ring based on this HASH value.

[0037] 2) Before the disk is full, if user requests continue to be received and the cache does not exist in the cache server, then the current disk write location is recorded from step 1 and the new cache file is continuously written, while the current disk write location is updated;

[0038] 3) After the disk is full, if a user request is received and the requested cache does not exist in the cache server, the server will continue writing from the current write position, directly overwriting the old cache at the current position;

[0039] 4) When the disk is full, if the currently requested file exists in the cache server, the file is read from the disk to complete the client response, and the forward distance between the current file position in the cache and the current write position is calculated;

[0040] When the disk is full, if the currently requested file exists in the cache server, the file will be read from the disk to complete the client response. At the same time, the forward distance between the current file's position in the cache and the current write position will be calculated. Then, based on the current distance and the preset parameters of the hotspot probability, if it is determined to be a hotspot, the currently requested cache file will be rewritten to the previously written position on the disk, and the current write position will be updated at the same time.

[0041] The forward distance is calculated as follows: if the current file storage position is after the write position, subtract the current write position from the current file storage position; if the current file storage position is before the current write position, the forward distance is calculated as: the end of the disk minus the current write position plus the current file storage position.

[0042] 5) Based on the hotspot probability obtained from the forward distance, determine whether the currently accessed cache is a hotspot. If so, read the currently accessed cache and rewrite it to the current write location. Otherwise, keep it in the original storage location.

[0043] Sets a curve for the forward distance of the cached file and its hotspot probability. This can be expressed as a percentage of the disk or a specific value. The smaller the forward distance, the greater the hotspot probability. If the hotspot probability calculated based on the forward distance is determined not to be a hotspot, the currently accessed cache remains in its original storage location.

[0044] With sequential writes to disk, the file at the current write location is the file that can be cached the longest, because if you want to overwrite the current write file, the cache must be filled up once and then returned to this location;

[0045] Files that have been accessed recently are considered relatively hot files and will be written to the current disk write location.

[0046] And set a relationship between the forward distance and the hotspot probability. The shorter the forward distance, the greater the hotspot probability. If the hotspot determination is passed, the file being accessed will be re-written in sequence to the current disk write location to avoid being overwritten by other requests.

[0047] The hotspot probability obtained according to the forward distance is as follows: the relationship between the forward distance and the hotspot probability has a corresponding hotspot curve, and the unique hotspot probability can be obtained from the hotspot curve according to the forward distance.

[0048] Determine whether the currently accessed cache is a hotspot as follows: Based on the forward distance, calculate the probability value from the probability curve. Use this probability value to determine whether the currently accessed cache is a hotspot. Set the forward distance vs. hotspot probability curve for the cached file to use a percentage of the disk. Set the forward distance vs. hotspot probability curve for the cached file to use a specific value.

[0049] Based on the forward distance, we can derive a probability value from the probability curve. For example, if the probability value is 95%, then approximately 95 out of 100 user access requests that meet this forward distance will be considered hotspots. For example, if the probability value is 30%, then approximately 30 out of 100 requests will be considered hotspots. (This is just an example; in actual calculations, each request is judged based on the probability value. As the number of requests increases, the probability distribution will generally be met.)

[0050] Example:

[0051] 1. First, the disk is not full yet, so the new cache files are written sequentially and the current write position is updated;

[0052] 2. When the disk is full, continue writing new cache files sequentially from the beginning;

[0053] 3. Set the curve of the forward distance and hotspot probability of the cache file. It can be a percentage of the disk or a specific size value. The smaller the forward distance, the greater the hotspot probability.

[0054] 4. For the currently requested file a, calculate its forward distance and then, based on the hotspot probability, determine whether it is a hotspot cache. If so, read the cache of a and rewrite it to the current write location.

[0055] 5. If the currently requested file a is not in the hotspot cache, calculate its forward distance and make a decision based on the hotspot probability. In this case, it is not rewritten and can be retained in the current cache location.

[0056] The cache server uses a sequential, cyclical write cache method. Maintaining the hotspot cache does not require additional memory for storing hotspot file information, nor does it require additional memory lookups, reducing CPU consumption and significantly improving the cache server's storage efficiency. The forward distance of the current cache is calculated. Based on the relationship between the preset forward distance and the hotspot probability, if a judgment is passed, the hotspot cache is rewritten to the current disk write location, effectively managing the hotspot cache files in the cache server.

[0057] This invention achieves the goal of maintaining recently accessed hot files, ensuring that hot files are stored on disk for the longest possible time, thereby improving cache utilization. Furthermore, since this invention does not require additional memory information to store statistics on hot files, it also avoids the need to search the hot file memory information table when determining hot files, greatly improving the processing performance of the cache server.

[0058] Those skilled in the art will understand that the above description is merely a preferred embodiment of the invention and is not intended to limit the invention. Although the invention has been described in detail with reference to the above examples, those skilled in the art can still modify the technical solutions described in the above examples or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, etc. made within the spirit and principles of the invention shall be included in the scope of protection of the invention. All technical features in this embodiment can be freely combined according to actual needs.

[0059] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A cache management method for a CDN network, characterized by: The specific steps include: Step 1: Upon receiving a user request, the cache server checks whether the cache file exists on the cache server. If not, it retrieves the file content from the origin server, selects a disk based on the configured HASH algorithm, and writes it to the cache disk in order. It also records the location of the file that has been written to the current disk. Step 2: Before the disk is full, if user requests continue to be received and the cache file does not exist on the cache server, the current disk write location is recorded from step 1 and new cache files are continuously written, while the current disk write location is updated. Step 3: After the disk is full, if a user request is received and the requested cache does not exist in the cache server, the server will continue writing from the current write position, directly overwriting the old cache at the current position. Step 4: When the disk is full, if the currently requested file exists in the cache server, the file is read from the disk to complete the client response, and the forward distance between the current file position in the cache and the current write position is calculated; Step 5: Determine whether the currently accessed cache is a hotspot based on the hotspot probability obtained from the forward distance. If so, read the currently accessed cache and rewrite it to the current write location. Otherwise, keep it in the original storage location.

2. A cache management method for a CDN network according to claim 1, characterized in that: In step 1, determine whether the cache file exists on the cache server, as follows: the cache server will save the index of all cached URLs, calculate the unique HASH value based on the URL requested by the user, and search in the cache index. If found, the file exists on the cache server.

3. The cache management method of a CDN network according to claim 1, characterized in that: In step 1, the cache server adopts a sequential circular writing method for the cache.

4. The cache management method of a CDN network according to claim 1, characterized in that: In step 1, the configured HASH algorithm adopts the consistent hashing algorithm, which includes the following steps: Step 1.1: Use the disk size as the weight to build a consistent hash ring for all disks. In step 1.2, a unique HASH value is calculated based on the URL requested by the user, and a specific disk is selected on the hash ring based on the HASH value.

5. The cache management method of a CDN network according to claim 1, characterized in that: In step 4, the forward distance is calculated as follows: if the storage position of the current file is after the writing position, the current writing position is subtracted from the current file storage position.

6. The cache management method of a CDN network according to claim 1, characterized in that: In step 4, the forward distance is calculated as follows: if the current file storage position is before the current write position, the forward distance is: the end position of the disk minus the current write position plus the current file storage position.

7. The cache management method of a CDN network according to claim 1, characterized in that: In step 5, the hotspot probability obtained according to the forward distance is as follows: the relationship between the forward distance and the hotspot probability has a corresponding hotspot curve, and a uniquely determined hotspot probability can be obtained from the hotspot curve according to the forward distance.

8. The cache management method of a CDN network according to claim 1, characterized in that: In step 5, it is determined whether the currently accessed cache is a hotspot, specifically as follows: a probability value is obtained in the probability curve according to the obtained forward distance, and whether the currently accessed cache is a hotspot is determined according to the probability value.

9. The cache management method of a CDN network according to claim 1, characterized in that: In step 5, the curve of the forward distance of the cache file and the hotspot probability is set to use the percentage of the disk.

10. The cache management method of a CDN network according to claim 1, characterized in that: In step 5, a curve of the forward distance of the cache file and the hotspot probability is set to a specific value.

Citation Information

Patent Citations

  • Memory caching method oriented to range querying on Hadoop

    CN103942289A

  • Hotspot data caching method and device, equipment and medium

    CN116028389A