High-concurrency storage method and system for massive image files, and readable medium

By adopting asynchronous processing and batch processing technologies in the distributed storage architecture, the performance bottleneck problem during high concurrent storage requests is solved, and the system throughput and efficiency improvement is achieved.

WO2025118631A1PCT designated stage expired Publication Date: 2025-06-12CRSC COMM & INFORMATION GRP CO LTD

Patent Information

Application Number
PCT/CN2024/107537
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-07-25
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In the prior art, distributed storage architectures are prone to performance bottlenecks when high concurrent storage requests, resulting in reduced system efficiency and waste of resources.

Method used

Asynchronous processing and batch processing technology are used to place user storage requests into message queues and distributed caches, and similar storage tasks are processed asynchronously through thread pools, and merge them into batch requests and sent to the server at one time.

Benefits of technology

Improves system throughput and efficiency, reduces request overhead and system load, and enhances storage processing speed and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024107537_12062025_PF_FP_ABST
    Figure CN2024107537_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A high-concurrency storage method and system for massive image files, and a readable medium, relating to the technical field of image file storage. The method comprises: putting a storage request of a user into a message queue and a distributed cache, and performing asynchronous processing on the storage request; taking out, from the message queue, the storage request having undergone asynchronous processing, and allocating the storage request to a thread pool; when threads in the thread pool take out storage tasks from the message queue, performing storage on the basis of the storage tasks; and combining similar storage tasks in the thread pool into a batch processing request, and sending the batch processing request to a server at one time. For a high-concurrency storage request, the throughput and efficiency of a system can be improved; batch processing of similar requests reduces the overheads of requests and increases the storage processing speed; and with the increase of data volumes, the capacity and the processing capability of a storage system are effectively and simply expanded.
Need to check novelty before this filing date? Find Prior Art

Description

A high-concurrency storage method, system and readable medium for massive image files Technical Field

[0001] The present invention relates to the field of, but is not limited to, big data processing technology, and belongs to the field of massive image file processing in image file storage. Background Art

[0002] Currently, image file storage primarily relies on local hard drives. This approach has limitations, such as limited storage capacity and access speeds that are affected by physical device performance. Furthermore, traditional image storage technologies often require manual operations for data backup, disaster recovery, and fault recovery, increasing management and maintenance complexity. Alternatively, image files rely on distributed storage architectures. Massive image storage typically employs a distributed storage architecture, distributing large amounts of image data across multiple nodes. This architecture offers higher storage capacity and access speed, as well as high availability and scalability. By distributing data across multiple nodes, system processing power and throughput are increased while reducing the load on individual nodes. However, this architecture can present performance bottlenecks, as each storage request must be processed immediately, which can lead to performance bottlenecks and reduce system efficiency. Furthermore, only one request can be processed at a time, which can waste resources because the system cannot share resources when processing multiple requests. This can also increase request latency, impacting user experience. SUMMARY OF THE INVENTION

[0003] In response to the above problems, the purpose of the present invention is to provide a high-concurrency storage method, system and readable medium for massive image files, which can improve the system's throughput and efficiency for high-concurrency storage requests, batch process similar requests, reduce request overhead, and improve storage processing speed. As the amount of data grows, the capacity and processing power of the storage system can be effectively and simply expanded. Technical issues

[0004] Existing architectures can experience performance bottlenecks, requiring immediate processing of each storage request. This can lead to performance bottlenecks and reduce system efficiency. Furthermore, only one request can be processed at a time, which can waste resources because the system cannot share resources when processing multiple requests. This can also increase request latency and impact the user experience. Technical Solutions

[0005] To achieve the above-mentioned objectives, the present invention proposes the following technical solutions: a high-concurrency storage method for massive image files, comprising the following steps: placing a user's storage request into a message queue and a distributed cache, and asynchronously processing the storage request; taking out the asynchronously processed storage request from the message queue, and assigning the storage request to a thread pool; after a thread in the thread pool takes out a storage task from the message queue, storing the task according to the storage task; and combining similar storage tasks in the thread pool into a batch request, which is sent to the server at one time.

[0006] Furthermore, the server is a Minio server or a Ceph server.

[0007] Furthermore, the distributed cache sets an expiration policy so that the data in the distributed cache is kept consistent with the data in the message queue.

[0008] Furthermore, the asynchronous processing is to put the storage request into a message queue after the storage request is issued, asynchronously obtain the storage request from the message queue, and perform subsequent processing operations.

[0009] Furthermore, the thread pool includes two or more threads, and the thread pool manages the life cycles of the threads so that the storage tasks are evenly distributed to each thread.

[0010] Furthermore, the thread pool evenly distributes storage tasks to each thread through a load balancing strategy of round-robin or least load.

[0011] Furthermore, the service uses a load balancer to evenly distribute the storage tasks to each server.

[0012] The present invention also discloses a high-concurrency storage system for massive image files, comprising: an asynchronous processing module for placing a user's storage request into a message queue and a distributed cache, and performing asynchronous processing on the storage request; a multi-threading module for taking out the asynchronously processed storage request from the message queue and assigning the storage request to a thread pool; a storage module for storing according to the storage task after the thread in the thread pool takes out the storage task from the message queue; and a batch processing module for combining similar storage tasks in the thread pool into a batch processing request and sending it to the server at one time.

[0013] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the high-concurrency storage method for massive image files described above. Beneficial effects

[0014] The present invention has the following advantages due to the adoption of the above technical solution:

[0015] 1) This invention targets highly concurrent storage requests, can improve system throughput and efficiency, batch process similar requests, reduce request overhead, increase storage processing speed, and effectively and simply expand the capacity and processing power of the storage system as the amount of data grows.

[0016] 2) This invention effectively improves the throughput and efficiency of massive image file storage systems through asynchronous and batch processing. Asynchronous processing fully utilizes resources, reduces latency, and enhances the system's concurrent processing capabilities. Batch processing consolidates similar requests, reducing request overhead and improving overall system performance.

[0017] 3) Compared with other complex object storage architectures, the massive image file architecture of the present invention is based on Minio, which may require more planning and management when expanding. Minio can easily add or delete server nodes, and data is automatically balanced between nodes. This design makes it very easy to expand. Storage capacity and I / O performance can be expanded by simply adding nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a schematic structural diagram of a wireless communication system for railway inspection equipment under extreme conditions according to an embodiment of the present invention. Modes for Carrying Out the Invention

[0019] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail through specific embodiments. However, it should be understood that the specific embodiments are provided only for a better understanding of the present invention and should not be construed as limiting the present invention. In the description of the present invention, it should be understood that the terms used are for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0020] To address the existing problems of processing only one request at a time, leading to resource waste because the system cannot share resources when processing multiple requests, increased request latency, and a negative impact on user experience, the present invention proposes a high-concurrency storage method, system, and readable medium for massive image files. This method, based on existing distributed storage, adds asynchronous processing and batch processing. In high-concurrency storage, asynchronous processing can be used to improve the throughput and performance of the storage system. Batch processing refers to the merging of multiple related processing tasks into a single batch for processing. In high-concurrency storage, batch processing can be used to merge multiple similar storage requests into a single batch for processing, thereby reducing request overhead and system load. This reduces unnecessary network communication and system overhead, improving the efficiency and performance of the storage system. Through asynchronous and batch processing, the throughput and efficiency of the massive image file storage system can be effectively improved. Asynchronous processing can fully utilize resources, reduce latency, and enhance the system's concurrent processing capabilities. Batch processing can merge similar requests, reducing request overhead and improving overall system performance. The present invention is described in detail below through examples with reference to the accompanying drawings.

[0021] Example 1

[0022] This embodiment discloses a high-concurrency storage method for massive image files, as shown in FIG1 , including the following steps:

[0023] S1 puts the user's storage request into the message queue and distributed cache, and processes the storage request asynchronously.

[0024] Asynchronous processing involves placing a storage request in a message queue and distributed cache instead of writing it directly to storage. The distributed cache can be configured with an appropriate expiration policy to ensure consistency between the data in the distributed cache and the data in the message queue. The storage request is asynchronously retrieved from the message queue and subsequent processing is performed. Meanwhile, the Kafka asynchronous storage service is responsible for asynchronously persisting cold data, ensuring sufficient resources for the high-performance middleware, enabling rapid responses to user requests and avoiding long wait times.

[0025] S2 takes the asynchronously processed storage request from the message queue and assigns the storage request to the thread pool.

[0026] The thread pool includes two or more threads. The specific number of threads can be set based on the system's processing capabilities and actual needs. The thread pool manages the thread lifecycle, avoiding the overhead of frequent thread creation and destruction, and evenly distributing storage tasks to each thread. In this embodiment, the thread pool uses load balancing strategies such as round-robin or least-load to evenly distribute storage tasks to each thread.

[0027] After the thread in the S3 thread pool takes the storage task from the message queue, it stores it according to the storage task.

[0028] S4 combines similar storage tasks in the thread pool into a batch request and sends it to the server at one time.

[0029] To improve efficiency, multiple storage requests are combined into a single batch request and sent to the Minio server at once. This approach reduces the number of network requests and improves data transmission efficiency. It also reduces the number of requests processed by the Minio server, reducing server load. In this embodiment, the service front-end uses a load balancer to evenly distribute storage tasks across servers, preventing excessive load on any one server. This embodiment monitors the system's operating status in real time, identifies and resolves system bottlenecks, and performs timely system optimization to ensure performance.

[0030] In this embodiment, the server is a Minio server or a Ceph server. However, the configuration and management of Ceph are relatively complex and have a certain threshold. The hardware requirements are high, and to ensure good performance, high-configuration hardware is required. The storage based on Minio has a simple architecture and is easy to deploy and operate and manage. The overall architecture has high performance and performs well for big data and AI workloads. Compared with other complex object storage architectures, the massive image file architecture based on Minio in this embodiment may require more planning and management when expanding. Minio can easily add or delete server nodes, and data will be automatically balanced between nodes. This design makes it very easy to expand. Storage capacity and I / O performance can be expanded by simply adding nodes.

[0031] Example 2

[0032] Based on the same inventive concept, this embodiment discloses a high-concurrency storage system for massive image files, including:

[0033] Asynchronous processing module, used to put the user's storage request into the message queue and distributed cache, and perform asynchronous processing on the storage request;

[0034] A multi-thread module is used to take the asynchronously processed storage request from the message queue and assign the storage request to a thread pool. The threads in the thread pool will obtain the storage task from the message queue and perform the storage operation according to the task;

[0035] A storage module, configured to store a storage task after a thread in the thread pool takes the storage task from the message queue;

[0036] The batch processing module is used to combine similar storage tasks in the thread pool into a batch processing request and send it to the server at one time.

[0037] The load balancing and persistence module is used to use a load balancer in front of the Minio service, which can evenly distribute requests to each Minio server to avoid excessive load on a certain server.

[0038] The storage system also includes a monitoring module, which is used to monitor the entire data storage process in real time.

[0039] Example 3

[0040] Based on the same inventive concept, this embodiment discloses a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the above-mentioned high-concurrency storage method for massive image files.

[0041] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0042] This application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0043] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0044] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be included within the scope of protection of the claims of the present invention. The above content is only a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the present application, which should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0046] Cross-reference to related applications:

[0047] This application claims priority to the prior Chinese patent application with application number 202311675038.4, which was filed on December 7, 2023 and is entitled: A high-concurrency storage method, system and readable medium for massive image files. The entire contents of the Chinese patent application are incorporated herein by reference. Industrial Applicability

[0048] The solution of the present invention can be widely used in the field of big data processing that requires massive image file processing. Currently, computer systems with such requirements are being used in many technical fields such as artificial intelligence monitoring. It can improve the throughput and efficiency of the system, batch process similar requests, reduce request overhead, and increase storage processing speed. As the amount of data grows, it can effectively and simply expand the capacity and processing power of the storage system, and has broad application prospects.

Claims

1. A high-concurrency storage method for massive image files, characterized in that: The following steps are involved: Put the user's storage request into the message queue and the distributed cache, and process the storage request asynchronously; Taking the asynchronously processed storage request out of the message queue and assigning the storage request to a thread pool; After the threads in the thread pool take out the storage tasks from the message queue, they store the tasks according to the storage tasks; Similar storage tasks in the thread pool are combined into a batch request and sent to the server at one time.

2. The high-concurrency storage method for massive image files according to claim 1, characterized in that: The server is a Minio server or a Ceph server.

3. The high-concurrency storage method for massive image files according to claim 1, characterized in that: The distributed cache sets an expiration policy so that the data in the distributed cache is consistent with the data in the message queue.

4. The high-concurrency storage method for massive image files according to claim 1, characterized in that: The asynchronous processing is to put the storage request into a message queue without processing it immediately after it is issued, asynchronously obtain the storage request from the message queue, and perform subsequent processing operations.

5. The high-concurrency storage method for massive image files according to claim 1, characterized in that: The thread pool includes two or more threads, and the thread pool manages the life cycle of the threads so that the storage tasks are evenly distributed to each thread.

6. The high-concurrency storage method for massive image files as claimed in claim 5, characterized in that: In the thread pool, storage tasks are evenly distributed to each thread through a load balancing strategy of polling or least load.

7. The high-concurrency storage method for massive image files according to claim 1, characterized in that: The front end of the server uses a load balancer to evenly distribute the storage tasks to each server.

8. A high-concurrency storage system for massive image files, characterized in that: include: An asynchronous processing module, used to put the user's storage request into the message queue and the distributed cache, and perform asynchronous processing on the storage request; A multithreading module, used for taking the asynchronously processed storage request out of the message queue and allocating the storage request to a thread pool; A storage module, used for storing according to the storage task after the thread in the thread pool takes out the storage task from the message queue; The batch processing module is used to combine similar storage tasks in the thread pool into a batch processing request and send it to the server at one time.

9. The high-concurrency storage system for massive image files as claimed in claim 8, characterized in that: The storage system also includes a monitoring module, which is used to monitor the entire data storage process in real time.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the high-concurrency storage method for massive image files as described in any one of claims 1-7.

Citation Information

Patent Citations

  • System for supporting high concurrent cache task queue and asynchronous batch operation method thereof

    CN103116634A

  • Regional traffic intelligent management system based on big data

    CN107729413A

  • Quasi-real-time asynchronous batch processing system, method and device and storage medium

    CN109582446A

  • Request processing method and system based on asynchronous request and interface batch calling

    CN117041351A

  • High-concurrency storage method and system for massive image files and readable medium

    CN117687571A

Cited By

  • Target detection and identification task processing method and device, electronic equipment and storage medium

    CN121053369A