A computing resource expansion system based on data storage
By processing computing tasks in parallel on multiple storage nodes and combining task scheduling algorithms, the problems of low resource utilization and data delay in traditional computer systems are solved, efficient computing resource expansion and data consistency are achieved, and product information processing efficiency of warehouse supermarkets is improved.
Patent Information
- Application Number
- CN202410441194.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-04-12
AI Technical Summary
Traditional computer systems have problems such as single point failure, performance bottlenecks and low resource utilization when dealing with computing resource requirements. Especially in the case of slow data calculation and slow cargo matching speed in warehouse supermarkets, the existing technology is difficult to effectively solve.
A computing resource expansion system based on data storage is designed. By processing computing tasks in parallel on multiple storage nodes, combining task scheduling algorithms and data consistency maintenance, the coordinated operation between storage nodes and the balance of computing capabilities is achieved, and a suitable storage node is selected for calculation and data allocation is used to select a suitable storage node.
It improves the utilization rate of computing resources and system performance, avoids data delays and slow calculations, ensures data consistency and reliability, and optimizes the efficiency of product information processing and system scalability.
Smart Images

Figure CN118567834B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of commodity computing, and specifically provides a computing resource expansion system based on data storage. Background Art
[0002] In modern society, the demand for computing resources is increasing continuously. However, traditional computer systems have certain limitations in meeting this demand. Traditional computer systems usually adopt a centralized architecture, where computing tasks are processed by a main server, and end users can only access the main server through terminals. This architecture has some problems, such as single point of failure, performance bottlenecks, and low resource utilization.
[0003] To solve these problems, a new computing resource expansion system has been proposed. This system is based on data storage and distributes computing tasks among multiple storage nodes for parallel processing. Different from the traditional centralized architecture, this system makes full use of the computing power of storage nodes, improving the utilization rate of computing resources and the performance of the system.
[0004] The core idea of this system is to closely combine computing tasks and data storage. By running computing tasks on storage nodes, the overhead of data transmission can be avoided, and multiple tasks can be processed in parallel. In addition, storage nodes can cooperate in task processing to further improve the performance of the system.
[0005] With the rapid development of technologies and markets such as mobile Internet, Internet of Things, and AI computing, various fields need to be integrated with computers to form stronger competitiveness. Warehouse-style hypermarkets are different from traditional supermarkets and are managed using modern computer technology, that is, through barcodes on commodities for fast payment settlement and scientific and reasonable control of the purchase, sales, and inventory of commodities. This not only facilitates people's shopping but also greatly improves the sales management level of the mall. However, with the expansion of warehouse-style supermarkets, the data generated by the purchase, sales, and inventory of commodities increases, making it more difficult to centralize and calculate the data in the storage center, resulting in slow data calculation and slow goods matching speed. Therefore, it is necessary to design a computing resource expansion system based on data storage. Summary of the Invention
[0006] The purpose of the present invention is to provide a computing resource expansion system based on data storage to solve the problems raised in the above background art.
[0007] To solve the above technical problems, the present invention provides the following technical solution: A computing resource expansion system based on data storage, including a cloud storage system, characterized in that the cloud storage system includes a data plane and a metadata plane, wherein the data plane is used to store user data, and the metadata plane is used to store the metadata corresponding to the data. The user data includes the data volume and the access volume. The increase in the data volume and access volume of the user will cause an increase in the number of entries and query volume stored in the metadata plane. The scalability of the metadata plane will directly affect the scalability of the entire cloud storage system;
[0008] The metadata plane includes a metadata storage base, a data directory unit one linking the metadata storage base, and a data directory unit two linking the metadata storage base;
[0009] The data directory unit one stores product information and at the same time supports the formation of a product information directory and the formation of a product information file;
[0010] The data directory unit two stores a location information list of product files and additionally adds the transmission path of each product file, providing a path for each product file to convert the computing system. Each store has an independent computing system;
[0011] Set one store as one storage node. Each storage node needs to have sufficient computing power and storage capacity to support the processing of computing tasks and the storage of data. The design of the storage node should consider the balance between computing and storage, as well as the communication and collaborative processing between nodes. Set each storage node as A1, A2, A3,..., A N 。
[0012] According to the above technical solution, a task scheduling algorithm is set in each of the storage nodes, that is, the computing power in each storage node is set as C. Here, the computing power is the expected computable power. C is divided according to the operation load of 0-100%. 0 means the storage node has the strongest computing power, and 100% means the storage node has the weakest computing power. Different load standards are set during the operation division of 0-100%:
[0013] 0-60%: The computing power of the store is excellent and meets the standard for fully starting the task scheduling algorithm;
[0014] 60%-90%: The computing power of the store is fully loaded and meets the standard for partially starting the task scheduling algorithm;
[0015] 90%-100%: The computing power of the store is overloaded and does not meet the standard for starting the task scheduling algorithm;
[0016] The computing power of each store is a fixed value, with a maximum of 100%. When the computing power of the A1 storage node is in the range of 0 - 60%, it means that the A1 storage node has the ability to calculate the commodity files of the remaining storage nodes. When the computing power of the A1 storage node is in the range of 60% - 90%, it means that the A1 storage node has the ability to calculate part of the commodity files of the remaining storage nodes. When the computing power of the A1 storage node is in the range of 90% - 100%, it means that the A1 storage node does not have the ability to calculate the commodity files of the remaining storage nodes. The remaining storage nodes and the A1 storage node are based on the same standard of the task scheduling algorithm.
[0017] According to the above technical solution, each file is divided into several file blocks, each file block is stored in a surface folder, and several surface folders form a bottom folder. The file part is the surface folder, and the number of surface folders is selected according to the computable ability of the selected storage node.
[0018] According to the above technical solution, since the commodity information processing is carried out on multiple storage nodes, and there may be data dependency relationships between commodity information, data consistency maintenance is set in the task scheduling algorithm. Data consistency maintenance can ensure that the execution results of commodity information on different nodes are consistent, and ensure the correctness and reliability of the data. When the calculation result is returned through the transmission path, this calculation result is independently stored and backed up, and this independently stored and backed up calculation result can be obtained and used. When the computing power of its own storage node is lower than 60%, the calculation result is verified. Data loss and other situations may occur during the transmission process of the calculation result. If the verification result is correct, the calculation result is continued to be used. If the verification result is incorrect, the calculation result is recalculated through its own storage node, and the system notifies it.
[0019] According to the above technical solution, the data volume of the user data includes the identity information entered when the user enters the store and the user information calculated from the data surface. The identity information of each user includes the frequency of the user entering the store and the purchase record of the commodity. The user information calculated from the data surface includes the average frequency of the user entering the store, the time since the last entry into the store, and the expected purchase of the commodity;
[0020] The expected purchased products are calculated based on the purchased product records. The products purchased by the user are marked for the first time. A purchase habit database of the user is formed through the first seven purchase records. The products purchased by the user multiple times are marked for the second time, and the first mark is deleted. The first mark is the mark for the formation of the user's purchase habit, and the marked products are not completely the products the user likes to use. The products marked for the second time are the products repeatedly purchased by the user. When the user logs in to the store app, the products marked for the second time are displayed on the home page for the user to directly obtain the purchase channels. At the same time, products related to the products marked for the second time are recommended on the home page or the second page to ensure the maximization of the utilization rate of the product information on the home page. To ensure the stability of the computing power.
[0021] According to the above technical solution, the method of marking deletion is used to reduce the proportion of data occupying the computing power range. The marks to be deleted will ultimately be divided into multiple small marks, and the information of the cancelled users also needs to be deleted. The user information to be deleted is divided into multiple information blocks. The small marks and information blocks are stored in the garbage data area. The method of deleting one by one is adopted to shorten the length of the continuous garbage data area. We disperse the pressure of data deletion, split the garbage data area into multiple small pieces, and at the same time enhance the ability of a single storage node to process garbage data, alleviating its impact on the query performance. The length of the garbage data is reduced by two divisions to ensure that the deletion speed is fast each time and it is convenient to intersperse other data calculation tasks in the middle.
[0022] According to the above technical solution, during the process of deleting garbage data, if the computing power of the storage node does not exceed 60%, the garbage data deletion task is suspended and other calculation tasks are interspersed to maximize the convenience for the user. If the computing power of the storage node exceeds 60%, the garbage data deletion task is started to ensure that the computing power is not affected by the garbage data, significantly improving the performance of the system.
[0023] According to the above technical solution, when the user enters the store, the products marked for the second time will be pushed to the user's mobile terminal in the form of text messages, so that the products marked for the second time are always preferentially selected by the user.
[0024] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: In the present invention, by setting a task scheduling algorithm, the number of surface folders is selected according to the computable power of the selected storage nodes, realizing the collaborative operation between storage nodes, avoiding the data delay and slow calculation phenomena brought by the unified calculation of storage nodes. In order to realize the parallel processing of product information, the task scheduling algorithm of the present application is designed. This algorithm should consider factors such as the dependency relationship of product information, the load situation of storage nodes, and network latency to achieve the efficient allocation and scheduling of product information processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:
[0026] Figure 1 is a schematic diagram of the system of the present invention. Detailed implementation manners
[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] Please refer to Figure 1 , the present invention provides a technical solution: a computing resource expansion system based on data storage, including a cloud storage system, characterized in that the cloud storage system includes a data plane and a metadata plane. The data plane is used to store user data, and the metadata plane is used to store the metadata corresponding to the data. The user data includes the data volume and the access volume. The increase in the data volume and access volume of the user will cause an increase in the number of entries and query volume stored in the metadata plane. The scalability of the metadata plane will directly affect the scalability of the entire cloud storage system;
[0029] The metadata plane includes a metadata storage base, a data directory unit one linking the metadata storage base, and a data directory unit two linking the metadata storage base;
[0030] The data directory unit one stores commodity information and supports the formation of a commodity information directory and the formation of a commodity information file. For example: file / directory creation, search, deletion, and renaming, etc. A file is listed for the goods information of each store, and the files relying on the metadata storage base are displayed as L1, L2, L3,..., L N , where N is the total number of stores. By presenting the commodity information of a single store in an array, it is convenient for data search and data transmission;
[0031] The data directory unit two stores a list of the location information of the commodity files and additionally adds the transmission path of each commodity file, providing a path for each commodity file to convert the computing system. Each store has an independent computing system. Taking a store as a storage node, each storage node needs to have sufficient computing power and storage capacity to support the processing of computing tasks and the storage of data. The design of the storage node should consider the balance between computing and storage, as well as the communication and collaborative processing between nodes. Each storage node is set as A1, A2, A3,..., A N , L1, L2, L3,..., L NCorresponds one by one with A1, A2, A3, ..., A N One-to-one correspondence;
[0032] A task scheduling algorithm is set in each storage node, that is, the computing power in each storage node is set to C. Here, the computing power is the expected computable power. C is divided according to the operation load of 0-100%. 0 means the strongest computing power of the storage node, and 100% means the weakest computing power of the storage node. Different load standards are set during the operation division of 0-100%:
[0033] 0-60%: The computing power of the store is excellent and meets the standard for the full start of the task scheduling algorithm;
[0034] 60%-90%: The computing power of the store is fully loaded and meets the standard for the partial start of the task scheduling algorithm;
[0035] 90%-100%: The computing power of the store is overloaded and does not meet the standard for the start of the task scheduling algorithm;
[0036] The computing power of each store is a fixed value, with a maximum of 100%. When the computing power of the A1 storage node is within the range of 0-60%, it means that the A1 storage node has the ability to calculate the commodity files of the other storage nodes. When the computing power of the A1 storage node is within the range of 60%-90%, it means that the A1 storage node has the ability to calculate part of the commodity files of the other storage nodes. When the computing power of the A1 storage node is within the range of 90%-100%, it means that the A1 storage node does not have the ability to calculate the commodity files of the other storage nodes. The other storage nodes are consistent with the A1 storage node based on the standard of the task scheduling algorithm;
[0037] When the A1 storage node meets the standard for the full start of the task scheduling algorithm, the other storage nodes can, when their own computing power is higher than 60%, transfer the commodity files originally belonging to themselves to the storage nodes within the range of 0-60%. Then the storage node calculates the files and returns the calculation results along the transmission path;
[0038] If the situation where its own computing power is higher than 90% occurs and the computing power of all other storage nodes is higher than 60%, the task scheduling algorithm preferentially selects the storage nodes in the range of 60%-90%. Then, according to the strength of the current computing power of the storage nodes, it selects the storage node with the strongest computing power, transfers part of the files of its own storage node to this storage node, uses this storage node for calculation, and then returns the calculation results along the transmission path;
[0039] Each file is divided into several file blocks, and each file block is stored in a surface folder. Several surface folders constitute a bottom folder. The file part in the present invention is the surface folder. The number of surface folders is selected according to the computable ability of the selected storage nodes to achieve the collaborative operation among storage nodes, avoiding the data delay and slow calculation caused by the unified calculation of storage nodes. In order to achieve the parallel processing of commodity information, the task scheduling algorithm of this application is designed. This algorithm should consider factors such as the dependency relationship of commodity information, the load of storage nodes, and network latency to achieve the efficient allocation and scheduling of commodity information processing.
[0040] Since the processing of commodity information is distributed on multiple storage nodes and there may be data dependency relationships among commodity information, data consistency maintenance is set in the task scheduling algorithm. Data consistency maintenance can ensure that the execution results of commodity information on different nodes are consistent and guarantee the correctness and reliability of data. When the calculation result is returned through the transmission path, this calculation result is independently stored and backed up. This independently stored and backed up calculation result can be obtained and used. When the computing power of its own storage node is lower than 60%, the calculation result is verified. Data loss and other situations may occur during the transmission of the calculation result. If the verification result is correct, the calculation result is continued to be used. If the verification result is incorrect, the result is recalculated through its own storage node and notified by the system.
[0041] The data volume of user data includes the identity information entered when the user enters the store and the user information calculated from the data surface. The identity information of each user includes the frequency of the user entering the store and the purchase record of goods. The user information calculated from the data surface includes the average frequency of the user entering the store, the time since the last entry into the store, and the estimated purchase of goods.
[0042] The expected purchased products are calculated based on the purchased product records. The products purchased by the user are marked for the first time. The purchase habit database of the user is formed through the first seven purchase records. The products purchased by the user multiple times are marked for the second time, and the first mark is deleted. The first mark is the mark for the formation of the user's purchase habit. The marked products are not completely the products that the user likes to use. The products marked for the second time are the products that the user repurchases. When the user logs in to the store app, the products marked for the second time are displayed on the home page for the user to directly obtain the purchase channels. At the same time, the products related to the products marked for the second time are recommended on the home page or the second page to ensure the maximum utilization rate of the product information on the home page. To ensure the stability of the computing power, the method of mark deletion is used to reduce the proportion of the data occupying the computing power range. The marks to be deleted will ultimately be divided into multiple small marks. The information of the cancelled users also needs to be deleted. The user information to be deleted is divided into multiple information blocks. The small marks and information blocks are stored in the garbage data area. The method of deleting one by one is used to shorten the length of the continuous garbage data area. We disperse the pressure of data deletion, split the garbage data area into multiple small pieces, and at the same time enhance the ability of a single storage node to process garbage data to alleviate its impact on the query performance. The length of the garbage data is reduced by two divisions to ensure that the deletion speed is fast each time and it is convenient to intersperse other data calculation tasks in the middle.
[0043] During the process of deleting garbage data, if the computing power of the storage node does not exceed 60%, the garbage data deletion task is suspended and other calculation tasks are interspersed to maximize the convenience for the user. If the computing power of the storage node exceeds 60%, the garbage data deletion task is started to ensure that the computing power is not affected by the garbage data and significantly improve the performance of the system.
[0044] When the user enters the store, the products marked for the second time will be pushed to the user's mobile terminal in the form of text messages, so that the products marked for the second time are always preferentially selected by the user.
[0045] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0046] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A computing resource expansion system based on data storage, including a cloud storage system, characterized in that, The cloud storage system includes a data plane and a metadata plane. The data plane is used to store user data, and the metadata plane is used to store the metadata corresponding to the data. The user data includes the data volume and the access volume. The increase in the data volume and access volume of the user will cause an increase in the number of entries and query volume stored in the metadata plane. The scalability of the metadata plane will directly affect the scalability of the entire cloud storage system; The metadata plane includes a metadata storage base, a data directory unit one linking the metadata storage base, and a data directory unit two linking the metadata storage base; The data directory unit one stores commodity information and supports the formation of a commodity information directory and the formation of a commodity information file; The data directory unit two stores a list of location information of commodity files and additionally adds the transmission path of each commodity file, providing a path for each commodity file to the computing system that can be converted. Each store has an independent computing system; Set a store as a storage node. Each storage node needs to have sufficient computing power and storage capacity to support the processing of computing tasks and the storage of data. The design of the storage node should take into account the balance between computing and storage, as well as the communication and collaborative processing between nodes. Set each storage node as A1, A2, A3, ..., A N; A task scheduling algorithm is set in each storage node, that is, the computing power in each storage node is set to C. Here, the computing power is the expected computable power. C is divided according to the operation load of 0-100%. 0 means the storage node has the strongest computing power, and 100% means the storage node has the weakest computing power. Different load standards are set during the operation division of 0-100%: 0-60%: The computing power of the store is excellent, meeting the standard for the full start of the task scheduling algorithm; 60%-90%: The computing power of the store is fully loaded, meeting the standard for the partial start of the task scheduling algorithm; 90%-100%: The computing power of the store is overloaded and does not meet the standard for the start of the task scheduling algorithm; The computing power of each store is a fixed value, with a maximum of 100%. When the computing power of the A1 storage node is within the range of 0-60%, it means that the A1 storage node has the ability to compute the commodity files of the other storage nodes. When the computing power of the A1 storage node is within the range of 60%-90%, it means that the A1 storage node has the ability to compute part of the commodity files of the other storage nodes. When the computing power of the A1 storage node is within the range of 90%-100%, it means that the A1 storage node does not have the ability to compute the commodity files of the other storage nodes. The other storage nodes are consistent with the A1 storage node based on the standard of the task scheduling algorithm; Each file is divided into several file blocks, and each file block is stored in a surface folder. Several surface folders form a bottom folder. The file part, that is, the surface folder, selects the number of surface folders according to the computable power of the selected storage node; Since the processing of commodity information is carried out on multiple storage nodes, and there may be data dependencies between commodity information, data consistency maintenance is set in the task scheduling algorithm. Data consistency maintenance can ensure that the execution results of commodity information on different nodes are consistent, and guarantee the correctness and reliability of the data. When the calculation result is returned through the transmission path, this calculation result is independently stored and backed up. This independently stored backup of the calculation result can be obtained and used. When the computing power of its own storage node is lower than 60%, the calculation result is verified. Data loss and other situations may occur during the transmission of the calculation result. If the verification result is correct, the calculation result is continued to be used. If the verification result is incorrect, the result is recalculated through its own storage node, and the system notifies it; The data volume of the user data includes the identity information entered when the user enters the store and the user information calculated from the data surface. The identity information of each user includes the frequency of the user entering the store and the purchase record of the commodity. The user information calculated from the data surface includes the average frequency of the user entering the store, the time since the last entry into the store, and the expected purchase of the commodity; The expected purchase of the commodity is calculated based on the purchase record. The commodities purchased by the user are marked for the first time. The purchase habit database of the user is formed through the first seven purchase records. The commodities purchased by the user multiple times are marked for the second time, and the first mark is deleted. The first mark is the mark for the formation of the user's purchase habit. The marked commodities are not completely the commodities that the user likes to use. The commodities marked for the second time are the commodities repeatedly purchased by the user. When the user logs in to the store app, the commodities marked for the second time are displayed on the home page for the user to directly obtain the purchase channel. At the same time, the commodities related to the commodities marked for the second time are recommended on the home page or the second page to ensure the maximum utilization rate of the commodity information on the home page. To ensure the stability of the computing power; The method of mark deletion is used to reduce the proportion of data occupying the computing power range. The marks to be deleted will ultimately be divided into multiple small marks. The information of the cancelled users also needs to be deleted. The user information to be deleted is divided into multiple information blocks. The small marks and information blocks are stored in the garbage data area. The method of deleting one by one is used to shorten the length of the continuous garbage data area. We disperse the pressure of data deletion, split the garbage data area into multiple small pieces, and at the same time enhance the ability of a single storage node to process garbage data, relieve its impact on the query performance, and use two divisions to reduce the length of the garbage data to ensure that the deletion speed is fast each time and it is convenient to intersperse other data calculation tasks; During the process of deleting garbage data, if the computing power of the storage node does not exceed 60%, the garbage data deletion task is suspended and other calculation tasks are interspersed to maximize the convenience for users. If the computing power of the storage node exceeds 60%, the garbage data deletion task is started to ensure that the computing power is not affected by the garbage data and significantly improve the performance of the system; When the user enters the store, the commodities marked for the second time will be pushed to the user's mobile terminal in the form of text messages, so that the commodities marked for the second time are always preferentially selected by the user.
Citation Information
Patent Citations
Method and system for dynamically managing metadata of distributed file system
CN101697168A
Cloud storage system and implementation method thereof
CN102307221A