Large-data-volume data storage method
By receiving and analyzing data in the data storage method of large data volume, using the number of threads and page records to form a discriminant formula, and choosing a suitable storage method, the problem of slow storage of large data volumes in the prior art is solved, and the rapid storage and efficient storage of large data volumes are achieved.
Patent Information
- Application Number
- CN202510108869.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is slow in the storage process of large data volumes, and lacks a unified storage method that can quickly save large data volumes.
By receiving and analyzing the data to be stored, the system's number of threads and page records are used to form a discriminant formula, the amount of data to be stored is judged, and the storage method is selected according to the size. The specific steps include receiving and parsing data, forming a discriminant formula, calculating the amount of data written to each thread, and writing in parallel through functions mmap and OpenMP.
It realizes the rapid storage of large data volumes, can receive data transmitted by 10 Gigabits of networks and performs rapid processing, improving the speed of traditional data storage at least 5 times.
Smart Images

Figure CN119988024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular to a method for storing large amounts of data. Background Art
[0002] The host computer software can obtain the collected data through various means, such as serial port, network (TCP / UDP protocol), PCIE interface and other communication methods. The communication speed of the serial port is about 150MB / s, the network transmission Gigabit broadband is about 128MB / S, the maximum transmission speed of the 10G network cable is 1.2GB / S, and the transmission speed of the PCIE 3.0 x4 port is about 4GB / s. The data transmission speed and data transmission method of each communication method are different.
[0003] The original data received by the host computer is usually large-volume unstructured data. The existing saved data is processed in a streaming manner and saved by calling the system function write. The processing is simple but the single data processing speed is slow.
[0004] In summary, how to quickly save data and provide a unified storage method has become a technical problem that current R&D personnel need to solve. Summary of the invention
[0005] In order to solve the above problems in the prior art, the present invention provides a method for storing large amounts of data, which solves the problem of slow storage of large amounts of data in the prior art.
[0006] In order to achieve the above object, the present invention adopts the following technical scheme: A method for storing large amounts of data, comprising the following steps: S1, receiving and parsing the data to be stored, and putting the parsed data to be stored into the memory queue; S2, using the number of system threads and the number of page records to form a discriminant for determining the amount of data to be stored; S3, when the amount of the data to be stored is greater than the discriminant, the amount of data written by each thread of the calculation system; when the amount of the data to be stored is not greater than the discriminant, the system calls a write function to write the data to be stored; S4. The system calls the corresponding function to write the data to be stored according to the calculation result in step S3.
[0007] In some embodiments of the present invention, the discriminant in step S2 is: Δ=total number of threads × number of page records & start flag.
[0008] In some embodiments of the present invention, the number of threads of the system is N, and the amount of data to be stored is divided into a first data amount and a second data amount, wherein the first data amount is written through the N threads of the system; and the second data amount is written through the main thread of the system.
[0009] In some embodiments of the present invention, the calculation formula of the first data amount is: Size1=sum-sum%(N×pagesize); Wherein, sum is the total amount of data to be stored; pagesize is the number of page records, % is the remainder calculation symbol, and N is the number of threads; The first amount of data is allocated to each thread. 1_N for: Size 1_N =Size1 / N.
[0010] In some embodiments of the present invention, the calculation formula of the second data amount is: Size2 = sum – Size1.
[0011] In some embodiments of the present invention, the receiving and parsing of the data to be stored in step S1 specifically includes: Receive the data to be stored by using multiple communication methods; and set the start flag of the data to be stored to 1; Perform protocol analysis.
[0012] In some embodiments of the present invention, the multiple communication modes include but are not limited to perforation, network port or PCIE.
[0013] In some embodiments of the present invention, step S4 specifically includes: Mapping the first amount of data to be stored into the system through the function mmap; Open up N threads through OpenMP; Each thread copies the corresponding data to the system mapped by the mmap function through the memcpy function.
[0014] In some embodiments of the present invention, a write function of the system is called for the data of the main thread to complete the writing of the second data volume.
[0015] In some embodiments of the present invention, the discriminant in step S2 is: Δ=n×total number of threads×number of page records&start flag; The function mmap is adapted to integer multiples of the number of page records.
[0016] The technical solution of the present invention has the following technical effects compared with the prior art: The data saving method of the present invention can receive data transmitted by a 10G network, perform calculations before receiving the data, and select a saving method according to the size of the data, thereby realizing rapid saving of large amounts of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0018] Figure 1 The figure is a flow chart of a method for storing large amounts of data described in an embodiment.
[0019] Figure 2 It is a structural schematic diagram of the electronic device.
[0020] Reference numerals: 100, electronic device; 110, processor; 120, memory; 130, transceiver. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0022] In the description of the present application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0023] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0024] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0025] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may include that the first and second features are in direct contact, or may include that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, a first feature being "above", "above" and "above" a second feature includes that the first feature is directly above and obliquely above the second feature, or simply indicates that the first feature is higher in level than the second feature. A first feature being "below", "below" and "below" a second feature includes that the first feature is directly below and obliquely below the second feature, or simply indicates that the first feature is lower in level than the second feature.
[0026] The disclosure below provides many different embodiments or examples to implement different structures of the present invention. In order to simplify the disclosure of the present invention, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present invention. In addition, the present invention can repeat reference numbers and / or reference letters in different examples, and this repetition is for the purpose of simplification and clarity, which itself does not indicate the relationship between the various embodiments and / or settings discussed.
[0027] Example 1: Reference Figure 1 As shown, a method for storing large amounts of data includes the following steps: S1, receiving and parsing the data to be stored, and putting the parsed data to be stored into the memory queue; Specifically, the host computer software receives data through a serial port, a network port or a PCIE, and sets the received data flag StartFlag to 1; Then the received data is parsed by protocol and the parsed data is put into the memory queue.
[0028] S2, using the number of system threads and the number of page records to form a discriminant for determining the amount of data to be stored; Determine the amount of data in an asynchronous manner; Use the discriminant Δ=N×pagesize&startflag; Among them, N is the total number of threads in the system, that is, the maximum number of threads supported by the current system; pagesize is the number of page records, and the default value is 4096 bytes.
[0029] First, determine whether the size of the data to be stored is greater than N×pagesize; then perform an AND operation on the determination result and startflag.
[0030] S3, when the amount of the data to be stored is greater than the discriminant, the amount of data written by each thread of the calculation system; when the amount of the data to be stored is not greater than the discriminant, the system calls a write function to write the data to be stored; Specifically, when the amount of data to be stored is not greater than the above discriminant, that is, when the current amount of data to be stored is insufficient to use all threads in the system to write records page by page, the system directly calls the write function write to append data to the end of the file.
[0031] When the amount of data to be stored is greater than the above discriminant, it means that the current amount of data to be stored is a large amount of data, and all threads of the system can be started to write data.
[0032] The amount of data allocated to each thread is calculated using the following formula: Assume that the system includes N threads, and the amount of data to be stored is divided into a first data amount Size1 and a second data amount Size2, wherein the first data amount Size1 is written through the N threads of the system; the second data amount Size2 is written through the main thread of the system, wherein the main thread is the 0th thread among the N threads.
[0033] The calculation formula of the first data volume Size1 is: Size1=sum-sum%(N×pagesize); Where N is the total number of threads in the system; pagesize is the number of page records, % is the remainder calculation symbol, and sum is the total amount of data to be stored; The first amount of data is allocated to each thread. 1_N for: Size 1_N =Size1 / N.
[0034] That is to say, the amount of data written by N threads is an integer multiple of pagesize. During the calculation process, the total amount of data to be stored is first divided by N to calculate the initial write amount of each thread from the 0th thread to the N-1th thread. In order to further speed up the writing speed, the write amount of each thread from the 1st thread to the N-1th thread is controlled to an integer multiple of the number of page records, that is, the initial write amount is used for division and remainder calculation, thereby obtaining the first data amount Size1.
[0035] During the data writing process, threads 0 to N-1 simultaneously write data according to Size. 1_N The amount of data to be written.
[0036] In addition, when judging the size of the data to be stored, the total number of threads N and the integer multiple of the number of page records can also be used as a discriminant, that is, the discriminant Δ=n×N×pagesize&startflag is used; Among them, n is a constant greater than zero, N is the total number of threads in the system, that is, the maximum number of threads supported by the current system; pagesize is the number of page records, and the default value is 4096 bytes.
[0037] First, determine whether the size of the data to be stored is greater than n×N×pagesize; then perform an AND operation on the determination result and startflag.
[0038] During the data writing process, the mmap function uses an integer multiple of the number of page records for writing and saving.
[0039] For the Nth thread, if the amount of data to be stored is a remainder obtained after division in the above process of calculating Size1, the second data amount Size2 represented by the remainder is the amount of data that the 0th thread needs to write.
[0040] The calculation formula of the second data volume is: Size2 = sum – Size1.
[0041] In some other embodiments, in the process of determining the size of the data to be stored, if the determination result is less than N×pagesize, it means that the data to be stored does not belong to a large amount of data. For the data to be stored in this embodiment, the system's write function is directly called to append data to the end of the file.
[0042] S4. The system calls the corresponding function to write the data to be stored according to the calculation result in step S3.
[0043] The data size calculated in step S3 1_N , take out the corresponding amount of data from the data to be stored, and map the data size through the function mmap 1_N The data to be stored in the system; Open up N threads through OpenMP; Each thread copies the corresponding data to the system mapped by the mmap function through the memcpy function to complete the writing.
[0044] For example, when the maximum number of threads allocated to the CPU is 10, that is, N=10, the size of the data allocated to be written by each thread is: The amount of data allocated to threads 0 to 9 is the data size 1_N Thread 0 continues to be allocated the second data amount Size2.
[0045] Each thread performs write operations in parallel.
[0046] In some other embodiments, the writing speed of the final file will also be limited by the disk IO, so the disk IO speed is required to be greater than the speed of receiving data. Generally, a disk supporting PCIE 4.0 x4 is selected, which basically meets the data transmission speed of the above-listed cases.
[0047] The technical solution of the present invention has the following technical effects compared with the prior art: The data storage method of the present invention can receive data transmitted via a 10 Gigabit network and data transmitted via a PCIE interface, that is, can receive data with a relatively fast data transmission speed.
[0048] Before receiving the data, calculation is performed first, and the saving method is selected according to the size of the data; the discriminant can be N×pagesize or an integer multiple thereof; thereby realizing fast saving of large amounts of data.
[0049] The parallel writing method through the mmap function is also not limited to data that is an integer multiple of N×pagesize. Because the mmap function operation is only applicable to data operations that are an integer multiple of the number of page records, the excess data is written to the file by calling the system write function to complete the entire file writing operation.
[0050] In actual use, the storage method of this embodiment can increase the data storage speed to at least 5 times that of traditional data.
[0051] Example 2: Reference Figure 2 As shown, in this embodiment, an electronic device 100 is provided, including: A processor 110, and a memory 120 and a transceiver 130 communicatively connected to the processor; The memory 120 stores computer-executable instructions; the transceiver 130 is used to send and receive data; The processor 110 executes the computer-executable instructions stored in the memory 120 to implement the saving method in Embodiment 1.
[0052] It should be understood that the electronic device 100 can be used to execute the corresponding steps and / or processes in the above method embodiments. Optionally, the memory 120 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory 120 may also include a non-volatile random access memory. For example, the memory 120 may also store information about the device type. The processor 110 may be used to execute the instructions stored in the memory 120, and when the processor 110 executes the instructions, the processor 110 may execute the corresponding steps and / or processes in the above method embodiments.
[0053] It should be understood that in the embodiment of the present application, the processor 110 may be a central processing unit (CPU), and the processor 110 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0054] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 110 or an instruction in the form of software. The steps of the method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in the processor 110 for execution. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor executes the instructions in the memory, and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.
[0055] Embodiment 3: In this embodiment, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the saving method in Embodiment 1.
[0056] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0057] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0058] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0059] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0060] In the description of the above embodiments, specific features, structures, materials or characteristics may be combined in a suitable manner in any one or more embodiments or examples.
[0061] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for storing large amounts of data, characterized in that: The following steps are involved: S1, receiving and parsing the data to be stored, and putting the parsed data to be stored into the memory queue; S2, using the number of system threads and the number of page records to form a discriminant for determining the amount of data to be stored; S3, when the amount of the data to be stored is greater than the discriminant, the amount of data written by each thread of the calculation system; When the amount of the data to be stored is not greater than the discriminant, the system calls a write function to write the data to be stored; S4. The system calls the corresponding function to write the data to be stored according to the calculation result in step S3.
2. The method for storing large amounts of data according to claim 1, characterized in that: The discriminant in step S2 is: Δ=total number of threads × number of page records & start flag.
3. The method for storing large amounts of data according to claim 1, characterized in that: The number of threads of the system is N, and the amount of data to be stored is divided into a first amount of data and a second amount of data, wherein the first amount of data is written through the N threads of the system; and the second amount of data is written through the main thread of the system.
4. The method for storing large amounts of data according to claim 1, characterized in that: The calculation formula of the first data volume is: Size1=sum-sum%(N×pagesize); Wherein, sum is the total amount of data to be stored; pagesize is the number of page records, % is the remainder calculation symbol, and N is the number of threads; The first amount of data is allocated to each thread. 1_N for: Size 1_N =Size1 / N。 5. The method for storing large amounts of data according to claim 1, characterized in that: The calculation formula of the second data volume is: Size2 = sum – Size1.
6. The method for storing large amounts of data according to claim 1, characterized in that: The receiving and parsing of the data to be stored in step S1 specifically includes: Receive the data to be stored by using multiple communication methods; and set the start flag of the data to be stored to 1; Perform protocol analysis.
7. The method for storing large amounts of data according to claim 1, characterized in that: The multiple communication modes include but are not limited to perforation, network port or PCIE.
8. The method for storing large amounts of data according to claim 1, characterized in that: The step S4 specifically includes: Mapping the first amount of data to be stored into the system through the function mmap; Open up N threads through OpenMP; Each thread copies the corresponding data to the system mapped by the mmap function through the memcpy function.
9. The method for storing large amounts of data according to claim 3, characterized in that: The write function of the system is called for the data of the main thread to complete the writing of the second data volume.
10. The method for storing large amounts of data according to claim 1, characterized in that: The discriminant in step S2 is: Δ=n×total number of threads×number of page records&start flag; The function mmap is adapted to integer multiples of the number of page records.
Citation Information
Patent Citations
I<2>C bus device with big data master device transmission function and communication method thereof
CN105718396A
Data dump method and system for database
CN114896335A
Data fusion method and system for multiple data sources in cross-cloud scene
CN119149614A
Storage system and control method of storage system
US20130132641A1
Methods, apparatuses, and computer program products for controlling write requests in storage system
US20190220201A1