A method, a processing device and a computing device for processing large files using threads

By using the first and second threads that run concurrently in the operating system to process large files, the problems of slow import speed and system instability caused by excessive memory usage in the prior art are solved, and more efficient large file data processing is achieved.

CN114510332BActive Publication Date: 2025-06-27UNIONTECH SOFTWARE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210111799.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-06-27
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

When processing super-large files, the prior art is prone to slow import speed, unstable system, and even crashes due to excessive memory usage.

Method used

By using the first and second threads running concurrently in the operating system, they are responsible for parsing and processing data sets in large files, respectively. The first thread parses the data set of a predetermined block size, obtains the field of interest, and sends it to the second thread for processing. The second thread only continues to process the next dataset after processing the field of interest in the current dataset.

Benefits of technology

It effectively controls the number of threads running synchronously, avoids excessive memory space due to running more threads, improves the reading and processing efficiency of large file data, and improves the performance of processing large files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114510332B_ABST
    Figure CN114510332B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, a processing device and a computing device for processing large files using threads. Among them, the method is executed in an operating system and includes the steps of: each time reading a data set of a predetermined block size from the file and starting a first thread to parse the data set of the predetermined block size to obtain fields of interest; and sending the fields of interest to a second thread that runs concurrently with the first thread to collect and process the fields of interest through the second thread. According to the technical solution of the present invention, it is possible to more reasonably utilize multiple threads to process large file data, improve the reading and processing efficiency of large file data while ensuring the stable operation of the system, and enhance the performance of processing large files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computers and operating systems, and particularly relates to a method, a processing device, and a computing device for processing large files using threads. Background Art

[0002] For an extremely large file containing hundreds of millions of lines of data, if one wants to process the data in the file and import the transformed data into a production database in a relatively short period of time, the prior art generally adopts methods such as splitting the file or multi-threaded import.

[0003] Among them, considering that when a program directly reads a large file, in the case of the program crashing suddenly, the progress of file reading will be lost and it is necessary to start reading again. By using the split command split to split the extremely large file into small files, for example, each file has 100,000 lines. After each small file is read, the small file can be moved to a specified folder. In this way, even if the application program crashes and restarts, when reading again, only the remaining files need to be read. And, after the file is split, a multi-node deployment method can be adopted for horizontal expansion, and each node reads a part of the file, so that the import speed can be doubled. However, the file splitting scheme is equivalent to loading all the data of the file into memory, resulting in excessive memory occupation.

[0004] The multi-threaded import method is a streaming reading method, reading data line by line, which requires creating too many threads and will also occupy too much memory. Once the memory occupation is excessive, not only will the import speed be too slow, resulting in the GC (Garbage Collector) mechanism being unable to clean the memory, but also the program or system will be unstable or even crash.

[0005] Therefore, a method for processing large files using threads is needed to solve the problems existing in the above technical solutions. Summary of the Invention

[0006] Therefore, the present invention provides a method and a processing device for processing large files using threads, in an attempt to solve or at least alleviate the problems existing above.

[0007] According to one aspect of the present invention, there is provided a method for processing a large file using threads, which is executed in an operating system and includes the steps of: each time reading a data set of a predetermined block size from the file, and starting a first thread to parse the data set of the predetermined block size to obtain fields of interest; and sending the fields of interest to a second thread that runs concurrently with the first thread to collect and process the fields of interest through the second thread.

[0008] Optionally, in the method for processing large files using threads according to the present invention, the dataset of the predetermined block size includes one or more lines of dataset, and the step of parsing the dataset of the predetermined block size includes: respectively parsing each line of dataset in the dataset of the predetermined block size, and sending the fields of interest in each line of dataset to one or more second threads running concurrently with the first thread, so as to collect and process the fields of interest in the one or more lines of dataset through the one or more second threads.

[0009] Optionally, in the method for processing large files using threads according to the present invention, the step of sending the fields of interest in each line of dataset to a second thread running concurrently with the first thread includes: respectively adding the fields of interest in each line of dataset to a structure, and respectively sending each structure to a second thread running concurrently with the first thread.

[0010] Optionally, in the method for processing large files using threads according to the present invention, the step of reading a dataset of the predetermined block size from the file each time and starting a first thread to parse the dataset of the predetermined block size includes: waiting until the one or more second threads have completed processing the fields of interest in the corresponding line of dataset in the dataset of the predetermined block size, and then starting a new first thread to parse the dataset of the predetermined block size read from the file next time.

[0011] Optionally, in the method for processing large files using threads according to the present invention, sending the fields of interest to a second thread running concurrently with the first thread includes: sending the fields of interest to a second thread running concurrently with the first thread through a data transmission channel.

[0012] Optionally, in the method for processing large files using threads according to the present invention, sending the fields of interest to a second thread running concurrently with the first thread includes: storing the fields of interest in shared memory and creating a mutex, so that each time a second thread accesses the fields of interest in the shared memory and processes them.

[0013] Optionally, in the method for processing large files using threads according to the present invention, when the second thread accesses the fields of interest in the shared memory, it is adapted to: request to acquire the mutex, if the mutex is acquired, access the fields of interest in the shared memory based on the mutex, and release the mutex after processing the fields of interest; if the mutex is not acquired, wait until the mutex is released and then request to acquire the mutex again.

[0014] Optionally, in the method for processing a large file using threads according to the present invention, before each time a data set of a predetermined block size is read from the file, it further includes the step of: determining whether all the data sets in the file have been read. If so, the final processing result is obtained based on the fields of interest processed by each second thread; if not, a data set of a predetermined block size is read from the file.

[0015] Optionally, in the method for processing a large file using threads according to the present invention, the thread is a goroutine.

[0016] According to one aspect of the present invention, there is provided a processing device resident in an operating system, including: a reading module adapted to read a data set of a predetermined block size from the file each time and start a first thread to parse the data set of the predetermined block size to obtain fields of interest; and a processing module adapted to send the fields of interest to a second thread running concurrently with the first thread to collect and process the fields of interest through the second thread.

[0017] According to one aspect of the present invention, there is provided a computing device including: at least one processor; and a memory storing program instructions, wherein the program instructions are configured to be executed by the at least one processor, and the program instructions include instructions for executing the method for processing a large file using threads as described above.

[0018] According to one aspect of the present invention, there is provided a readable storage medium storing program instructions, which, when read and executed by a computing device, cause the computing device to execute the method as described above.

[0019] According to the technical solution of the present invention, there is provided a method for processing a large file using threads. By separating the reading and parsing of the data set in the large file from the collection and processing of the data, and respectively executing them in parallel through the concurrently running first thread and second thread, this is beneficial to improving the reading and processing efficiency of the large file data. Further, in the present invention, after each row of the fields of interest in the currently read data set of a predetermined block size has been processed by one or more second threads, the parsing and processing of the next data set of a predetermined block size will continue to be executed. In this way, the number of threads running synchronously can be effectively controlled, avoiding occupying too much memory space and causing pressure on the CPU due to running too many threads. It can be seen that the present invention can more reasonably utilize multi-threads to process large file data, improve the reading and processing efficiency of large file data while ensuring the stable operation of the system, and enhance the performance of processing large files.

[0020] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented in accordance with the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To achieve the above and related purposes, certain illustrative aspects are described herein in connection with the following description and the accompanying drawings, which indicate various ways in which the principles disclosed herein may be practiced, and all aspects and their equivalent aspects are intended to fall within the scope of the claimed subject matter. By reading the following detailed description in conjunction with the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become more apparent. Throughout the present disclosure, like reference numerals generally refer to like components or elements.

[0022] Figure 1 FIG. shows a schematic diagram of a computing device 100 according to an embodiment of the present invention;

[0023] Figure 2 FIG. shows a flowchart of a method 200 for processing large files using threads according to an embodiment of the present invention; and

[0024] Figure 3 FIG. shows a schematic diagram of a processing device 300 according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0026] Figure 1 is a schematic block diagram of an exemplary computing device 100.

[0027] As Figure 1 shown, in the basic configuration 102, the computing device 100 typically includes a system memory 106 and one or more processors 104. A memory bus 108 can be used for communication between the processor 104 and the system memory 106.

[0028] Depending on the desired configuration, processor 104 can be any type of processor, including but not limited to: a microprocessor (UP), a microcontroller (UC), a digital information processor (DSP), or any combination thereof. Processor 104 can include one or more levels of cache, such as level one cache 110 and level two cache 112, a processor core 114, and registers 116. Exemplary processor core 114 can include an arithmetic logic unit (ALU), a floating point unit (FPU), a digital signal processing core (DSP core), or any combination thereof. Exemplary memory controller 118 can be used with processor 104, or in some implementations, memory controller 118 can be an internal part of processor 104.

[0029] Depending on the desired configuration, system memory 106 can be any type of memory, including but not limited to: volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.), or any combination thereof. System memory 106 can include an operating system 120, one or more applications 122, and program data 124. In some embodiments, applications 122 can be arranged to execute instructions on the operating system by one or more processors 104 using program data 124.

[0030] Computing device 100 further includes storage device 132, which includes removable storage 136 and non-removable storage 138.

[0031] Computing device 100 can also include a storage interface bus 134. Storage interface bus 134 enables communication from storage device 132 (e.g., removable storage 136 and non-removable storage 138) via bus / interface controller 130 to basic configuration 102. At least a portion of operating system 120, applications 122, and data 124 can be stored on removable storage 136 and / or non-removable storage 138, and when computing device 100 is powered on or an application 122 is to be executed, it is loaded into system memory 106 via storage interface bus 134 and executed by one or more processors 104.

[0032] The computing device 100 may also include an interface bus 140 that facilitates communication from various interface devices (e.g., output device 142, peripheral interface 144, and communication device 146) to the basic configuration 102 via the bus / interface controller 130. Example output devices 142 include an image processing unit 148 and an audio processing unit 150. They may be configured to facilitate communication with various external devices such as a display or speakers via one or more A / V ports 152. Example peripheral interfaces 144 may include a serial interface controller 154 and a parallel interface controller 156, which may be configured to facilitate communication with external devices such as input devices (e.g., keyboard, mouse, pen, voice input device, touch input device) or other peripherals (e.g., printer, scanner, etc.) via one or more I / O ports 158. Example communication device 146 may include a network controller 160, which may be arranged to facilitate communication with one or more other computing devices 162 via one or more communication ports 164 over a network communication link.

[0033] The network communication link may be an example of a communication medium. A communication medium can generally embody computer-readable instructions, data structures, program modules in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. A "modulated data signal" may be a signal in which one or more of its data sets or its changes can be encoded in the signal in a way that encodes information. As a non-limiting example, the communication medium may include wired media such as a wired network or a dedicated line network, and various wireless media such as sound, radio frequency (RF), microwave, infrared (IR), or other wireless media. The term computer-readable medium as used herein may include both storage media and communication media.

[0034] The computing device 100 may be implemented as a personal computer including desktop computer and notebook computer configurations. Of course, the computing device 100 may also be implemented as part of a small-sized portable (or mobile) electronic device, which may be such as a cellular phone, a digital camera, a personal digital assistant (PDA), a personal media player device, a wireless network browsing device, a personal head-mounted device, an application-specific device, or a hybrid device that may include any of the above functions. It may even be implemented as a server, such as a file server, a database server, an application server, and a WEB server, etc. Embodiments of the present invention are not limited thereto.

[0035] In an embodiment according to the present invention, the computing device 100 is configured to execute the method 200 for processing large files using threads according to the present invention. Among them, the operating system 120 of the computing device 100 contains multiple program instructions for executing the method 200 for processing large files using threads according to the present invention. These program instructions can instruct the processor to execute the method 200 for processing large files using threads according to the present invention, so that the computing device 100 executes the method 200 for processing large files using threads according to the present invention to improve the processing efficiency of large files.

[0036] According to an embodiment of the present invention, a processing device 300 is deployed in the operating system 120. The processing device 300 contains multiple program instructions for executing the method 200 for processing large files using threads according to the present invention, so that the method 200 for processing large files using threads according to the present invention can be executed in the processing device 300.

[0037] Figure 2 The flowchart of the method 200 for processing large files using threads according to an embodiment of the present invention is shown. The method 200 for processing large files using threads can be executed in the operating system of a computing device (such as the aforementioned computing device 100).

[0038] Figure 2 As shown, the method 200 for processing large files using threads starts from step S210.

[0039] In step S210, a data set of a predetermined block size is read from the file each time, and a first thread is started to parse the data set of the predetermined block size to obtain the fields of interest in the data set. Here, the fields of interest are target fields that need to be processed determined according to the actual application scenario. The present invention does not limit the specific types of the fields of interest. For example, when a user searches for a certain word in a large file, the searched word can be used as the field of interest. Another example is that when a user needs to replace or mark one or more words in a large file, these one or more words can be used as the fields of interest.

[0040] That is to say, by sequentially reading each line of the data set in the file and starting a first thread each time after reading a data set of a predetermined block size, the first thread is used to parse the fields of interest in the data set of the predetermined block size.

[0041] It should be noted that the dataset with a predetermined block size contains one or more lines of dataset. In the present invention, a first thread is started after reading one or more lines of dataset to parse the read dataset. Preferably, the dataset with a predetermined block size contains multiple lines of dataset. In this way, the present invention does not start a first thread for each line of dataset respectively, but starts a first thread for parsing after reading multiple lines of dataset, so as to minimize the number of created threads and avoid occupying too much memory space due to creating more threads.

[0042] In one implementation, the predetermined block size is, for example, 64K, but the present invention is not limited thereto. The predetermined block size can be set by those skilled in the art according to actual needs.

[0043] Subsequently, in step S220, the interesting fields parsed from the dataset with a predetermined block size are sent to a second thread that runs concurrently with the first thread, so as to collect and process the interesting fields through the second thread.

[0044] In one implementation, the first thread serves as the main thread, and the first thread can be implemented as a goroutine based on Golang. Correspondingly, the second thread is a goroutine that runs concurrently with the first thread.

[0045] It should be pointed out that according to the embodiments of the present invention, the reading and parsing of the dataset in the large file are separated from the collection and processing of the data, and are respectively executed in parallel by the first thread and the second thread that are decoupled from each other and run concurrently. In this way, it is beneficial to improve the efficiency of reading and processing the large file data.

[0046] According to an embodiment of the present invention, the dataset with a predetermined block size includes one or more lines of dataset. The specific implementation of parsing the dataset with a predetermined block size by the first thread can be:

[0047] Parse each line of the dataset with a predetermined block size respectively, and send the interesting fields in each line of the dataset to a second thread that runs concurrently with the first thread, so as to collect and process the interesting fields in one or more lines of the dataset through one or more second threads. That is to say, the interesting fields in the dataset with a predetermined block size are collected and processed through one or more second threads. Here, the interesting fields in each line of the dataset can be added to a structure respectively, and each structure is sent to a second thread that runs concurrently with the first thread. It should be pointed out that the present invention does not limit the specific implementation of the structure.

[0048] It should be understood that when reading a data set from a file, a first thread is started only after reading a data set of a predetermined block size each time. The first thread is used to parse the fields of interest in the data set of the predetermined block size. The data set of the predetermined block size contains one or more lines of data sets. The first thread sequentially parses each line of the data set of the predetermined block size, and after parsing each line of the data set, sends the fields of interest obtained by parsing each line of the data set to a second thread. That is to say, each second thread is responsible for collecting and processing the fields of interest in one line of the data set. In this way, the first thread sequentially sends the fields of interest in one or more lines of the data set contained in the data set of the predetermined block size to one or more second threads that run concurrently with the first thread, so as to collect and process the fields of interest in one or more lines of the data set through one or more second threads. Among them, each second thread collects and processes the fields of interest in the corresponding one line of the data set.

[0049] According to an embodiment of the present invention, each time a data set of a predetermined block size is read from a file, and a first thread is started to parse the data set of the predetermined block size, which is specifically executed according to the following method:

[0050] After waiting for one or more second threads to complete the collection and processing of the fields of interest in the corresponding one line of the data set of the predetermined block size (that is, after waiting for one or more second threads to complete the collection and processing of the fields of interest in the data set of the predetermined block size read from the file this time), start a new first thread to parse the data set of the predetermined block size read from the file next time. Furthermore, send the fields of interest obtained by parsing the data set of the predetermined block size read next time to one or more second threads that run concurrently with the new first thread, and each second thread collects and processes the fields of interest in each line of the data set.

[0051] In one implementation, a WaitGroup can be used to implement waiting for a group of second threads to finish execution, that is, use the WaitGroup to wait for one or more second threads to complete the collection and processing of the fields of interest in the corresponding one line of the data set of the predetermined block size. After that, a new first thread will be started to continue the parsing of the data set of the next predetermined block size.

[0052] According to an implementation of the present invention, after each line of the interesting fields in the currently read dataset of a predetermined block size has been processed by one or more second threads, the parsing and processing of the next dataset of a predetermined block size will continue to be executed downward. In this way, the number of threads running synchronously can be effectively controlled, thereby avoiding occupying too much memory space and causing pressure on the CPU due to running a large number of threads. By more reasonably using threads to process large file data, the reading and processing efficiency of large file data can be improved while ensuring the stable operation of the system.

[0053] In one embodiment, based on the mutually decoupled first thread and second thread, the interesting fields parsed by the first thread can be sent to the second thread running concurrently with the first thread through a data transmission channel.

[0054] In another embodiment, the method of call waiting can be used to process the interesting fields by calling one second thread each time. Specifically, the interesting fields parsed by the first thread are stored in the shared memory and a mutex is created, so that only one second thread can access the interesting fields in the shared memory based on the mutex each time, and process the interesting fields in the corresponding line of the dataset. Here, using the mutex can ensure that only one second thread can access the interesting fields in the shared memory and process the interesting fields in the corresponding line of the dataset each time, and ensure that multiple second threads execute the processing process of the interesting fields in sequence. In one implementation, the mutex can be created by the Mutex.Lock() method.

[0055] Further, when the second thread accesses the interesting fields in the shared memory, it first requests to obtain the mutex. If the mutex is obtained, it accesses the interesting fields in the shared memory based on the mutex and processes the interesting fields (the interesting fields in the corresponding line of the dataset), and releases the mutex after processing the interesting fields.

[0056] If the lock is not obtained, it means that the mutex is obtained and used by other second threads, and it waits until the mutex is released and then requests to obtain the mutex again. Until the mutex is obtained, it accesses the interesting fields in the shared memory based on the mutex and processes the interesting fields (the interesting fields in the corresponding line of the dataset).

[0057] In this way, each second thread can process the interesting fields in the corresponding line of the dataset in sequence, enabling multiple second threads to synchronously and orderly execute the processing of the interesting fields in each line of the dataset, and finally complete the processing of the interesting fields in the dataset of a predetermined block size.

[0058] In one implementation, the present invention uses a resource pool to manage multiple threads. For example, by calling async.Pool, threads are allocated for a dataset of a predetermined block size read each time, so as to parse or process the dataset through the threads. By setting the number of threads managed by the resource pool to a predetermined number, no new threads will be created after the number of threads in the resource pool reaches the predetermined number, and new threads will be created only after the existing threads end to perform data parsing or processing. In this way, excessive memory occupation in the case of high concurrency of multiple threads can be further avoided.

[0059] In one embodiment, before reading a dataset of a predetermined block size from a file each time, it is also determined whether all the datasets in the file have been read. If all have been read, the final processing result can be obtained based on the fields of interest processed by each second thread, and the final processing result is output. If not, continue to read the next dataset of a predetermined block size from the file.

[0060] Figure 3 FIG. 300 is a schematic diagram of a processing device according to an embodiment of the present invention. The processing device 300 resides in the operating system of a computing device (such as the aforementioned computing device 100) and is adapted to execute the method 200 for processing large files using threads of the present invention.

[0061] The processing device 300 includes a connected reading module 310 and a processing module 320. Among them, the reading module 310 reads a dataset of a predetermined block size from a file each time and starts a first thread to parse the dataset of the predetermined block size to obtain fields of interest. The processing module 320 sends the fields of interest to a second thread that runs concurrently with the first thread to collect and process the fields of interest through the second thread.

[0062] It should be noted that the reading module 310 is used to execute the foregoing step S210, and the processing module 320 is used to execute the foregoing step S220. Here, for the specific execution logics of the reading module 310 and the processing module 320, refer to the descriptions of steps S210 to S220 in the foregoing method 200, which will not be elaborated here.

[0063] Method 200 for processing large files using threads according to the present invention separates the reading and parsing of data sets in large files from the collection and processing of data, and executes them in parallel through a first thread and a second thread running concurrently. This is conducive to improving the efficiency of reading and processing large file data. Further, in the present invention, after each line of the interesting fields in the data set of the currently read predetermined block size has been processed by one or more second threads, the parsing and processing of the next data set of the predetermined block size will continue to be executed downward. In this way, the number of threads running synchronously can be effectively controlled, avoiding excessive memory space occupation and CPU pressure caused by running too many threads. It can be seen that the present invention can more reasonably utilize multiple threads to process large file data, improve the efficiency of reading and processing large file data while ensuring the stable operation of the system, and enhance the performance of processing large files.

[0064] The various technologies described herein can be implemented in combination with hardware or software, or a combination thereof. Thus, the method and apparatus of the present invention, or certain aspects or portions of the method and apparatus of the present invention, may take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, a USB flash drive, a floppy disk, a CD-ROM, or any other machine-readable storage medium, wherein when the program is loaded into a machine such as a computer and executed by the machine, the machine becomes an apparatus for practicing the present invention.

[0065] In the case where the program code is executed on a programmable computer, the computing device generally includes a processor, a processor-readable storage medium (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device. Among them, the memory is configured to store the program code; the processor is configured to execute the method for processing large files using threads of the present invention according to the instructions in the program code stored in the memory.

[0066] By way of example and not limitation, the readable medium includes a readable storage medium and a communication medium. The readable storage medium stores information such as computer-readable instructions, data structures, program modules, or other data. The communication medium generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery medium. A combination of any of the above is also included within the scope of the readable medium.

[0067] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems may also be used in conjunction with the examples of the present invention. Based on the above description, the structure required to construct such systems will be apparent. Additionally, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.

[0068] In the specification provided herein, numerous specific details are set forth. However, it can be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0069] Similarly, it should be understood that, in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present invention.

[0070] Those skilled in the art should understand that the modules or units or components of the devices in the examples disclosed herein may be arranged in the devices as described in the embodiments, or alternatively may be located in one or more devices different from those in the examples. The modules in the foregoing examples may be combined into one module or may be further divided into multiple sub-modules.

[0071] Those skilled in the art can understand that the modules in the devices of the embodiments can be adaptively changed and arranged in one or more devices different from those of the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition can be divided into multiple sub-modules or sub-units or sub-components. Except for the fact that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0072] In addition, those skilled in the art will appreciate that, although some of the embodiments described herein include certain features included in other embodiments and not others, combinations of features of different embodiments are meant to be within the scope of the present invention and form different embodiments.

[0073] In addition, some of the embodiments are described herein as a method or a combination of method elements that can be implemented by a processor of a computer system or by other devices performing the functions. Accordingly, a processor having the necessary instructions for implementing the method or method elements forms a means for implementing the method or method elements. In addition, the elements described herein of the apparatus embodiments are examples of the apparatus for performing the functions performed by the elements for the purpose of implementing the present invention.

[0074] As used herein, unless otherwise specified, the use of ordinal numbers such as "first", "second", "third", etc. to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects so described must have a given order in terms of time, space, ranking, or in any other manner.

[0075] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art within the technology will appreciate that other embodiments can be contemplated within the scope of the present invention as thus described. In addition, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and not for the purpose of explaining or limiting the subject matter of the present invention. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative, not restrictive, and the scope of the present invention is defined by the appended claims.

Claims

1. A method for processing large files using threads, which is executed in an operating system and includes the steps: Each time a data set of a predetermined block size is read from the file, a first thread is started to parse the data set of the predetermined block size to obtain fields of interest, where, The data set of the predetermined block size includes one or more lines of data sets; Parsing the data set of the predetermined block size includes: respectively parsing each line of the data set of the predetermined block size, and sending the fields of interest in each line of the data set to one or more second threads that run concurrently with the first thread, so as to collect and process the fields of interest in the one or more lines of the data set through the one or more second threads; after waiting for the one or more second threads to finish processing the fields of interest in the corresponding line of the data set of the predetermined block size, starting a new first thread to parse the data set of the predetermined block size read from the file next time; And Sending the fields of interest to a second thread that runs concurrently with the first thread, so as to collect and process the fields of interest through the second thread.

2. The method according to claim 1, wherein, The step of sending the fields of interest in each line of the data set to a second thread that runs concurrently with the first thread includes: Adding the fields of interest in each line of the data set to a structure respectively, and sending each structure to a second thread that runs concurrently with the first thread respectively.

3. The method according to claim 1 or 2, wherein Sending the fields of interest to a second thread that runs concurrently with the first thread includes: Sending the fields of interest to a second thread that runs concurrently with the first thread through a data transmission channel.

4. The method according to claim 1 or 2, wherein Sending the fields of interest to a second thread that runs concurrently with the first thread includes: Storing the fields of interest in a shared memory and creating a mutex, so that each time a second thread accesses the fields of interest in the shared memory and processes them.

5. The method according to claim 4, wherein, When accessing the fields of interest in the shared memory, the second thread is adapted to: Request to acquire the mutex. If the mutex is acquired, access the fields of interest in the shared memory based on the mutex, and release the mutex after processing the fields of interest; If the lock is not acquired, wait until the mutex is released and then request to acquire the mutex again.

6. The method according to claim 1 or 2, wherein, Before each time of reading the data set of the predetermined block size from the file, it further includes the step: Judging whether all the data sets in the file have been read. If so, obtaining the final processing result based on the fields of interest processed by each second thread; If not, reading the data set of the predetermined block size from the file.

7. The method according to claim 1 or 2, wherein The thread is a goroutine.

8. A processing device, which resides in an operating system and is adapted to process large files using threads. The processing device includes: A reading module, adapted to read a data set of a predetermined block size from the file each time, and start a first thread to parse the data set of the predetermined block size to obtain fields of interest, wherein the data set of the predetermined block size includes one or more lines of data sets; parsing the data set of the predetermined block size includes: respectively parsing each line of the data set of the predetermined block size, and sending the fields of interest in each line of the data set to one or more second threads running concurrently with the first thread, so as to collect and process the fields of interest in the one or more lines of data sets through the one or more second threads; the reading module is further adapted to: after the one or more second threads have all processed the fields of interest in the corresponding line of the data set of the predetermined block size, start a new first thread to parse the data set of the predetermined block size read next time from the file; and A processing module, adapted to send the fields of interest to a second thread running concurrently with the first thread, so as to collect and process the fields of interest through the second thread.

9. A computing device, comprising: At least one processor; And A memory storing program instructions, wherein the program instructions are configured to be executed by the at least one processor, and the program instructions include instructions for executing the method according to any one of claims 1-7.

10. A readable storage medium storing program instructions, which, when read and executed by a computing device, cause the computing device to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Data processing method and device, terminal equipment and computer storage medium

    CN111935535A