Grid file parallel processing method and device, computer equipment and storage medium

By dividing data into segments according to file size and performing header information identification and completion in large-scale Fluent mesh file processing, combined with inter-process communication, the problem of low I/O efficiency in existing technologies is solved, efficient parallel processing is achieved, and the preprocessing efficiency of aero-engine simulation is improved.

CN121029104AActive Publication Date: 2025-11-28AERO ENGINE ACAD OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511563060.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2025-11-28
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing technologies suffer from low I/O efficiency and flow limitations in parallel file systems when processing large-scale Fluent mesh files due to repeated traversal mechanisms, which prevent the effective utilization of parallel computing resources and severely impact simulation process efficiency.

Method used

By dividing the grid file into multiple data segments according to file size, and performing header information identification and completion in each process, and combining inter-process communication to build a global header information set, a one-time large-block data reading and integration can be achieved.

Benefits of technology

It significantly reduces disk access frequency and data throughput pressure, shortens data loading time, avoids file system throttling and network congestion, and improves parallel scalability and preprocessing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029104A_ABST
    Figure CN121029104A_ABST
Patent Text Reader

Abstract

The invention relates to a grid file parallel processing method and device, computer equipment and a storage medium. Comprising the following steps: respectively controlling a plurality of target processes to read corresponding file data segments, and controlling each target process to identify header information contained in the read file data segments after successful reading; for any file data segment, when the header information in the file data segment is determined to be incomplete according to the header information identification result, complementing the header information; controlling communication among the plurality of target processes, and constructing a global header information set according to a communication result and a header information identification result of each target process; and according to the global header information set, analyzing the grid data in the file data segment of each target process to obtain the grid data of each file data segment, and integrating the grid data of each file data segment to obtain the target grid file, so that a mechanism for avoiding repeated traversal greatly shortens the data loading time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of grid processing, and particularly relates to a parallel processing method and device of a grid file, a computer device and a storage medium. BACKGROUND

[0002] With the in-depth application of Computational Fluid Dynamics (CFD) in the design and optimization of an aero-engine, higher requirements are put forward for the scale and complexity of the computational grid.

[0003] In the related art, the entire Fluent grid file can be read first, and the information starting position and data length of all header information are identified and recorded; then each process calculates the file offset and data segment range to be read according to the allocated data segment; finally, the corresponding data is read into the memory of each process through parallel file I / O. In this process, in order to identify and locate the header information of the data segment, at least two complete I / O operations must be performed on the Fluent grid file: the first time is globally scanned by the master process to extract the header information, and the second time is to perform actual data reading according to the header information by each process. This repeated traversal mechanism leads to a significant reduction in I / O efficiency. SUMMARY

[0004] Therefore, the parallel processing method and device of a grid file, the computer device and the storage medium are provided to solve the problems in the related art.

[0005] In a first aspect, the present disclosure provides a parallel processing method of a grid file, the method comprising: obtaining a file size of a grid file to be read, and determining a plurality of target process data segments to be read according to the file size; respectively controlling a plurality of target processes to read the corresponding file data segments, and after reading successfully, controlling each target process to identify the header information contained in the read file data segment; for any file data segment, when it is determined that the header information in the file data segment is incomplete according to the identification result of the header information, the header information is completed; controlling the communication between the plurality of target processes, and constructing a global header information set according to the communication result and the header information identification result of each target process; according to the global header information set, the grid data in the file data segment of each target process is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain a target grid file.

[0006] In a second aspect, the present disclosure provides a parallel processing device for a grid file, which is applied to the parallel processing selection method for the grid file as described in the first aspect. The device comprises: an obtaining module configured to obtain a file size of a grid file to be read, and determine file data segments to be read by a plurality of target processes according to the file size; a control module configured to control the plurality of target processes to read corresponding file data segments respectively, and control each target process to identify header information contained in the read file data segment after successful reading; a completion module configured to, for any file data segment, complete the header information when it is determined that the header information in the file data segment is incomplete according to the identification result of the header information; a construction module configured to control communication between the plurality of target processes, and construct a global header information set according to the communication result and the identification result of the header information of each target process; and an analysis module configured to analyze grid data in the file data segment of each target process according to the global header information set, to obtain grid data of each file data segment, and integrate the grid data of each file data segment to obtain a target grid file.

[0007] In a third aspect, the present disclosure provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the parallel processing method for the grid file.

[0008] In a fourth aspect, the present disclosure provides a computer readable storage medium, which stores a computer program / instruction. The computer program / instruction is executed by a processor to implement the steps of the parallel processing method for the grid file.

[0009] In a fifth aspect, the present disclosure provides a computer program product, which stores a computer program / instruction. The computer program / instruction is executed by a processor to implement the steps of the parallel processing method for the grid file.

[0010] The above at least one technical solution adopted by the embodiments of the present disclosure can achieve the following beneficial effects: the file size of a grid file to be read is obtained, and file data segments to be read by a plurality of target processes are determined according to the file size; the plurality of target processes are controlled to read corresponding file data segments respectively, and each target process is controlled to identify header information contained in the read file data segment after successful reading; for any file data segment, the header information is completed when it is determined that the header information in the file data segment is incomplete according to the identification result of the header information; communication between the plurality of target processes is controlled, and a global header information set is constructed according to the communication result and the identification result of the header information of each target process; and grid data in the file data segment of each target process is analyzed according to the global header information set, to obtain grid data of each file data segment, and the grid data of each file data segment is integrated to obtain a target grid file.

[0011] It can be seen that, in the initial stage, the grid file to be read is divided into multiple file data segments according to the file size, and each target process is assigned a corresponding file data segment, so that each target process only needs to read the corresponding file data segment through one reading operation. Subsequently, each target process identifies the header information of the respective file data segment and completes the incomplete header information, and then constructs a global header information set in combination with the communication results between the target processes. Finally, according to the global header information set, the grid data in the file data segment of each target process is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain the target grid file. In this way, multiple target processes can be controlled to read their respective corresponding file data segments in parallel at one time, thereby significantly reducing the access frequency of the disk and the total data throughput pressure. Especially in large-scale grid file processing scenarios, this mechanism of avoiding repeated traversal not only greatly shortens the data loading time, but also effectively avoids the risk of file system throttling and network congestion caused by intensive operations. While ensuring data integrity and parsing accuracy, nearly linear parallel scalability is achieved, greatly improving the preprocessing efficiency of engineering applications such as aero-engine combustion chamber simulation in a supercomputing environment. BRIEF DESCRIPTION OF DRAWINGS

[0012] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which:

[0013] Figure 1 A flowchart of a parallel processing method of a grid file according to an embodiment of the present disclosure; Figure 2 A structure diagram of data segment reading in a grid file according to an embodiment of the present disclosure; Figure 3 A structure diagram of a Fluent grid file according to an embodiment of the present disclosure; Figure 4 A diagram illustrating the truncation of header information according to an embodiment of the present disclosure; Figure 5 A flowchart of another parallel processing method of a grid file according to an embodiment of the present disclosure; Figure 6 A structure diagram of a parallel processing device of a grid file according to an embodiment of the present disclosure; Figure 7 A structure diagram of an electronic device according to an embodiment of the present disclosure; Figure 8 A structural schematic diagram of a computer system provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 9 A structural schematic diagram of a computer program product provided by an embodiment of the present disclosure is shown in FIG. 2. DETAILED DESCRIPTION

[0014] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0015] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.

[0016] As used herein, the term "includes" and its variants are open-ended, meaning "includes but is not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined in the description that follows. It should be noted that the concepts mentioned in the present disclosure are merely used for distinguishing different apparatuses, modules or units and are not intended to limit the functions of the apparatuses, modules or units.

[0017] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0018] The names of the messages or information exchanged between the plurality of apparatuses in the embodiments of the present disclosure are merely used for illustrative purposes and are not intended to limit the scope of the messages or information.

[0019] With the in-depth application of Computational Fluid Dynamics (CFD) in the design and optimization of aero-engines, higher requirements are put forward for the size and complexity of the computational grid, which is particularly reflected in the representation of complex geometric configurations. For example, in Fluent, the node number and element number of the unstructured grid can reach hundreds of millions or even billions. The excessive grid size not only greatly increases the demand for computing resources, but also makes grid generation and preprocessing the key link that restricts the efficiency of the entire simulation process. Therefore, an efficient and reliable grid processing method has become one of the core challenges to improve the overall performance of CFD simulation. Here, CFD is a discipline that integrates computer science, numerical methods, and fluid mechanics theory to simulate the flow, heat transfer, and related physical phenomena of fluids through numerical calculation; Fluent is a globally leading commercial CFD simulation software.

[0020] Generally, when using Fluent to carry out simulation calculation, the pre-processing stage needs to first read the Fluent grid file and load various topological data such as node coordinates, element connection relationships, and boundary identifiers in the file into the memory of the current processing node to form structured data objects that can be called for calculation. Among them, the Fluent grid file can use a text format specific to Fluent, and different categories of data are identified by specific header information starting position, and open bracket "(" start and close bracket ")" end as boundary. However, generally, various topological data are randomly distributed in the Fluent grid file, lacking a fixed order, which significantly increases the complexity of reading and parsing.

[0021] To solve the above problems, the related art usually reads various topological data in the Fluent grid file in the following way: Method one, serial reading method, that is, by scanning the brackets to confirm the start and end of the header information and data segment, the data is read into the memory in bytes one by one, but this method cannot utilize parallel computing resources and has low reading efficiency.

[0022] Method two, multi-process parallel reading method, that is, first, the main process sequentially reads the entire Fluent grid file, identifies and records the information start position and data length of all header information; then each process calculates the file data segment and data segment range to be read according to the allocated header information metadata; finally, the corresponding data is read into the memory of each process through parallel file I / O. Although this method realizes multi-process cooperative reading and has certain parallel ability, it still has two obvious performance bottlenecks in super large-scale grid processing because it must rely on the main process for global scanning, and each process still needs to independently initiate multiple I / O requests, which leads to two obvious performance bottlenecks: one is that the global pre-reading operation introduces additional I / O overhead and synchronization delay, and the other is that a large number of scattered I / O requests can easily cause access conflicts and system throttling in the parallel file system, which seriously restricts the overall reading efficiency and scalability.

[0023] It can be seen that in the related art, there are several fundamental performance bottlenecks that seriously limit the efficiency of super large-scale Fluent grid processing. First, in order to identify and locate the header information of the data segment, at least two complete I / O operations must be performed on the Fluent grid file: the first is global scanning by the main process to extract the header information, and the second is actual data reading operation by each process according to the header information. This repeated traversal mechanism significantly reduces I / O efficiency. Secondly, since the Fluent grid file usually adopts a special text format with parentheses as boundaries, and the data segments are distributed in disorder, the traditional parallel method needs to parse byte by byte to identify the structure boundary, which will trigger a large number of fine-grained network and storage I / O requests in the supercomputing environment. This not only causes high system load, but also easily triggers the throttling mechanism of the parallel file system, and even causes the computing nodes to fail due to I / O timeout. In addition, although the global scanning in the related art avoids redundant reading, it increases the number of I / Os, while the segmented reading reduces the number of times but may introduce redundant data. This limitation makes it unable to adapt to the processing needs of hundreds of millions to billions of grids, seriously restricting the efficiency of the entire simulation process and the scalability of the system.

[0024] To solve the above problems, the embodiment of the present disclosure provides a parallel processing method of grid files. The method can distribute all target processes in a byte-level load balancing manner to read the reading range of the grid file to be read, and introduce a mechanism of buffering data segments when each target process reads to ensure the integrity of the header information. Meanwhile, the method initiates a bulk data reading operation by using MPI-IO, loads all file data segments into the memory of the corresponding target process at one time, and stores the identifiable data into the corresponding data structure by identifying the header information. Then, the method shares the global header information set through inter-process communication of the target processes, and completes the missing or truncated header information of each target process. Finally, the method writes the processed grid data into a regular format file through one-time bulk I / O, so that the efficient parallel analysis and distribution of the entire grid can be completed through only one file reading and minimum I / O calls.

[0025] Figure 1 A flowchart of a parallel processing method of grid files provided by an embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the method specifically includes the following steps. Figure 1 S101, obtaining the file size of the grid file to be read, and calculating the file data segment to be read by each target process through a uniform division logic. Here, the grid file to be read can be a Fluent grid file, and the number of target processes is usually dynamically allocated by the master process or the scheduling system in the program initialization stage.

[0026] In some embodiments, all target processes can be controlled to synchronously obtain the file size of the grid file to be read, i.e., the global byte size. Then, in combination with a byte-level complete uniform division algorithm, the byte range of the file data segment to be read by each target process in the grid file to be read is determined.

[0027] Specifically, the file size of the grid file to be read can be determined first, which can be denoted as n bytes. The number of target processes is denoted as m. At this time, the theoretical reading amount of the file data segment to be read by each target process is n / m bytes. Here, when n cannot be divided by m, the remainder r can be uniformly divided to each target process again to ensure load balancing. For example, when n=100 bytes and m=4, each target process reads 25 bytes, and no remainder needs to be allocated additionally. When n=102 bytes and m=4, each target process reads 25 bytes, but there are r=2 bytes left. At this time, the remaining 2 bytes can be uniformly divided to each target process, i.e., each target process is finally allocated 25.5 bytes.

[0028] ​In some embodiments, after determining the amount of data each target process needs to read, the byte range of the target data segment corresponding to each target process in the grid file to be read can also be accurately calculated according to the global number of each target process, which can be expressed as an interval (start, end), where the start position start and the end position end can be located in bytes.

[0029] For example, for a file with a size of 100 bytes and 4 target processes: target process 0 reads (0, 25), target process 1 reads (25, 50) byte range data, target process 2 reads (50, 75) byte range data, and target process 3 reads (75, 100) byte range data, so as to achieve continuous and non-overlapping division of the entire grid file to be read. Here, the global number of each target process refers to a unique and continuous integer identifier assigned by the runtime system to each target process participating in the calculation in a parallel computing environment, which is used to uniquely distinguish and locate different target processes in a global range, and is the basis for inter-process communication, data division and cooperative work.

[0030] In the parallel reading process based on byte division, the end position of the target data segment may be located in the middle of a row of data, which will cause the header information in the target data segment to be missing. Figure 2 A structural diagram of data segment reading in a grid file is provided for an embodiment of the present disclosure. As shown in Figure 2 The first target data segment 201 to be read by target process 1, the buffer data segment 202 of target process 1, and the second target data segment 203 to be read by target process 2 are included. The end position of the first target data segment 201 is located between "10" and "1" in the data "(12 (6f 1 10 1 3))", which causes the data to be split into two incomplete parts between "10" and "1", one part in the first target data segment 201 and the other part in the second target data segment 203. If such truncation is not handled, the header information or topology of target process 1 and target process 2 will not be recognized, which will cause subsequent data parsing errors.

[0031] In some embodiments, to eliminate the truncation problem of the header information in the file data segment, the embodiments of the present disclosure can add a buffer data segment in the file data segment required to be read by each target process, at this time, the file data segment can include the target data segment originally required to be read and the additionally added buffer data segment. Here, the buffer data segment can be a segment of data of a specific length read continuously after the target data segment originally required to be read by the target process, so that when the content of the header information of the target process is incomplete, the data missing from the header information can be directly obtained from the corresponding buffer data segment without frequent inter-target process communication to complete the content of the header information of the current target process. It can be seen that the design of the buffer data segment not only ensures the integrity of the header information in the to-be-read grid file, but also reduces the communication overhead between target processes for coordinating the integrity of the file data segment, and improves the overall reading efficiency of the to-be-read grid file. Here, the header information is a continuous byte sequence with a fixed or variable length in the to-be-read grid file.

[0032] Specifically, the embodiments of the present disclosure introduce a buffer data segment for each target process, that is, an additional buffer data segment with a length of a is read at the end of the target data segment of each target process, so that an overlapping data segment with a length of a is formed between adjacent target processes. Here, the data length a of the buffer data segment can be greater than the maximum byte number of a single row of data in the target file, and a can usually be set to be greater than or equal to 100, so that the reading requirements of most grid files can be met. It should be understood that for the target process with the last global number, since the file data segment to be read by the last target process has reached the end of the file, the buffer data segment can not be set.

[0033] Based on this, after determining the byte range required to be read by each target process, a buffer data segment with a length of a can be extended backward to construct a file data segment composed of a target data segment and a buffer data segment for each target process. For example, when a to-be-read grid file with a total size of n is read using 4 target processes, target process 0 reads (0, n / 4+a), target process 1 reads (n / 4, 2n / 4+a), target process 2 reads (2n / 4, 3n / 4+a), and target process 3 reads (3n / 4, n).

[0034] It can be seen that the embodiments of the present disclosure can ensure the integrity of the header information of the to-be-read grid file while significantly reducing the communication overhead between target processes by using the byte-level equal division and buffer data segment mechanism, and high-efficiency and scalable parallel reading is achieved.

[0035] S102, control a plurality of target processes to read corresponding file data segments, identify the header information in the plurality of file data segments, and store the identified header information in a preset data structure.

[0036] In some embodiments, the grid file to be read can be opened using MPIIO, and read through a large block continuous I / O read operation, so that each target process acquires and stores the file data segment it is responsible for processing, and after the target process reads the corresponding file data segment, the read successful file data segment can be stored in the staging area corresponding to the target process. Here, the staging area is not the final data structure used for calculation, and its content is the original byte sequence without parsing, which needs to be further parsed and stored in a data array with a clear logical structure. The staging area can include a local area and a buffer area, the local area is used to store the target data segment that the target process needs to process, and the size of the local area is the file size of the grid file to be read divided by the total number of processes; and the buffer area is used to store the buffer data segment, and the size of the buffer area is a. Based on this, the total amount of bytes actually read by each target process can be determined as n / m+a. Here, MPIIO refers to a set of programming interfaces and specifications that allow multiple target processes to read and write the same or different files in parallel and cooperatively in a parallel computing environment, such as a high-performance computing cluster or a supercomputer.

[0037] After the target process completes the data read-in, all subsequent data processing is performed in the memory space, i.e., each target process first parses the original bytes of the target data segment in the staging area in the corresponding memory space to identify and extract the header information in the target data segment. The header information can include the total number of points, the total number of faces, the total number of bodies, and the starting position information of the point data block, the starting position information of the face data block, and the starting position information of the body data block. These information together with the grid dimension information constitute the header information. Here, one target process corresponds to one independent and private memory space, and the memory space includes a staging area for temporarily storing data, and one target process can include at least one header information.

[0038] Figure 3 A structure diagram of a Fluent grid file is provided for an embodiment of the present disclosure. As shown in Figure 3 The basic structure information of the Fluent grid file is shown, and four key data types and their storage formats contained in the grid file are specifically presented, including string annotation 301, dimension 302, size information 303, basic information of point data 304, specific information of unit topology 305, specific information of face topology 306, and identification information of face topology 307. Among them, the size information 303 includes three header information starting with an open bracket and ending with a close bracket. The first digit in each header information is used to identify the data type corresponding to the header information, and the data type can include point coordinate marker 10, face topology marker 12, and body topology marker 13. Here, the core role of the header information is to define the data type and data structure of the subsequent data block.

[0039] In some embodiments, when each target process parses the target data segment in its staging area, it can identify zero or at least one header information, at which point, all header information owned by the target process can be collected and encapsulated into a data structure named Patterns, which contains the number of the target process to which the header information belongs, the data type of the header information, the byte range of the header information in the target data file, whether the header information is truncated due to partitioning, the specific location of the truncation of the header information, and the byte range read in the buffer to complete the header information. Here, one target process corresponds to one Patterns.

[0040] As shown in Figure 2 , the target process 1 first identifies multiple complete header information from the first target data segment 201 and records it in the Patterns, and when parsing to the end of the first target data segment 201, the target process 1 finds that one of the header information is truncated, resulting in incomplete header information in the target process 1, at which point, the corresponding data in the target process 1's buffer data segment 202 can be obtained to complete the incomplete header information. For example, continuing to read "1 3 ))" makes the header information in the first target data segment 201 complete. At this point, the Patterns will mark this header information as truncated and record the byte range of the complete reading in the buffer data segment to ensure that the subsequent data processing can correctly identify and splice the data segment.

[0041] In some embodiments, when the processed header information is the start position information of specific information, special processing is required. Here, the start position information of specific information is a special type of header information, which functions to declare the data type immediately following it, and the specific format can be (a (b c d...)). At this point, if the truncation occurs within the start position information of specific information, the data parsing process should not terminate immediately upon encountering a carriage return.

[0042] Figure 4 A schematic diagram of a truncated header information is provided for an embodiment of the present disclosure. As shown in Figure 4 , because different versions of grid files can insert carriage returns between parentheses, if only reading to the carriage return, the complete structure information can not be obtained. Therefore, the correct processing method can be to continue reading the buffer data segment in the buffer area until the start opening parenthesis of the next header information is captured, and mark all contents before the opening parenthesis as the complete part of the current header information, while recording its position in the buffer to ensure the completeness and accuracy of the header information structure.

[0043] When the target process successfully resolves and obtains the starting position information of at least one complete point, face or volume, the data type of the subsequent data segment can be determined according to the starting position information. For this purpose, each target process creates three associated containers of point information, face topology information and volume topology information respectively. Based on this, all identifiable data segments in the staging area can be parsed and stored in the corresponding container according to the determined header information type.

[0044] The reading of specific data can take the complete row in the grid file to be read as the minimum unit, wherein the point information corresponds to the two-dimensional or three-dimensional coordinates of a point; the face topology information contains all point indexes constituting the face and their adjacent left and right cell identifiers; the volume topology information needs to distinguish two cases: case one, if the grid type is explicitly specified in the header information, the topology data mapping can be directly generated according to the declaration; case two, if it is not specified, each row of data in the local area needs to be parsed in order, and if a row of data is truncated at the partition boundary, it needs to be read in the buffer until the end of line character is reached, so as to complete the data item and record the truncated position. It should be understood that the processing rules here are different from the header information completion strategy, and the reading of data row takes the end of line character as the explicit termination boundary.

[0045] S103, control multiple target processes to communicate with each other, and when no header information is identified in the file data segment of a certain target process, identify the file data segment of the target process according to the communication result.

[0046] In some embodiments, after all target processes complete the data collection of the corresponding Patterns, all target processes can be controlled to communicate with each other to construct a global header information set AllPatterns according to the Patterns of each target process, so that all target processes can synchronously obtain all header information. Since the number of header information contained in each target process is different, the communication can be divided into two steps: first, exchange the length of the Patterns data structure of each target process, and then synchronously transmit the actual data body.

[0047] After the communication is completed, each target process can parse the residual data in the staging area based on AllPatterns, which can be specifically performed by traversing AllPatterns to find the place where the data type changes between the two header information before and after it, and determining the type of the data located in the interval. Here, if the current file data segment is truncated by the previous target process at the partition, it must be read from the truncated point recorded in the Patterns of the previous target process to ensure the continuity of the data; if the data segment is complete, it is directly parsed from the starting position of the file data segment of the target process. In the parsing process, the program automatically skips spaces, carriage returns and other format characters, and directly extracts valid numerical information.

[0048] For example, assume that there is a piece of residual data in the staging area of the target process 2, which cannot be determined as a type, and the byte range of the residual data corresponds to the 150-170 bytes of the global file. After obtaining the AllPatterns, the target process 2 traverses the AllPatterns and finds that the last header information of the target process 1 (ending at 149 bytes) is marked as the face topology data, and the first header information of the target process 3 (starting at 171 bytes) is marked as the volume topology data. Thus, it is determined that the residual data between 150-170 bytes should be the information for the transition from the face topology to the volume topology. Since the end of the data of the target process 1 is recorded in the Patterns as being truncated at 149 bytes, the target process 2 starts to parse from 150 bytes (after the truncation point), automatically skips the line feed, continuously reads the numerical sequence, and correctly parses the numerical sequence as the topology connection information of a certain face unit according to the context information, and finally fills the topology connection information into the face topology data structure.

[0049] Finally, each target process fills the mesh data corresponding to each header information into the pre-set point topology structure container, face topology structure container, or volume topology structure container according to the type. After the data filling is completed, each target process has completed the entire analysis work in the local staging area, and the global mesh topology is formed by the aggregation of the local data of all processes. Subsequently, each target process writes the complete file data segment in the memory into the standardized target file through a batch aggregation I / O operation. This process ensures that the output data has good structural properties and reading and writing efficiency, and facilitates the subsequent calculation task to be directly and efficiently called. Here, the header information marks the starting position and data type of each data segment in the mesh file.

[0050] S104, integrating the mesh data of each target process according to the communication result and the global header information set to obtain a target mesh file.

[0051] In some embodiments, before the mesh data of each target process is integrated, the number of point data, the number of face data, and the number of volume data held by each target process can be summarized from the AllPatterns. Since the types and quantities of data processed by each target process are unevenly distributed, the above information can be synchronized through inter-target process communication to ensure that all target processes know the global data distribution state, which is the basis for subsequent parallel writing operations.

[0052] Before parallel writing, the writing offset of each target process in different data files can be calculated, which depends on the sum of the amount of data to be written by all processes before the target process, and the calculation is based on the inherent format of each data type, for example, point data contains 2 or 3 coordinate values per record; body topology contains only 1 type identifier per record; surface topology contains 4 values (2 point indexes and left and right cell numbers) in two dimensions, and 6 values (4 point indexes and left and right cell numbers) in three dimensions, and the insufficient is filled with -1. In this way, by unifying the length of each record, the writing position and total amount of data of each target process can be accurately calculated, laying a foundation for subsequent parallel IO. Based on this, the file data segments of multiple target processes can be simultaneously and parallelly written into the target file according to the data type through MPIIO, so as to obtain a target grid file containing point grid files, surface grid files and body grid files. Thus, when reading the grid file next time, the target grid file can be directly read to generate point grid files, surface grid files and body grid files.

[0053] It can be seen that, in the initial stage, the grid file to be read is divided into multiple file data segments according to file size and distributed to each target process. Each target process only needs to read the corresponding file data segment through one reading operation, then each target process identifies the header information of the file data segment and completes the incomplete header information, and then constructs a global header information set in combination with the communication results between target processes. Finally, according to the global header information set, the grid data in the file data segment of each target process is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain a target grid file. In this way, multiple target processes can be controlled to read their corresponding file data segments in parallel at one time, thereby significantly reducing the access frequency of the disk and the total data throughput pressure. Especially in large-scale grid file processing scenarios, this mechanism avoids repeated traversal, not only greatly shortens the data loading time, but also effectively avoids the risk of file system throttling and network congestion caused by intensive operations. While ensuring data integrity and parsing accuracy, nearly linear parallel scalability is achieved, greatly improving the preprocessing efficiency of engineering applications such as aero-engine combustion chamber simulation in a supercomputing environment.

[0054] The parallel processing method of the grid file provided by the embodiments of the present disclosure can be executed by a terminal or a chip applied to the terminal.

[0055] Exemplarily, the terminal can include one or more of a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a Personal Digital Assistant (PDA), and a wearable device based on augmented reality (AR) and / or virtual reality (VR) technology, and the like, and can further include, but is not limited to, a remote control device, a wearable device, a street lamp, a smart terminal of a household appliance, and the like, and the embodiments of the present disclosure do not make specific limitations thereto.

[0056] Figure 5 A flowchart of another parallel processing method of a grid file provided by an embodiment of the present disclosure is shown in FIG. 5. As shown in FIG. 5, the method specifically includes the following steps. Figure 5 S501, obtaining a file size of a grid file to be read, and determining file data segments to be read by a plurality of target processes according to the file size.

[0057] In some embodiments, first, the main process obtains the total size of the grid file to be read through a file system call. For example, for a grid file (Fluent grid file in the above) to be read with a size of 1.2 GB, according to a preset number of 4 target processes, the file size is divided by the number of processes to obtain a theoretical load of 300 MB for each target process. In order to process the remainder (200 MB) that cannot be divided, a dynamic allocation strategy is adopted: the first two target processes (target process 0 and target process 1) each allocate an additional 50 MB, i.e., 350 MB is read respectively, and the starting offset is 0 and 350 MB respectively; and the last two target processes (target process 2 and target process 3) read 250 MB respectively, and the starting offset is 700 MB and 950 MB respectively, to ensure that all data is completely covered and no overlap occurs.

[0058] In this way, by dispersing the remainder to the previous target processes, it is avoided that a single target process handles too much remaining data and becomes a performance bottleneck, so that the I / O operation time of each target process tends to be consistent. At the same time, the explicit division of the continuous file blocks eliminates the reading conflict, significantly improves the parallel reading efficiency, and especially for large grid files, can greatly shorten the overall data loading time, and lays a high-efficiency data foundation for subsequent parallel computation.

[0059] S502, respectively controlling a plurality of target processes to read corresponding file data segments, and after reading is successful, controlling each target process to identify the header information contained in the read file data segment.

[0060] ​In some embodiments, after the master process broadcasts the respective file data segment division information (byte range of the file data segment) to the four target processes, the master process controls the target processes to read the specified file data segment in parallel. For example, after target process 0 reads 0-350MB of data, it immediately identifies the fixed-format header information (such as grid dimension, variable quantity, and other metadata) at the beginning of the file data segment in the corresponding memory; the same operation is performed synchronously for other target processes to ensure that each target process can correctly understand the structure of the data segment it is responsible for.

[0061] In this way, each target process can independently verify the validity of the data without waiting for the entire grid file to be read, greatly reducing the I / O blocking time. At the same time, local processing of the header information avoids frequent query dependence on the master process, improving processing efficiency and laying the foundation for subsequent independent processing of the main grid data by each target process based on the metadata, forming an efficient staged parallel pipeline.

[0062] S503, for any file data segment, when it is determined according to the identification result of the header information that the header information in the file data segment is incomplete, the header information is completed.

[0063] In some embodiments, suppose that target process 2 discovers that the header information is truncated (e.g., only contains the grid dimension but lacks the variable definition) when reading the data segment starting at 700MB, the target process 2 will immediately send a request to the adjacent target process 1 (responsible for the previous data segment) to obtain a complete copy of the header information contained in the tail of the file data segment of target process 1. Subsequently, target process 2 will splice the received supplementary information with its incomplete header information to reconstruct the complete file header structure, thereby ensuring correct parsing of the subsequent grid data.

[0064] In this way, information completion is achieved through lightweight communication between adjacent target processes, avoiding the bottleneck formed by concentrating all header information processing in the master process, and maintaining the overall efficiency of parallel reading. This design enhances the robustness of parallel file reading, ensuring that each target process can obtain the metadata required for complete parsing even if the data segment division is not ideal, improving the fault tolerance and adaptability of the system.

[0065] S504, control the communication between the plurality of target processes, and construct a global header information set according to the communication result and the identification result of the header information of each target process.

[0066] In some embodiments, under the coordination of the master process, each target process can send the respective identified header information to all other target processes through a communication operation. For example, target process 0 reports that its file data segment contains X-axis 1-100 grid points, target process 1 reports 101-200 points, and the master process collects this information to generate a global header information set describing the entire grid domain (X-axis 1-400 points), and broadcasts it to all target processes, ensuring that each target process obtains a unified global data view.

[0067] In this way, each target process can not only process local file data segments, but also clearly understand the overall grid structure, laying the foundation for computing operations that require global information. At the same time, the communication operation avoids the serial bottleneck of the master process, achieving efficient synchronization of information. Although a small amount of communication overhead is introduced, the correctness and collaboration ability of the entire parallel program are significantly improved, which is a key step in building efficient large-scale parallel applications.

[0068] S505, according to the global header information set, the grid data in each target process file data segment is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain a target grid file.

[0069] In some embodiments, in aircraft flow simulation, the header information defines a complete aerodynamic grid (including wing surface grid faces, far-field boundary grid faces, and space flow field volume grids). At this time, target process 0 parses its corresponding file data segment to obtain the surface grid points (points) in the wing leading edge region and the triangular grid faces formed thereby; target process 1 parses the wing trailing edge and part of the near-field space, including surface grid and space volume grid elements; target process 2 parses the volume grid in the far-field region. After each target process completes the parsing, the master process splices the wing leading edge surface grid of target process 0 and the trailing edge surface grid of target process 1 according to the global coordinates to restore the complete wing surface grid, and then assembles the space volume grids of target process 1 and target process 2 according to the topological relationship, and finally integrates them into a target grid file containing complete surface and space volume grids for CFD calculation.

[0070] In this way, through systematic parsing and integration operations, the grid data scattered in the file data segments stored in different target processes can be efficiently and completely reconstructed into a target grid file with unified structure and standardized format, which not only significantly improves the utilization and sharing of data resources, but also greatly facilitates subsequent reading and application; Since the data has been integrated into a single file and follows a standard structure, subsequent software or analysis programs can directly and quickly read and parse, avoiding the complexity of handling multiple heterogeneous data sources, thereby improving the efficiency and reliability of the overall data processing process.

[0071] It can be seen that, in the initial stage, the grid file to be read is divided into multiple file data segments according to the file size, and each target process is assigned a corresponding file data segment, so that each target process only needs to read the corresponding file data segment through one reading operation. Subsequently, each target process identifies the header information of the respective file data segment and completes the incomplete header information, and then constructs a global header information set in combination with the communication results between the target processes. Finally, according to the global header information set, the grid data in the file data segment of each target process is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain the target grid file. In this way, the multiple target processes can be controlled to read their respective corresponding file data segments in parallel at one time, thereby significantly reducing the access frequency of the disk and the total data throughput pressure. Especially in a large-scale grid file processing scenario, this mechanism of avoiding repeated traversal not only greatly shortens the data loading time, but also effectively avoids the risk of file system throttling and network congestion caused by intensive operations. While ensuring data integrity and parsing accuracy, nearly linear parallel scalability is achieved, greatly improving the preprocessing efficiency of engineering applications such as aero-engine combustion chamber simulation in a supercomputing environment.

[0072] In some embodiments, after each target process identifies the header information contained in the read file data segment, the method further comprises: according to the header information identification result, classifying the multiple file data segments of the multiple target processes into identifiable file data segments and unidentified data segments; for the identifiable file data segments, when it is determined according to the header information identification result that the header information in the identifiable file data segment is incomplete, completing the header information; and for the unidentified data segments, determining the header information of the unidentified data segment according to the header information identification result of the file data segment adjacent to the unidentified data segment.

[0073] Specifically, after completing the header information identification, each target process's file data segment can be classified according to the identification result. For example, for target process 0, target process 1, target process 2 and target process 3, target process 0, target process 1 and target process 3 are marked as identifiable file data segments because they successfully identify key metadata such as grid dimensions; while the data segment of target process 2 fails to parse valid header information because its starting position falls exactly at the gap between two grid blocks, and is marked as an unidentified data segment. For the incomplete variable type information found by target process 1, the missing descriptor can be requested and completed from target process 0. For target process 2, the header information of its adjacent target process 1 and target process 3 data segments can be used to infer that this segment should be pure grid node coordinate data, and accordingly the correct header information is assigned to it.

[0074] In some embodiments, there are multiple cases for the target process to identify the header information of its file data segment: if a target process fails to identify any header information, it can directly infer and determine the header information of the file data segment of the target process according to the header information identified by the previous target process; if a target process successfully identifies the header information in the middle of the file data segment, the data between the header information and the beginning of the file data segment is likely to be incomplete due to the partition boundary, and the corresponding header information cannot be determined independently. At this time, the tail header information record of the previous target process still needs to be relied on to determine the actual type of the incomplete data, so as to ensure that the residual data before the header information can also be correctly parsed when the header information appears in the middle of the data segment. Here, the specific implementation can also refer to the related content of S102 and S103, which will not be described in detail here.

[0075] In this way, by distinguishing between "identifiable" and "unidentified" file data segments, and respectively adopting the "adjacent completion" and "adjacent inference" strategies, various complex file cutting situations can be intelligently handled, especially when the division boundary falls on the metadata area or the junction of different data types. This not only avoids the interruption of the entire process due to the failure of individual file data segment identification, but also reduces the dependence on perfect data segment division, ensuring that the data can be correctly reconstructed even under non-ideal file division, and improving the fault tolerance and adaptability of the entire method.

[0076] In some embodiments, according to the communication results, the header information in the identifiable file data segment after the completion of the header information and the header information in the unidentified data segment after the identification of the header information can be integrated to construct a global header information set.

[0077] Specifically, assuming that the target process 1 completes the header information to the target process 0, and the target process 2 identifies the header information of its data segment from the target process 1 and the target process 3, the main process initiates global communication to collect the updated header information of all target processes, including the complete grid parameters of the target process 0 and the target process 1, the newly inferred node coordinate data description of the target process 2, and the boundary condition identification of the target process 3. Based on this, the main process can integrate the completed information from the identifiable file data segment and the newly identified information from the unidentified file data segment, and finally construct a complete and consistent global header information set that accurately describes the organization structure of the entire grid file, all variable types, and the physical meaning of each data segment. Here, the global header information set is AllPatterns in the above, and the construction process of AllPatterns can refer to the related description in the above, which will not be described in detail here.

[0078] In this way, the reliable information after completion and the reasonable information after inference are fused based on the result of the previous classification processing, and the accuracy and integrity of the global header information set are ensured. This design makes the construction process of the global set have strong fault tolerance. Even if the initial division leads to the fact that part of the data segment cannot be independently identified, a unified global cognition can be finally formed, which lays a solid foundation for the correct parsing and integration of subsequent data and ensures the final quality of the parallel processing process.

[0079] In some embodiments, the file data segment includes a target data segment and a buffer data segment, wherein, in the grid file to be read, the buffer data segment is located at the end of the target data segment and overlaps with part of the data of the next adjacent file data segment; for any file data segment, when it is determined according to the header information recognition result that the header information in the file data segment is incomplete, the header information is completed, including: when it is determined according to the header information recognition result that the header information in the file data segment is incomplete, the header information is completed according to the buffer data segment of the file data segment.

[0080] Specifically, the file data segment read by each target process is designed to include two parts: a core target data segment and a buffer data segment located at the end and overlapping with the starting part of the next target process. For example, target process 1 is responsible for reading data from 350MB to 700MB, wherein 350-699MB is the target data segment of target process 1, and 699-701MB is the buffer data segment, which is 2MB in length and overlaps completely with the 0-2MB at the beginning of the target data segment of target process 2. At this time, when target process 1 finds that the header information at the end of its own target data segment is truncated, target process 1 does not need to request target process 2 and can directly extract the header of the next data block from the local buffer data segment (i.e., the 699-701MB region), thereby completing the missing header information. Here, the relevant content can be referred to the related description of S102 in the foregoing, and will not be described in detail here.

[0081] In some embodiments, the buffer data segment can be used not only to complete the truncated header information but also to complete the truncated data body. Specifically, when a target process parses the data body after a complete header information (such as (66 2)), if the data body is truncated due to being located at the end of the file data segment (for example, the header information indicates that there should be 66 face unit data, but the actual file data segment contains only 64), the target process will continue to read data from the buffer data segment until the complete 66 data units are obtained, thereby ensuring the integrity of a single data body (such as a complete face topology sequence).

[0082] In this way, by sacrificing a small amount of redundant disk reading (each target process reads a small amount of overlapping data), the synchronization communication waiting and delay caused by incomplete header information in the parsing stage are completely avoided. This space-time trade-off strategy is particularly suitable for high-performance computing environments, significantly reducing the coupling and coordination costs between target processes, enabling each target process to independently and efficiently complete header information parsing and completion, greatly improving the efficiency and smoothness of the overall parallel reading.

[0083] In some embodiments, the header information includes an information start identifier and an information end identifier, and for any file data segment, when it is determined that the header information in the file data segment is incomplete according to the identification result of the header information, the header information is completed, and the method comprises: when only the information start identifier or the end identifier of the header information exists in the file data segment, it is determined that the header information in the file data segment is incomplete.

[0084] Specifically, assuming that the header information of the grid file starts with a specific character (START) and ends with (END). Target process 1 finds (START) in the middle of its file data segment, but does not find the corresponding (END) in the data after (START), at which point it can be determined that the header information in the file data segment is incomplete. For incomplete header information, information can be completed from the cache data segment corresponding to target process 1. For details, please refer to the related content in the above Figure 4 .

[0085] In this way, the specific situation of the truncated header information can be accurately identified, providing accurate basis for subsequent targeted completion strategy, avoiding misjudgment based on complex content analysis, and enhancing the robustness and automation level of the entire identification and completion process.

[0086] In some embodiments, according to the global header information set, the grid data in the file data segment of each target process is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain the target grid file, comprising: according to the global header information set, determining the data type and data range of the grid data included in each file data segment; for each file data segment, according to the data type and data range of the grid data included in the file data segment, the grid data in the file data segment is parsed into structured data, wherein the structured data includes at least one of point coordinate data, face topology data and volume topology data; the structured data of the plurality of file data segments is integrated to obtain the target grid file.

[0087] Specifically, in the simulation of an aero-engine, the global header information set defines the composition of the entire model: the surface mesh of the fan blade (face topology data), the volume mesh of the compressor passage (volume topology data), and the point coordinate data of the combustion chamber wall. At this time, it can be assumed that target process 0 parses the face mesh of the 1st-3rd fan blade according to the data range of the allocated file data segment; target process 1 parses the volume mesh unit of the 4th-6th compressor according to the data range of the allocated file data segment; and target process 2 parses the point cloud data of the combustion chamber region according to the data range of the allocated file data segment. At this time, after each target process completes the local structured data conversion, the host process integrates the blade face mesh, the compressor volume mesh, and the combustion chamber point cloud data according to the engine aerodynamic path, and finally generates a complete engine flow passage target mesh file that can be used for CFD calculation.

[0088] In this way, by performing parallel parsing of heterogeneous mesh components (points, faces, and volumes) according to global information, the loading efficiency of multi-component and multi-precision hybrid meshes is greatly improved. Based on the integration of data semantics and spatial range, the geometric continuity and topological consistency of the mesh data of each component at the interface are ensured, providing an accurate mesh model for high-fidelity engine flow field simulation, and significantly shortening the data preparation time for large-scale parallel computing.

[0089] In some embodiments, the file size of the mesh file to be read is obtained, and the file data segments to be read by the plurality of target processes are determined according to the file size, including: dividing the file size by the number of target processes to obtain the theoretical reading byte number of each target process; and determining the file data segment to be read by each target process in the mesh file to be read according to the preset order of each target process and the theoretical reading byte number.

[0090] In some embodiments, the size of the aero-engine whole mesh file to be read is 8GB, and the number of target processes is 4. The system first divides the file size to obtain the theoretical reading byte number of each target process, which is 2GB. Then, the file data segments are allocated strictly according to the preset order of the target process ID (0, 1, 2, 3): target process 0 reads the data segment from offset 0 to 2GB, target process 1 reads the data segment from 2GB to 4GB, target process 2 reads the data segment from 4GB to 6GB, and target process 3 reads the data segment from 6GB to 8GB, ensuring that the entire file is continuously and non-overlappingly covered.

[0091] In this way, the most basic load balancing is achieved, ensuring that the I / O data volume of each target process is basically consistent, and avoiding idle waiting of the target process caused by uneven data distribution. Moreover, this method has simple logic and extremely small overhead, and is suitable for files with uniformly distributed mesh data.

[0092] Table 1. Time consumed in processing grid files with different core numbers

[0093] As shown in Table 1, the parallel processing method of the grid file provided by the embodiments of the present disclosure has a significant advantage compared with the method of the related art. For example, for a grid unit of 15.8 million, as the core number increases from 8 to 48, the processing time of the method provided by the embodiments of the present disclosure can be reduced from 11.71 seconds to 2.509 seconds, showing good parallel scalability, while the processing time of the method of the related art always maintains more than 21 seconds, which cannot effectively utilize multi-core resources. More importantly, when processing a super large grid of 840 million, the method of the related art cannot complete reading in a supercomputing environment due to I / O and memory problems, while the method provided by the embodiments of the present disclosure successfully processes only using 1000 cores in 50.13 seconds, which shows excellent large-scale processing capability and reliability.

[0094] As shown in Table 1, the method provided by the embodiments of the present disclosure shows excellent capability when processing a grid of 15.8 million units, and the time consumed decreases nearly linearly with the increase of the core number, and when processing a super large grid of 840 million, the method shows excellent capability. It can be seen that the method provided by the embodiments of the present disclosure effectively solves the problems of not being able to read and performance bottleneck caused by I / O and memory problems in the related art.

[0095] It can be seen that in the initial stage, the grid file to be read is divided into multiple file data segments according to the file size, and each target process is assigned to a corresponding file data segment. Each target process only needs to read the corresponding file data segment through one reading operation, and then each target process identifies the header information of the corresponding file data segment and completes the incomplete header information, and then combines the communication results between the target processes to construct a global header information set. Finally, according to the global header information set, the grid data in the file data segment of each target process is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain the target grid file. In this way, multiple target processes can read their corresponding file data segments in parallel at one time, thereby significantly reducing the access frequency of the disk and the total data throughput pressure. Especially in the large-scale grid file processing scenario, this mechanism avoids repeated traversal, not only greatly shortens the data loading time, but also effectively avoids the risk of file system throttling and network congestion caused by intensive operations. While ensuring data integrity and parsing accuracy, it realizes nearly linear parallel scalability, and greatly improves the preprocessing efficiency of engineering applications such as aero-engine combustion chamber simulation in a supercomputing environment.

[0096] The above describes the scheme provided by the embodiments of the present disclosure from the perspective of the server. It can be understood that the server comprises hardware structures and / or software modules corresponding to each function in order to implement the above functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.

[0097] The embodiments of the present disclosure can divide the functional units of the server according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated in one management module. The above integrated module can be implemented in the form of hardware or software functional module. It should be noted that the division of the module in the embodiments of the present disclosure is illustrative, and is only a logical functional division. Actual implementation can have another division manner.

[0098] In the case of dividing each functional module according to each function, the exemplary embodiments of the present disclosure provide a grid file parallel processing device, which can be a server or a chip applied to a server. Figure 6 A structural schematic diagram of a grid file parallel processing device provided by an embodiment of the present disclosure is shown in FIG. 6. As shown in FIG. 6, the grid file parallel processing device 600 comprises: Figure 6 An acquisition module 601 is configured to acquire a file size of a grid file to be read, and determine file data segments to be read by a plurality of target processes according to the file size, respectively; A control module 602 is configured to control the plurality of target processes to read corresponding file data segments, respectively, and control each target process to identify header information contained in the read file data segment after successful reading; A completion module 603 is configured to, for any file data segment, complete the header information when it is determined that the header information in the file data segment is incomplete according to the identification result of the header information; A construction module 604 is configured to control communication between the plurality of target processes, and construct a global header information set according to the communication result and the identification result of the header information of each target process; ​The parsing module 605 is configured to parse the grid data in the file data segment of each target process according to the global header information set to obtain the grid data of each file data segment, and integrate the grid data of each file data segment to obtain a target grid file.

[0099] In an optional manner, after the control of each target process to identify the header information contained in the read file data segment, the method further comprises: According to the header information identification result, the file data segments of the plurality of target processes are classified into identifiable file data segments and unidentified data segments; for the identifiable file data segments, when it is determined according to the header information identification result that the header information in the identifiable file data segment is incomplete, the header information is completed; for the unidentified data segments, the header information of the unidentified data segments is determined according to the header information identification result of the file data segment adjacent to the unidentified data segment.

[0100] In an optional manner, according to the communication result, the header information in the identifiable file data segment after the header information is completed and the unidentified data segment after the header information is identified is integrated to construct a global header information set.

[0101] In an optional manner, the file data segment includes a target data segment and a buffer data segment, wherein, in the grid file to be read, the buffer data segment is located at the end of the target data segment and overlaps with part of the data of the next adjacent file data segment; the completion of the header information for any file data segment when it is determined according to the header information identification result that the header information in the file data segment is incomplete comprises: according to the header information identification result, when it is determined that the header information in the file data segment is incomplete, the header information is completed according to the buffer data segment of the file data segment.

[0102] In an optional manner, the header information includes information start identifier and information end identifier, and the completion of the header information for any file data segment when it is determined according to the header information identification result that the header information in the file data segment is incomplete comprises: when only the information start identifier or the end identifier of the header information exists in the file data segment, it is determined that the header information in the file data segment is incomplete.

[0103] In an optional manner, the parsing, according to the global header information set, of mesh data in the file data segment of each target process to obtain mesh data of each file data segment, and the integration of the mesh data of each file data segment to obtain a target mesh file, comprises: determining, according to the global header information set, a data type and a data range of mesh data included in each file data segment; for each file data segment, parsing the mesh data in the file data segment into structured data according to the data type and the data range of the mesh data included in the file data segment, wherein the structured data comprises at least one of point coordinate data, face topology data and volume topology data; and integrating the structured data of a plurality of file data segments to obtain the target mesh file.

[0104] In an optional manner, the obtaining of a file size of the mesh file to be read, and the determining of file data segments to be read by a plurality of target processes according to the file size, comprises: dividing the file size by the number of the plurality of target processes to obtain a theoretical reading byte number of each target process; and determining, according to a preset order of each target process and the theoretical reading byte number, the file data segments to be read by each target process in the mesh file to be read.

[0105] The embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; a memory for storing at least one processor-executable instruction; wherein the at least one processor is configured to execute the instructions to implement the steps of the above-mentioned method disclosed by the embodiments of the present disclosure.

[0106] Figure 7 The electronic device provided by an embodiment of the present disclosure is shown in a structural schematic diagram. As shown in the figure, the electronic device 700 comprises at least one processor 701 and a memory 702 coupled to the processor 701, and the processor 701 can execute the corresponding steps in the above-mentioned method disclosed by the embodiments of the present disclosure. Figure 7

[0107] ​The processor 701 can also be referred to as a central processing unit (CPU), which can be an integrated circuit chip that has the processing capability of signals. Each step in the method disclosed in the embodiments of the present disclosure can be completed by the integrated logic circuit of hardware or the instructions in the form of software in the processor 701. The processor 701 can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as hardware code processing for execution, or executed by a combination of hardware and software modules in the code processing. The software module can be located in the memory 702, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, and other mature storage media in the art. The processor 701 reads the information in the memory 702, and completes the steps of the above method in combination with the hardware thereof.

[0108] In addition, various operations / processes according to the present disclosure, when implemented by software and / or firmware, can be downloaded from a storage medium or a network to a computer system with a dedicated hardware structure, for example, Figure 8 The computer system 800 shown is installed with programs constituting the software, and when various programs are installed, the computer system can perform various functions, including functions such as those described above. Figure 8 The structural schematic diagram of the computer system provided for an embodiment of the present disclosure.

[0109] The computer system 800 is intended to represent various forms of digital electronic computer devices, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown in the figures, their connections and relationships, and their functions, are merely examples, and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0110] As Figure 8As shown, the computer system 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded into a random access memory (RAM) 803 from a storage unit 808. Various programs and data required for the operation of the computer system 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0111] A plurality of components in the computer system 800 are connected to the I / O interface 805, including an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information to the computer system 800, and can receive inputted digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 807 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 809 allows the computer system 800 to exchange information / data with other devices through a network, such as the Internet, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0112] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above. For example, in some embodiments, the above-described methods disclosed by embodiments of the present disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the above-described methods disclosed by embodiments of the present disclosure by any other appropriate means, such as by means of firmware.

[0113] The embodiment of the present disclosure further provides a computer readable storage medium, wherein when instructions in the computer readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the above method disclosed by the embodiment of the present disclosure.

[0114] The computer readable storage medium in the embodiment of the present disclosure can be a tangible medium, which can contain or store programs for use by or in connection with an instruction execution system, apparatus or device. The above computer readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specifically, the above computer readable storage medium can include one or more wire-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0115] The above computer readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0116] Figure 9 A schematic diagram of a computer program product provided by an embodiment of the present disclosure is shown. As shown in the figure, the computer program product 900 includes a computer program 901, wherein the computer program 901 is executed by a processor to implement the above method disclosed by the embodiment of the present disclosure. Figure 9

[0117] In the embodiments of the present disclosure, the computer program code for performing the operations of the present disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on a user computer, partially on a user computer, as an independent software package, partially on a user computer and partially on a remote computer, or entirely on a remote computer or server. In the case involving a remote computer, the remote computer can be connected to the user computer through any kind of network (including local area network (LAN) or wide area network (WAN)), or can be connected to an external computer.

[0118] ​The flow and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0119] The modules, components or units described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the name of the module, component or unit does not constitute a limitation on the module, component or unit itself.

[0120] The functions described above in the specification of the present disclosure can be performed by one or more hardware logic components. For example, and without limitation, examples of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0121] The above description is merely some embodiments of the present disclosure and a description of principles of technology used. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0122] Although some specific embodiments of the present disclosure have been described in detail by way of examples, it should be understood that the above examples are merely for illustration and not intended to limit the scope of the present disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A method for parallel processing of grid files, characterized in that, include: Obtain the file size of the grid file to be read, and determine the file data segments to be read by multiple target processes based on the file size; Multiple target processes are controlled to read the corresponding file data segments, and after successful reading, each target process is controlled to identify the header information contained in the read file data segment; For any of the file data segments, if it is determined from the header information identification result that the header information in the file data segment is incomplete, the header information shall be completed. Control communication between multiple target processes, and construct a global header information set based on the communication results and the header information identification results of each target process; Based on the global header information set, the grid data in the file data segment of each target process is parsed to obtain the grid data of each file data segment, and the grid data of each file data segment is integrated to obtain the target grid file.

2. The method according to claim 1, characterized in that, After controlling each target process to identify the header information contained in the read file data segment, the method further includes: Based on the header information identification results, the multiple file data segments of the multiple target processes are classified into identifiable file data segments and unidentified data segments; For the identifiable file data segment, when it is determined from the header information identification result that the header information in the identifiable file data segment is incomplete, the header information is completed. For the unidentified data segment, the header information of the unidentified data segment is determined based on the header information identification result of the file data segment adjacent to the unidentified data segment.

3. The method according to claim 2, characterized in that, The method further includes: Based on the communication results, the header information in the identifiable file data segment after the header information is completed is integrated with the header information in the unidentified data segment where the header information has been identified, in order to construct a global header information set.

4. The method according to claim 1, characterized in that, The file data segment includes a target data segment and a cached data segment. In the grid file to be read, the cached data segment is located at the end of the target data segment and partially overlaps with the next adjacent file data segment. When it is determined, based on the header information identification result, that the header information in any of the file data segments is incomplete, the header information is supplemented, including: When it is determined that the header information in the file data segment is incomplete based on the header information recognition result, the header information is completed based on the buffer data segment of the file data segment.

5. The method according to claim 1, characterized in that, The header information includes a start identifier and an end identifier. For any file data segment, when it is determined based on the header information identification result that the header information in the file data segment is incomplete, the header information is completed. The method includes: If only the header information start identifier or end identifier exists in the file data segment, the header information in the file data segment is determined to be incomplete.

6. The method according to claim 1, characterized in that, The step of parsing the grid data in the file data segment of each target process according to the global header information set to obtain the grid data of each file data segment, and integrating the grid data of each file data segment to obtain the target grid file, includes: Based on the global header information set, determine the data type and data range of the grid data included in each file data segment; For each file data segment, the grid data in the file data segment is parsed into structured data according to the data type and data range of the grid data included in the file data segment. The structured data includes at least one of point coordinate data, surface topology data, and volume topology data. The structured data of multiple file data segments are integrated to obtain the target grid file.

7. The method according to claim 1, characterized in that, The process of obtaining the file size of the grid file to be read and determining the file data segments to be read by multiple target processes based on the file size includes: The file size is divided equally according to the number of multiple target processes to obtain the theoretical number of bytes that each target process can read; Based on the preset order of each target process and the theoretical number of bytes to be read, the file data segment to be read by each target process in the grid file to be read is determined.

8. A parallel processing apparatus for grid files, characterized in that, include: The acquisition module is used to acquire the file size of the grid file to be read, and to determine the file data segments to be read by multiple target processes based on the file size; The control module is used to control multiple target processes to read the corresponding file data segments, and after successful reading, control each target process to identify the header information contained in the read file data segment; The completion module is used to complete the header information of any file data segment when it is determined from the header information recognition result that the header information in the file data segment is incomplete. A construction module is used to control communication between multiple target processes and to construct a global header information set based on the communication results and the header information identification results of each target process. The parsing module is used to parse the grid data in the file data segment of each target process according to the global header information set, so as to obtain the grid data of each file data segment, and to integrate the grid data of each file data segment to obtain the target grid file.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Grid parallel reading method and device, terminal equipment and readable storage medium

    CN116820772A

  • Method and device for identifying boundary type of structured grid

    CN118568798A

  • Information extraction method and device, electronic equipment, storage medium and program product

    CN118886411A

  • Parallel reading method based on GNNS hybrid grid

    CN119203641A

  • Set partitioning for encoding file system allocation metadata

    US20090300084A1