Parallel processing system, parallel processing method, and parallel processing device
The parallel processing system improves efficiency by executing processes in parallel and selectively excluding data that meets a condition, reducing overall process count and resource usage.
Patent Information
- Application Number
- PCT/JP2023/046786
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-07-03
AI Technical Summary
Existing systems fail to efficiently execute multiple processes in parallel, leading to suboptimal execution efficiency and resource utilization.
A parallel processing system that acquires multiple data sets and executes a predetermined process in parallel for each data set while determining if a condition is met, selectively excluding data that satisfies the condition to reduce unnecessary processing.
This approach reduces the total number of processes required, enhancing execution efficiency and resource utilization by focusing on data that meets the condition.
Smart Images

Figure JP2023046786_03072025_PF_FP_ABST
Abstract
Description
Parallel processing system, parallel processing method, and parallel processing device
[0001] The present disclosure relates to a parallel processing system, a parallel processing method, and a parallel processing device.
[0002] Systems for comparing data have been developed. For example, Patent Literature 1 discloses a technology for detecting changes in the shape of a measurement object by comparing a three-dimensional point cloud stored in a point cloud database with a three-dimensional point cloud obtained by measurement using laser light.
[0003] Japanese Patent Application Laid-Open No. 2004-163200
[0004] In Patent Document 1, a comparison is made for each of a plurality of points between a 3D point cloud stored in a point cloud database and a 3D point cloud obtained by measurement using laser light. Thus, when the same process is performed on a plurality of data, it is conceivable to execute the plurality of processes in parallel. However, Patent Document 1 does not mention executing processes in parallel. The present disclosure has been made in consideration of the above-mentioned problems, and one of its purposes is to provide a technology for improving the execution efficiency of a plurality of processes.
[0005] The parallel processing system of the present disclosure includes an acquisition means for acquiring a plurality of target data items each including a first data item and a plurality of second data items, an execution control means for executing, in parallel, a predetermined process for each of the target data items, the first data item and the second data item, with a predetermined number or less of the second data items being the target data items, and a determination means for determining, based on a result of the predetermined process, whether or not any of the second data items that have been the target of the predetermined process satisfy a condition with the first data item. For each of the target data items for which it has been determined that none of the second data items that have been the target of the predetermined process satisfy the condition with the first data item, the execution control means executes, in parallel, the predetermined process for each of the predetermined number or less of the second data items that have not yet been the target of the predetermined process.
[0006] A parallel processing method according to the present disclosure is executed by a computer. The parallel processing method includes: an acquisition step of acquiring multiple pieces of target data, each piece of target data including first data and multiple pieces of second data; an execution control step of executing, in parallel, a predetermined process for the first data and the second data, with a predetermined number or less of the second data being the target data; and a determination step of determining, based on a result of the predetermined process, whether any of the second data that have been the target of the predetermined process satisfy a condition with the first data. For each piece of target data for which it has been determined in the execution control step that none of the second data that have been the target of the predetermined process satisfy the condition with the first data, the predetermined process is executed in parallel with each of the predetermined number or less of the second data that have not yet been the target of the predetermined process.
[0007] The parallel processing device of the present disclosure includes an acquisition means for acquiring a plurality of target data items each including a first data item and a plurality of second data items, an execution control means for executing, in parallel, a predetermined process for each of the target data items, the first data item and the second data item, with a predetermined number or less of the second data items being the target data items, and a determination means for determining, based on a result of the predetermined process, whether or not any of the second data items that have been the target of the predetermined process satisfy a condition with the first data item. For each of the target data items for which it has been determined that none of the second data items that have been the target of the predetermined process satisfy a condition with the first data item, the execution control means executes, in parallel, the predetermined process for each of the predetermined number or less of the second data items that have not yet been the target of the predetermined process.
[0008] According to the present disclosure, a technique for improving the efficiency of executing multiple processes is provided.
[0009] FIG. 1 is a diagram illustrating an example of an overview of the operation of a parallel processing system. FIG. 2 is a diagram conceptually illustrating a state in which predetermined processing is executed in parallel on a plurality of target data. FIG. 3 is a diagram illustrating an example of an advantage of a parallel processing system. FIG. 4 is a block diagram illustrating an example of the functional configuration of a parallel processing system. FIG. 5 is a block diagram illustrating an example of the functional configuration of a parallel processing device. FIG. 6 is a block diagram illustrating an example of the hardware configuration of a computer that realizes the parallel processing system. FIG. 7 is a flowchart illustrating an example of the flow of processing executed by the parallel processing system. FIG. 8 is a diagram illustrating an example of threads and warps. FIG. 9 is a diagram illustrating an example of the functional configuration of a parallel processing system that further includes a generation unit. FIG. 10 is a diagram illustrating an example of an overview of a method for generating target data. FIG. 11 is a flowchart illustrating an example of the flow of processing to generate target data.
[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and duplicate explanations will be omitted as necessary for clarity. Furthermore, unless otherwise specified, predetermined values such as predetermined values and threshold values are stored in advance in a storage device accessible from a device that uses the values. Furthermore, unless otherwise specified, the storage unit is composed of one or any number of storage devices.
[0011] <Overview> Fig. 1 is a diagram illustrating an example of an overview of the operation of a parallel processing system 2000. Note that Fig. 1 is a diagram for facilitating understanding of the overview of the parallel processing system 2000, and the operation of the parallel processing system 2000 is not limited to the operation shown in Fig. 1.
[0012] The parallel processing system 2000 executes a plurality of predetermined processes in parallel on two data sets. For example, the predetermined processes are processes for calculating the similarity between two data sets.
[0013] The parallel processing system 2000 acquires a plurality of target data 10 as targets for the predetermined processing. Each target data 10 includes first data 20 and a plurality of second data 30. Each second data 30 included in the target data 10 is data that is to be subjected to the predetermined processing together with the target data 10 included in that target data 10. In other words, the predetermined processing is executed with the first data 20 and the second data 30 as targets.
[0014] The parallel processing system 2000 attempts to detect, for each of the plurality of target data 10, second data 30 that satisfies a predetermined condition with the first data 20. Here, it is sufficient that one second data 30 that satisfies the predetermined condition with the first data 20 is detected for each target data 10.
[0015] To this end, the parallel processing system 2000 determines, for each of the plurality of target data 10, whether or not there is second data 30 in the target data 10 that satisfies a predetermined condition with the first data 20. This determination is made based on the result of a predetermined process.
[0016] For example, assume that the predetermined process is the process of calculating the similarity between the first data 20 and the second data 30 described above. In this case, for example, the predetermined condition is a condition of "similar to the first data 20." The parallel processing system 2000 determines, for each piece of target data 10, whether or not there is second data 30 that is similar to the first data 20, based on the similarity between the first data 20 and the second data 30 calculated by the predetermined process.
[0017] As described above, the parallel processing system 2000 executes a plurality of predetermined processes in parallel. Specifically, the parallel processing system 2000 executes a predetermined process in parallel for each of a plurality of target data 10, with a predetermined number of second data 30 or less as the target. Hereinafter, this predetermined number will be referred to as the number of parallel processes. Hereinafter, the phrase "performing a predetermined process on the second data 30" means "performing a predetermined process on the second data 30 and the first data 20 included in the target data 10 that includes the second data 30."
[0018] 2 is a conceptual diagram illustrating a situation in which a predetermined process is executed in parallel for a plurality of target data 10. In FIG. 2, target data 10-i represents the ith target data 10. Ai represents the first data 20 included in the target data 10-i. Bij represents the jth second data 30 included in the target data 10-i. Furthermore, f(Ai, Bij) represents the predetermined process executed for Ai and Bij.
[0019] 2, the number of parallel processes is k. Therefore, for each target data 10-i, the parallel processing system 2000 executes k predetermined processes in parallel, such as a predetermined process f(Ai,Bi1) for Ai and Bi1, a predetermined process f(Ai,Bi2) for Ai and Bi2, ..., a predetermined process f(Ai,Bik) for Ai and Bik. For example, for target data 10-1, f(A1,B11), f(A1,B12), ..., f(A1,Bik) are executed.
[0020] Here, if the target data 10 includes the parallel number or more of second data 30 that have not yet been subjected to the predetermined processing, the predetermined processing is executed on the parallel number of second data 30 that have not yet been subjected to the predetermined processing for the target data 10. On the other hand, if the target data 10 does not include the parallel number or more of second data 30 that have not yet been subjected to the predetermined processing, the predetermined processing is executed on all of the second data 30 that have not yet been subjected to the predetermined processing for the target data 10.
[0021] Thereafter, the parallel processing system 2000 determines, for each piece of target data 10, whether or not to exclude that piece of target data 10 from the target of the predetermined processing, based on the result of the predetermined processing executed on each piece of second data 30. Here, if, as a result of the predetermined processing, second data 30 that satisfies a predetermined condition with respect to the first data 20 is detected in a certain piece of target data 10, there is no need to further execute the predetermined processing on that piece of target data 10. On the other hand, if, as a result of the predetermined processing, second data 30 that satisfies a predetermined condition with respect to the first data 20 is not detected in a certain piece of target data 10, there is a need to further execute the predetermined processing on that piece of target data 10.
[0022] Therefore, the parallel processing system 2000 uses the results of the predetermined processing to determine, for each target data 10, whether any of the second data 30 that were the target of the predetermined processing satisfy a predetermined condition with the first data 20. If it is determined that any of the second data 30 satisfy the predetermined condition with the first data 20, the target data 10 is excluded from the target of the predetermined processing. On the other hand, if it is determined that none of the second data 30 satisfy the predetermined condition with the first data 20, the target data 10 is not excluded from the target of the predetermined processing.
[0023] Furthermore, if the predetermined processing has already been performed on all of the second data 30 included in the target data 10, there is no second data 30 remaining to be subjected to the predetermined processing for that target data 10. Therefore, the parallel processing system 2000 also excludes the target data 10 for which all of the second data 30 have been subjected to the predetermined processing from the target of the predetermined processing.
[0024] The parallel processing system 2000 executes the specified processing in parallel for each target data 10 that is the target of the specified processing (in other words, each target data 10 that has not been excluded from the target of the specified processing), targeting each of the second data 30 that has not yet been the target of the specified processing and that is equal to or less than the parallel number.
[0025] Similarly, as long as there is target data 10 to be subjected to the predetermined processing (in other words, until all target data 10 are excluded from the target of the predetermined processing), the parallel processing system 2000 repeats the following two processes: (1) For each target data 10 to be subjected to the predetermined processing, execute the predetermined processing in parallel for each second data 30 up to the parallel number; and (2) Based on the result of the predetermined processing, determine whether or not to exclude each target data 10 from the target of the predetermined processing.
[0026] As described above, the condition for the target data 10 to be excluded from the predetermined processing is that either of the following two conditions is satisfied: (1) Second data 30 that satisfies the predetermined condition with the first data 20 is detected, or (2) The predetermined processing is executed for all of the second data 30.
[0027] <Example of Action and Effect> According to the parallel processing system 2000 of this embodiment, for each of the target data 10 that is the target of the predetermined processing, a predetermined processing is executed in parallel for each of the parallel number of second data 30. In this manner, parallel execution of a plurality of predetermined processing is repeatedly executed. Here, the target data 10 for which the second data 30 that satisfies the predetermined condition with the first data 20 has been detected, and the target data 10 for which the predetermined processing has already been executed for all of the second data 30 are excluded from the target of the predetermined processing.
[0028] In this way, by executing the predetermined process in parallel for each of the parallel number of second data 30, the total number of predetermined processes to be executed can be reduced. As a result, the predetermined processes can be executed more efficiently. The reason for this will be explained below.
[0029] 3 is a diagram illustrating an example of the advantages of the parallel processing system 2000. The target data 10 in FIG. 3 includes one first data item A1 and eight second data items B11 to B18. The second data item B13 satisfies a predetermined condition with the first data item A1.
[0030] The upper part of Fig. 3 shows a case where parallel processing is performed simultaneously on all second data 30 included in the target data 10. On the other hand, the lower part of Fig. 3 shows a case where the number of parallel processes is four. That is, in the case of the lower part, predetermined processing is first performed on four second data 30, namely, second data B11 to B14.
[0031] In this example, the execution control unit 2040 detects that the second data B13 satisfies a predetermined condition with the first data A1. Therefore, the predetermined processing is not executed for the second data B15 to B18. In this way, by limiting the number of second data 30 to be processed in parallel at one time, rather than all second data 30, the total number of predetermined processing operations to be executed can be reduced. By reducing the total number of predetermined processing operations to be executed, the predetermined processing operations can be executed more efficiently. As a result, the time required for all target data 10 to be excluded from the processing targets, i.e., the time required for the parallel processing system 2000 to complete the processing, is shortened. Furthermore, the utilization efficiency of computer resources by the parallel processing system 2000 can be improved.
[0032] The parallel processing system 2000 of this embodiment will be described in more detail below.
[0033] <Example of Functional Configuration> FIG. 4 is a block diagram illustrating an example of the functional configuration of the parallel processing system 2000. The parallel processing system 2000 includes an acquisition unit 2020, an execution control unit 2040, and a determination unit 2060. The acquisition unit 2020 acquires multiple pieces of target data 10. For each piece of target data 10 to be processed, the execution control unit 2040 executes, in parallel, a predetermined process targeting each of the second data 30 that are not yet processed in parallel, but are equal to or less than the parallel number. The determination unit 2060 determines whether each piece of second data 30 that is the target of the predetermined process satisfies a predetermined condition with the first data 20. For each piece of target data 10 for which it is determined that none of the second data 30 that is the target of the predetermined process satisfies the predetermined condition with the first data 20, the execution control unit 2040 executes, in parallel, a predetermined process targeting each of the second data 30 that are not yet processed in parallel, but are equal to or less than the parallel number.
[0034] Here, the parallel processing system 2000 may be realized by a plurality of devices or by a single device. One device that realizes the parallel processing system 2000 is called a parallel processing device. FIG. 5 is a block diagram illustrating an example of the functional configuration of a parallel processing device. The parallel processing device 3000 has an acquisition unit 2020 and an execution control unit 2040.
[0035] <Example of Hardware Configuration> Each functional component of the parallel processing system 2000 may be realized by hardware that realizes the functional component (e.g., a hardwired electronic circuit, etc.), or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it, etc.). Below, a case where each functional component of the parallel processing system 2000 is realized by a combination of hardware and software will be further described.
[0036] 6 is a block diagram illustrating an example of the hardware configuration of a computer 1000 that implements the parallel processing system 2000. The computer 1000 is any computer. For example, the computer 1000 is a stationary computer such as a PC (Personal Computer) or a server machine. Alternatively, the computer 1000 may be a portable computer such as a smartphone or a tablet terminal. The computer 1000 may be a dedicated computer designed to implement the parallel processing system 2000, or may be a general-purpose computer.
[0037] For example, by installing a predetermined application on the computer 1000, the computer 1000 realizes each function of the parallel processing system 2000. The application is configured with a program for realizing each functional component of the parallel processing system 2000. Note that any method for acquiring the program is possible. For example, the program can be acquired from a storage medium (such as a DVD disk or USB memory) on which the program is stored. Alternatively, the program can be acquired by downloading the program from a server device that manages the storage device on which the program is stored.
[0038] The computer 1000 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 to transmit and receive data to and from each other. However, the method of connecting the processor 1040 and the like to each other is not limited to bus connection.
[0039] The processor 1040 is a processor such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device realized using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device realized using a hard disk, a solid state drive (SSD), a memory card, a read only memory (ROM), or the like.
[0040] The input / output interface 1100 is an interface for connecting the computer 1000 to an input / output device. For example, the input / output interface 1100 is connected to an input device such as a keyboard and an output device such as a display device.
[0041] The network interface 1120 is an interface for connecting the computer 1000 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).
[0042] The storage device 1080 stores programs (programs that realize the above-mentioned applications) that realize the various functional components of the parallel processing system 2000. The processor 1040 reads these programs into the memory 1060 and executes them to realize the various functional components of the parallel processing system 2000.
[0043] The parallel processing system 2000 may be realized by one computer 1000 or by multiple computers 1000. In the former case, it can be said that the parallel processing device 3000 is realized by one computer 1000 that realizes the parallel processing system 2000. In the latter case, the configurations of the computers 1000 do not need to be the same, and can be different from each other.
[0044] 7 is a flowchart illustrating the flow of processing executed by the parallel processing system 2000. The acquisition unit 2020 acquires a plurality of target data 10 (S102).
[0045] Steps S104 to S114 constitute a loop process L1, which is repeatedly executed while target data 10 to be subjected to the predetermined process exists.
[0046] In S104, the execution control unit 2040 determines whether or not there is target data 10 to be subjected to the predetermined process. Note that in the first execution of the loop process L1, all of the target data 10 is subjected to the predetermined process.
[0047] If there is no target data 10 to be subjected to the predetermined processing, the execution control unit 2040 ends the execution of the loop processing L1. If there is target data 10 to be subjected to the predetermined processing, step S106 is executed.
[0048] The execution control unit 2040 executes the predetermined processing in parallel for each of the target data 10 to be processed, for each of the second data 30 that is not yet processed but is equal to or less than the parallel number (S106). The determination unit 2060 determines whether or not a predetermined condition is satisfied between each of the second data 30 that has been processed and the first data 20 (S108). The execution control unit 2040 excludes from the target of the predetermined processing any of the target data 10 for which it has been determined that one or more of the second data 30 satisfies the predetermined condition between the second data 30 and the first data 20 (S110). The execution control unit 2040 excludes from the target of the predetermined processing any of the target data 10 for which the predetermined processing has already been performed for all of the second data 30 (S112).
[0049] Step S114 is the end of the loop process L1, so step S104 is executed again.
[0050] <Acquisition of Target Data 10: S102> The acquisition unit 2020 acquires a plurality of pieces of target data 10 (S102). The acquisition unit 2020 can acquire the target data 10 by any method. For example, the target data 10 is stored in advance in a storage unit accessible from the parallel processing system 2000. In this case, the acquisition unit 2020 acquires the target data 10 by reading the target data 10 from the storage unit.
[0051] Alternatively, for example, the target data 10 may be transmitted from another device (e.g., a device that generated the target data 10) to the parallel processing system 2000. In this case, the acquisition unit 2020 acquires the target data 10 by receiving the target data 10 transmitted from the other device.
[0052] As will be described later, the target data 10 may be generated by the parallel processing system 2000. In this case, the parallel processing system 2000 acquires the target data 10 by acquiring the target data 10 that it has generated itself.
[0053] <Execution of Predetermined Processing: S106> The execution control unit 2040 executes, in parallel, the predetermined processing for each of the second data 30 that is not yet the target of the predetermined processing, up to the parallel number, for the target data 10 that is the target of the predetermined processing (S104). The target data 10 that is the target of the predetermined processing is managed, for example, using a list. Hereinafter, this list will be referred to as a target list. The target list is data that indicates the identifier of each of the target data 10 that is the target of the predetermined processing. Note that the method for managing the target data 10 that is the target of the predetermined processing is arbitrary and is not limited to management using a list.
[0054] The execution control unit 2040 executes a predetermined process on each of the target data 10 shown in the target list. To do so, the execution control unit 2040 extracts the second data 30 to be used as the target of the predetermined process from each of the target data 10 shown in the target list.
[0055] It is assumed that the number of second data 30 that have not yet been subjected to the predetermined processing in the target data 10 is equal to or greater than the parallel number. In this case, the execution control unit 2040 extracts, from the target data 10, the same number of second data 30 that have not yet been subjected to the predetermined processing as the parallel number.
[0056] For example, suppose the number of parallel processes is 5. Also, suppose the target data 10 includes 10 pieces of second data 30 that have not yet been subjected to the predetermined processing. In this case, the execution control unit 2040 extracts five pieces of second data 30 from the target data 10, out of the ten pieces of second data 30 that have not yet been subjected to the predetermined processing, as the pieces of second data 30 to be subjected to the predetermined processing.
[0057] Any method may be used to determine the second data 30 to be extracted from the target data 10. For example, the execution control unit 2040 may extract the second data 30 in a random order from the target data 10. Alternatively, the execution control unit 2040 may extract the second data 30 in a predetermined order from the target data 10. In the latter case, for example, if the second data 30 is presented in an ordered order in the target data 10, the execution control unit 2040 extracts the second data 30 in accordance with that order.
[0058] It is assumed that the number of second data 30 that have not yet been subjected to the predetermined processing is less than the parallel number in the target data 10. In this case, the execution control unit 2040 extracts all second data 30 that have not yet been subjected to the predetermined processing from the target data 10.
[0059] For example, suppose the number of parallel processes is 5. Also, suppose the target data 10 includes three pieces of second data 30 that have not yet been subjected to the predetermined processing. In this case, the execution control unit 2040 extracts all three pieces of second data 30 that have not yet been subjected to the predetermined processing from the target data 10 as targets for the predetermined processing.
[0060] The execution control unit 2040 executes the predetermined processing in parallel on the second data 30 extracted from each target data 10 that is the target of the predetermined processing. A method for executing the predetermined processing in parallel will be specifically described below. Here, a case where a GPU is used to execute the predetermined processing will be described.
[0061] The execution control unit 2040 generates a thread for each pair of first data 20 and second data 30. An operand and a process to be executed on the operand are assigned to the thread. The operands assigned to the thread are the first data 20 and the second data 30. The process assigned to the thread is a predetermined process to be executed on the first data 20 and the second data 30. Using the notation in FIG. 2, the operands assigned to the thread are Ai and Bij, and the process assigned to the thread is f.
[0062] The execution control unit 2040 generates a thread for each of all second data 30 extracted as targets for the predetermined process. Specifically, for each piece of second data 30 extracted as targets for the predetermined process, the execution control unit 2040 generates a thread by pairing the first data 20 corresponding to that second data 30 with that second data 30.
[0063] The execution control unit 2040 then causes the GPU to execute all threads, thereby executing the predetermined process in parallel for all of the second data 30 extracted as the target of the predetermined process.
[0064] In a GPU, scheduling is performed in units of groups of threads called warps. If the number of generated threads is greater than the warp size (the number of threads included in one warp), the execution control unit 2040 divides the generated threads into multiple warps.
[0065] Here, threads included in a warp access consecutive areas in the GPU memory (performing coreless memory access), thereby improving memory access efficiency. The first data 20 is common to multiple threads generated for one target data 10. Therefore, by including all multiple threads generated for one target data 10 in the same warp, memory access efficiency can be improved. Therefore, it is preferable for the execution control unit 2040 to include multiple threads generated for one target data 10 in the same warp.
[0066] 8 is a diagram illustrating threads and warps. In FIG. 8, the number of target data 10 to be processed is m. For each target data 10, k threads 40 are generated.
[0067] 8, the warp size is 2k, so one warp 50 includes threads 40 generated from two target data 10, respectively.
[0068] <Regarding the Number of Parallel Processes> As will be explained below, the number of parallel processes can affect the efficiency of processing by the parallel processing system 2000. First, by reducing the number of parallel processes, the number of second data 30 executed in parallel at one time for each target data 10 decreases. Therefore, by reducing the number of parallel processes, the total number of predetermined processes executed can be reduced.
[0069] For example, in the example of Figure 3, the number of parallel processes is four, so the predetermined process is executed four times. However, in this example, by setting the number of parallel processes to three, the number of times the predetermined process is executed can be reduced to three. This is because, even when the number of parallel processes is three, f(A1, B13) is executed in the first parallel execution of the predetermined process, so it is determined that B13 satisfies the predetermined condition with A1.
[0070] On the other hand, increasing the number of parallel processes increases the efficiency of memory access in each warp. For example, assume that the warp size is 16 and the number of parallel processes is 2. In this case, the number of first data 20 accessed in one warp is 8 (16 / 2). On the other hand, assume that the warp size is 16 and the number of parallel processes is 4. In this case, the number of first data 20 accessed in one warp is 4 (16 / 4). In this way, in the latter case where the number of parallel processes is larger, the range of memory access in the execution of the warp is narrower, and therefore the efficiency of memory access is higher.
[0071] In view of the above, it is preferable to determine the number of parallel processes by considering the balance between the effect of reducing the total number of specified processes to be executed by reducing the number of parallel processes, and the effect of improving memory access efficiency by increasing the number of parallel processes.
[0072] For example, the number of parallel processes may be statically set in advance. For example, the number of parallel processes may be set to a multiple of one warp size (e.g., two or three times the warp size). Alternatively, an appropriate number of parallel processes may be determined through test execution of the parallel processing system 2000.
[0073] Here, if the number of parallel threads is larger than the warp size, multiple threads generated from one target data 10 are divided into multiple warps. For example, assume that the warp size is w and the number of parallel threads is k=2w.
[0074] Under this assumption, it is assumed that the number of first data 20 that have not yet been subjected to predetermined processing in the target data 10 is 2w or more. In this case, the 2w threads generated for the target data 10 will fill exactly two warps.
[0075] On the other hand, suppose that the number of first data 20 that has not yet been subjected to the predetermined processing in the target data 10 is greater than w and less than 2w. In this case, one warp that is satisfied by the thread generated for the target data 10 and one warp that is not satisfied by the thread generated for the target data 10 are generated.
[0076] As in the latter case, when there is a warp that cannot be satisfied by threads generated for one piece of target data 10, for example, the execution control unit 2040 causes the warp to be satisfied by including threads generated for target data 10 other than the target data 10 in the warp. As another example, the execution control unit 2040 may not assign threads generated for target data 10 other than the target data 10 to a warp that cannot be satisfied by threads generated for one piece of target data 10. In this case, the number of threads assigned to the warp is less than the warp size.
[0077] The number of parallel processes may be dynamically determined. For example, the execution control unit 2040 increases the number of parallel processes as the number of target data 10 excluded from the target of the predetermined process increases.
[0078] For example, the number of parallel connections is determined by the following equation (1). In equation (1), k0 is the initial value of the number of parallel processes. k0 is set in advance. r is the ratio of the number of target data 10 acquired by the acquisition unit 2020 to the number of target data 10 that are the target of the predetermined process.
[0079] If the number of target data 10 that are the target of the predetermined processing is represented by x and the number of target data 10 acquired by the acquisition unit 2020 is represented by y, then r=y / x. The smaller the denominator x, i.e., the fewer the target data 10 that are the target of the predetermined processing, the larger r becomes.
[0080] The method for dynamically determining the number of parallel processes is not limited to the method using equation 1. For example, the execution control unit 2040 may determine the number of parallel processes based on the distribution of neighboring points.
[0081] <Exclusion of Target Data 10 from Targets of Predetermined Processing: S108 to S112> The execution control unit 2040 determines whether a predetermined condition is satisfied between each of the second data 30 that was subject to the predetermined processing in S106 and the corresponding first data 20 (S108). For example, the predetermined condition may be "similar to the first data 20." In this case, the determination unit 2060 determines whether each of the second data 30 that was subject to the predetermined processing is similar to the first data 20.
[0082] The execution control unit 2040 deletes from the target list the target data 10 for which it has been determined that one or more pieces of second data 30 satisfy the predetermined condition with the first data 20. As a result, the target data 10 for which it has been determined that one or more pieces of second data 30 satisfy the predetermined condition with the first data 20 is excluded from the target of the predetermined process (S110).
[0083] Here, the remaining determination processes that have not yet been performed may be omitted for target data 10 in which second data 30 that satisfies the predetermined condition with the first data 20 has been detected. For example, among the second data 30 that have been subjected to the predetermined process in certain target data 10, the second data 30 that has been subjected to the determination process second is determined to satisfy the predetermined condition with the first data 20. In this case, since it has already been determined that the target data 10 should be excluded from the target of the predetermined process, the execution control unit 2040 does not need to determine whether the third and subsequent second data 30 satisfy the predetermined condition with the first data 20.
[0084] Furthermore, the execution control unit 2040 deletes from the target list the target data 10 for which the predetermined process has already been executed on all of the first data 20. As a result, the target data 10 for which the predetermined process has already been executed on all of the first data 20 is excluded from the target list (S112).
[0085] <Output of Results> The parallel processing system 2000 outputs information (hereinafter, output information) indicating the results of processing performed on the target data 10. For example, the output information indicates each second data 30 determined to satisfy a predetermined condition with the first data 20, in association with the first data 20. Alternatively, for each second data 30 determined to satisfy a predetermined condition with the first data 20, the output information may indicate an identifier of the second data 30 in association with an identifier of the target data 10 including the second data 30.
[0086] The output information may indicate the identifier of the target data 10 for which no second data 30 that satisfies a predetermined condition between the target data 10 and the first data 20 has been detected.
[0087] There are various methods for outputting the output information. For example, the parallel processing system 2000 stores the output information in an arbitrary storage unit. Alternatively, for example, the parallel processing system 2000 displays the output information on an arbitrary display device. Alternatively, for example, the parallel processing system 2000 transmits the output information to another device.
[0088] <Generation of Target Data 10> The generation of the target data 10 may be performed by the parallel processing system 2000, or may be performed by a system other than the parallel processing system 2000. Hereinafter, the functional configuration unit that generates the target data 10 will be referred to as a generation unit. The generation unit is provided in the system that generates the target data 10.
[0089] 9 is a diagram illustrating an example of the functional configuration of a parallel processing system 2000 that further includes a generation unit. In FIG. 9, the parallel processing system 2000 has a generation unit 2080. When the parallel processing system 2000 is realized by a parallel processing device 3000, the parallel processing device 3000 has the generation unit 2080.
[0090] Below, specific examples of the target data 10 generated by the generation unit 2080 will be exemplified together with methods of using and generating the target data 10. However, the target data 10 handled by the parallel processing system 2000 is not limited to the examples exemplified below.
[0091] FIG. 10 is a diagram illustrating an example of an outline of a method for generating target data 10. In the example of FIG. 10, target data 10 is generated from point cloud data (point cloud data 60 and point cloud data 70) generated at two different times. Both point cloud data 60 and point cloud data 70 are generated by measuring the distance to a measurement target using a distance measuring device such as a LiDAR (Light Detection and Ranging). Therefore, point cloud data 60 and point cloud data 70 each represent the three-dimensional shape of a target object.
[0092] The measurement target may be any facility (e.g., a logistics warehouse, a factory, a power plant, or the like) that includes multiple objects.
[0093] The point cloud data 60 and the point cloud data 70 are generated at time t1 and time t2, respectively. Time t2 is a time later than time t1. In other words, the point cloud data 70 is generated at a later time than the point cloud data 60.
[0094] The target data 10 is used to detect the difference between the point cloud data 60 and the point cloud data 70. In other words, the target data 10 is used to detect a location of the measurement target whose position has changed from time t1 to time t2.
[0095] Specifically, for each point data 62 in the point cloud data 60, the parallel processing system 2000 detects point data 72 in the point cloud data 70 that represents a neighboring point of the point data 62. Then, the point data 62 is compared with the point data 72 that represents a neighboring point of the point data 62.
[0096] 10, point data 72-1, point data 72-2, point data 72-3, and point data 72-4 are detected as point cloud data 70 representing neighboring points of point data 62-1. Therefore, point data 62-1 is compared with each of point data 72-1 to point data 72-4.
[0097] To compare the point data 62 and the point data 72, the parallel processing system 2000 calculates a feature vector representing the spatial features of each of the point data 62 and the point data 72. If the two point data represent the same location on the measurement object, the feature vectors of these two point data will be similar to each other. On the other hand, if the two point data represent different locations on the measurement object, the feature vectors of these two point data will be dissimilar to each other.
[0098] Therefore, the parallel processing system 2000 compares the feature vector of the point data 62 with the feature vector of the point data 72 representing a neighboring point of the point data 62 to determine whether the location represented by the point data 62 and the location represented by the point data 72 are the same location. For example, in the example of Figure 10, the feature vector of the point data 62-1 is compared with the feature vector of each of the point data 72-1 to 72-4. This determines whether or not there is point data 72 representing the same location as the location represented by the point data 62 among the point data 72-1 to 72-4.
[0099] Similarly, the parallel processing system 2000 compares each point data 72 included in the point cloud data 70 with point data 62 representing neighboring points of the point data 72 detected from the point cloud data 60 .
[0100] In order to compare the point cloud data 60 and the point cloud data 70, the parallel processing system 2000 generates target data 10. FIG. 11 is a flowchart illustrating the flow of a process for generating the target data 10. Steps S202 to S208 constitute a loop process L2. In the loop process L2, the target data 10 is generated in which the feature vector of the point data 62 is the first data 20 and the feature vector of the point data 72 is the second data 30.
[0101] In one execution of the loop process L2, any one of the point data 62 is selected as a candidate for the first data 20. The generation unit 2080 executes the loop process L2 for each of all the point data 62, in which the point data 62 is selected as a candidate for the first data 20.
[0102] In step S202, the generation unit 2080 determines whether or not all of the point data 62 have already been selected as candidates for the first data 20. If all of the point data 62 have already been selected as candidates for the first data 20, the generation unit 2080 ends the execution of the loop process L2.
[0103] If there is point data 62 that has not yet been selected as a candidate for the first data 20, the generation unit 2080 selects one of the point data 62 that has not yet been selected as a candidate for the first data 20 as a candidate for the first data 20. Here, the point data 62 selected as a candidate for the first data 20 is represented as point data Pi.
[0104] The generating unit 2080 detects neighboring points of the point data Pi from the point cloud data 70 (S204). The neighboring points of the point data Pi are, for example, point data 72 whose distance from the point data Pi is equal to or less than a predetermined threshold.
[0105] The generation unit 2080 calculates feature vectors for each of the point data Pi and the point cloud data 70 of neighboring points of the point data Pi (S206). Here, in the second and subsequent loop processing L2, it is not necessary to calculate feature vectors again for point data for which feature vectors have already been calculated in step S206 executed so far.
[0106] The generation unit 2080 generates target data 10 in which the feature vector of point data Pi is the first data 20 and the feature vectors of each neighboring point are the second data 30 (S208). For example, in the example of Fig. 10, the target data 10 is generated in which the feature vector of point data 62-1 is the first data 20 and the feature vectors of point data 72-1, point data 72-2, point data 72-3, and point data 72-4 are the second data 30.
[0107] Since step S210 is the end of loop processing L2, next, step S202 is executed again.
[0108] Steps S212 to S220 constitute loop processing L3. In loop processing L3, target data 10 is generated in which the feature vector of point data 72 is first data 20 and the feature vector of point data 62 is second data 30. Loop processing L3 is processing in which the positions of point cloud data 60 and point cloud data 70 are swapped in loop processing L2.
[0109] In one execution of the loop process L3, any one of the point data 72 is selected as a candidate for the first data 20. The generation unit 2080 executes the loop process L3 for each of all the point data 72, in which the point data 72 is selected as a candidate for the first data 20.
[0110] In step S212, the generation unit 2080 determines whether or not all of the point data 72 have already been selected as candidates for the first data 20. If all of the point data 72 have already been selected as candidates for the first data 20, the generation unit 2080 ends the execution of the loop process L3.
[0111] If there is point data 72 that has not yet been selected as a candidate for the first data 20, the generation unit 2080 selects one of the point data 72 that has not yet been selected as a candidate for the first data 20 as a candidate for the first data 20. Here, the point data 72 selected as a candidate for the first data 20 is represented as point data Qi.
[0112] The generation unit 2080 detects neighboring points of point data Qi from the point cloud data 60 (S214). The generation unit 2080 calculates feature vectors for each of point data Qi and its neighboring points (S216). Here, for point data for which feature vectors have already been calculated in step S206 or any of the steps S216 executed up to this point, there is no need to calculate feature vectors again.
[0113] The generating unit 2080 generates target data 10 in which the feature vector of the point data Qi is the first data 20 and the feature vectors of each neighboring point are the second data 30 (S218).
[0114] Since step S220 is the end of loop processing L3, next, step S212 is executed again.
[0115] 11, the parallel processing system 2000 detects changes in the measurement target from time t1 to time t2. To this end, the parallel processing system 2000 uses the target data 10 generated in S208. The predetermined condition is that the target data 10 be "similar to the first data 20." The predetermined process is a process of calculating a value representing the similarity between the first data 20 and the second data 30 (for example, the distance between the feature vectors).
[0116] First, the parallel processing system 2000 uses the target data 10 generated in the loop process L2. If there is second data 30 that satisfies a predetermined condition with the first data 20, the feature vector of the point data Pi and the feature vector of the neighboring point of the point data Pi are similar to each other. Therefore, it can be seen that the location of the measurement target represented by the point data Pi at time t1 has not changed between time t1 and time t2.
[0117] On the other hand, if there is no second data 30 that satisfies the predetermined condition with the first data 20, there is no neighboring point having a feature vector similar to the feature vector of the point data Pi. Therefore, the parallel processing system 2000 detects the location of the measurement object represented by the point data Pi at time t1 as a location that changed between time t1 and time t2.
[0118] Next, the parallel processing system 2000 uses the target data 10 generated in loop processing L3. If there is second data 30 that satisfies a predetermined condition with the first data 20, the feature vector of point data Qi and the feature vector of a neighboring point of point data Qi are similar to each other. Therefore, it can be seen that the location of the measurement target represented by point data Qi at time t2 has not changed between time t1 and time t2.
[0119] On the other hand, if there is no second data 30 that satisfies the predetermined condition with the first data 20, there is no neighboring point that has a feature vector similar to the feature vector of the point data Qi. Therefore, the parallel processing system 2000 detects the location of the measurement object represented by the point data Qi at time t2 as a location that has changed between time t1 and time t2.
[0120] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0121] In the above examples, the program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray (registered trademark) disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.
[0122] Each drawing is merely an example for describing one or more embodiments. Each drawing may not relate to only one particular embodiment, but may also relate to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.
[0123] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes: (Supplementary Note 1) A parallel processing system comprising: acquisition means for acquiring a plurality of target data including first data and a plurality of second data; execution control means for executing, in parallel, a predetermined process targeting the first data and the second data for each of the target data, with a predetermined number or less of the second data as the target; and determination means for determining, based on a result of the predetermined process, for each of the target data, whether any of the second data that have been subjected to the predetermined process satisfy a condition with the first data, wherein the execution control means executes, in parallel, the predetermined process targeting the predetermined number or less of the second data that have not yet been subjected to the predetermined process for each of the target data for which it has been determined that none of the second data that have been subjected to the predetermined process satisfy a condition with the first data. (Supplementary Note 2) The parallel processing system according to Supplementary Note 1, wherein the execution control means excludes from the target of the predetermined processing each of the target data for which it is determined that one or more of the second data subjected to the predetermined processing satisfies a condition with the first data, and each of the target data for which the predetermined processing has already been performed for all of the plurality of second data. (Supplementary Note 3) The parallel processing system according to Supplementary Note 1 or 2, wherein the predetermined processing is processing for calculating a value representing the degree of similarity between the first data and the second data, and the condition is a condition that the second data is similar to the first data. (Supplementary Note 4) The parallel processing system according to Supplementary Note 1 or 2, wherein the execution control means increases the predetermined number as the number of target data subjected to the predetermined processing decreases. (Supplementary Note 5) The parallel processing system according to Supplementary Note 4, wherein the execution control means determines the predetermined number based on a ratio between the number of target data subjected to the predetermined processing and the number of target data acquired by the acquisition means.(Supplementary Note 6) The parallel processing system according to Supplementary Note 1 or 2, wherein the execution control means generates, for each of the target data, a thread that uses the second data and the first data as operands for each of the second data to be subjected to the predetermined processing and that executes the predetermined processing, and causes a GPU (Graphics Processing Unit) to execute each of the generated threads. (Supplementary Note 7) The parallel processing system according to Supplementary Note 6, wherein the execution control means assigns multiple threads generated for the same target data to the same warp. (Supplementary Note 8) The parallel processing system according to Supplementary Note 1 or 2, wherein the first data represents feature vectors of point data included in first point cloud data generated by measuring a measurement object at a first time point, and the second data represents feature vectors of point data that represent neighboring points of point data in the first point cloud data that corresponds to the first data, among point data included in second point cloud data generated by measuring the measurement object at a second time point. (Supplementary Note 9) A parallel processing method executed by a computer, comprising: an acquisition step of acquiring multiple pieces of target data including first data and multiple pieces of second data; an execution control step of executing, for each of the target data, a predetermined process targeting the first data and the second data in parallel, targeting a predetermined number or less of the second data; and a determination step of determining, for each of the target data, based on a result of the predetermined process, whether any of the second data that have been targeted for the predetermined process satisfy a condition with the first data, wherein, for each of the target data for which it has been determined in the execution control step that none of the second data that have been targeted for the predetermined process satisfy a condition with the first data, the predetermined process is executed in parallel, targeting each of the predetermined number or less of the second data that have not yet been targeted for the predetermined process.(Supplementary Note 10) The parallel processing method according to Supplementary Note 9, wherein, in the execution control step, each of the target data for which it is determined that one or more of the second data targeted for the predetermined processing satisfies a condition with the first data, and each of the target data for which the predetermined processing has already been executed for all of the plurality of second data, are excluded from the target of the predetermined processing. (Supplementary Note 11) The parallel processing method according to Supplementary Note 9 or 10, wherein the predetermined processing is processing for calculating a value representing a degree of similarity between the first data and the second data, and the condition is a condition that the second data is similar to the first data. (Supplementary Note 12) The parallel processing method according to Supplementary Note 9 or 10, wherein, in the execution control step, the fewer the target data targeted for the predetermined processing, the greater the predetermined number is set. (Supplementary Note 13) The parallel processing method according to Supplementary Note 12, wherein, in the execution control step, the predetermined number is determined based on a ratio between the number of target data targeted for the predetermined processing and the number of target data acquired in the acquisition step. (Supplementary Note 14) The parallel processing method according to Supplementary Note 9 or 10, wherein in the execution control step, for each of the second data to be subjected to the predetermined processing, a thread is generated, the second data and the first data being operands and the processing to be executed is the predetermined processing, and each of the generated threads is executed by a GPU (Graphics Processing Unit). (Supplementary Note 15) The parallel processing method according to Supplementary Note 14, wherein in the execution control step, multiple threads generated for the same set of target data are assigned to the same warp. (Supplementary Note 16) The parallel processing method according to Supplementary Note 9 or 10, wherein the first data represents feature vectors of point data included in first point cloud data generated by measuring a measurement object at a first time point, and the second data represents feature vectors of point data representing neighboring points of point data in the first point cloud data corresponding to the first data, among point data included in second point cloud data generated by measuring the measurement object at a second time point.(Supplementary Note 17) A parallel processing device comprising: an acquisition means for acquiring a plurality of target data including first data and a plurality of second data; an execution control means for executing, for each of the target data, a predetermined process targeting the first data and the second data in parallel, targeting a predetermined number or less of the second data; and a determination means for determining, for each of the target data, based on a result of the predetermined process, whether any of the second data that have been subject to the predetermined process satisfy a condition with the first data, wherein the execution control means executes in parallel the predetermined process targeting each of the predetermined number or less of the second data that have not yet been subject to the predetermined process, for each of the target data for which it has been determined that none of the second data that have been subject to the predetermined process satisfy the condition with the first data. (Supplementary Note 18) The parallel processing device according to Supplementary Note 17, wherein the execution control means excludes, from the target data for which it has been determined that one or more of the second data that have been subject to the predetermined process satisfy the condition with the first data, and each of the target data for which the predetermined process has already been performed for all of the plurality of second data. (Supplementary Note 19) The parallel processing device according to Supplementary Note 17 or 18, wherein the predetermined processing is processing for calculating a value representing a degree of similarity between the first data and the second data, and the condition is a condition that the second data is similar to the first data. (Supplementary Note 20) The parallel processing device according to Supplementary Note 17 or 18, wherein the execution control means increases the predetermined number as the number of target data to be subjected to the predetermined processing decreases. (Supplementary Note 21) The parallel processing device according to Supplementary Note 20, wherein the execution control means determines the predetermined number based on a ratio between the number of target data to be subjected to the predetermined processing and the number of target data acquired by the acquisition means. (Supplementary Note 22) The parallel processing device according to Supplementary Note 17 or 18, wherein the execution control means generates, for each of the target data, a thread that uses the second data and the first data as operands and that executes the predetermined processing for each of the second data to be subjected to the predetermined processing, and causes a GPU (Graphics Processing Unit) to execute each of the generated threads.(Supplementary Note 23) The parallel processing device according to Supplementary Note 22, wherein the execution control means assigns the plurality of threads generated for the same target data to the same warp. (Supplementary Note 24) The parallel processing device according to Supplementary Note 17 or 18, wherein the first data represents a feature vector of point data included in first point cloud data generated by measuring the measurement object at a first time point, and the second data represents a feature vector of point data representing a neighboring point of the point data of the first point cloud data corresponding to the first data, among point data included in second point cloud data generated by measuring the measurement object at a second time point.
[0124] REFERENCE SIGNS LIST 10 Target data 20 First data 30 Second data 40 Thread 50 Warp 60 Point cloud data 62 Point data 70 Point cloud data 72 Point data 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface 2000 Parallel processing system 2020 Acquisition unit 2040 Execution control unit 2060 Determination unit 2080 Generation unit 3000 Parallel processing device
Claims
1. An acquisition means for acquiring a plurality of target data including first data and a plurality of second data; an execution control means for, for each of the target data, executing in parallel a predetermined process targeting the first data and the second data for each of the second data of a predetermined number or less; and a determination means for determining, based on the result of the predetermined process, whether any of the second data targeted by the predetermined process satisfies a condition with the first data for each of the target data. The execution control means executes in parallel the predetermined process targeting each of the second data of the predetermined number or less that has not yet been targeted by the predetermined process for each of the target data for which it is determined that none of the second data targeted by the predetermined process satisfies the condition with the first data. A parallel processing system.
2. The execution control means excludes from the targets of the predetermined process each of the target data for which it is determined that any one or more of the second data targeted by the predetermined process satisfies the condition with the first data, and each of the target data for which the predetermined process has already been executed for all of the plurality of the second data. The parallel processing system according to claim 1.
3. The predetermined process is a process for calculating a value representing the degree of similarity between the first data and the second data, and the condition is the condition that the second data is similar to the first data. The parallel processing system according to claim 1 or 2.
4. The execution control means increases the predetermined number as the number of the target data targeted by the predetermined process decreases. The parallel processing system according to claim 1 or 2.
5. The execution control means determines the predetermined number based on the ratio between the number of the target data targeted by the predetermined process and the number of the target data acquired by the acquisition means. The parallel processing system according to claim 4.
6. The execution control means generates, for each of the target data, a thread that uses the second data and the first data as operands and whose process to be executed is the predetermined process for each of the second data targeted by the predetermined process, and causes each of the generated threads to be executed by a GPU (Graphics Processing Unit). The parallel processing system according to claim 1 or 2.
7. The parallel processing system according to claim 6, wherein the execution control means assigns a plurality of the threads generated for the same target data to the same warp.
8. The first data represents a feature vector of point data included in first point cloud data generated by measuring a measurement target at a first time point, and the second data represents a feature vector of point data included in second point cloud data generated by measuring the measurement target at a second time point, and represents a feature vector of point data representing a neighboring point of the point data of the first point cloud data corresponding to the first data. The parallel processing system according to claim 1 or 2.
9. An acquisition step of acquiring a plurality of target data including first data and a plurality of second data; an execution control step of, for each of the target data, executing in parallel a predetermined process targeting the first data and the second data for each of the second data equal to or less than a predetermined number; and a determination step of determining, based on the result of the predetermined process, whether any of the second data targeted by the predetermined process satisfies a condition with the first data for each of the target data. In the execution control step, for each of the target data determined that none of the second data targeted by the predetermined process satisfies the condition with the first data, the predetermined process targeting each of the second data equal to or less than the predetermined number that has not yet been targeted by the predetermined process is executed in parallel. A parallel processing method executed by a computer.
10. In the execution control step, for each of the target data determined that any one or more of the second data targeted by the predetermined process satisfies the condition with the first data, and for each of the target data for which the predetermined process has already been executed for all of the plurality of the second data, the target of the predetermined process is excluded. The parallel processing method according to claim 9.
11. The predetermined process is a process of calculating a value representing the degree of similarity between the first data and the second data, and the condition is the condition that the second data is similar to the first data. The parallel processing method according to claim 9 or 10.
12. In the execution control step, the larger the number of the target data targeted by the predetermined process is reduced, the larger the predetermined number is. The parallel processing method according to claim 9 or 10.
13. The parallel processing method according to claim 12, wherein in the execution control step, the predetermined number is determined based on a ratio between the number of target data to be subjected to the predetermined process and the number of target data acquired by the acquisition step.
14. The parallel processing method according to claim 9 or 10, wherein in the execution control step, for each of the target data, for each of the second data to be subjected to the predetermined process, a thread is generated with the second data and the first data as operands and the process to be executed being the predetermined process, and each of the generated threads is caused to be executed by a GPU (Graphics Processing Unit).
15. A parallel processing apparatus comprising: an acquisition means for acquiring a plurality of target data including first data and a plurality of second data; an execution control means for, for each of the target data, executing in parallel a predetermined process targeting the first data and the second data for each of the second data equal to or less than a predetermined number; and a determination means for determining, for each of the target data, whether any of the second data targeted by the predetermined process satisfies a condition with the first data based on a result of the predetermined process, wherein the execution control means executes in parallel the predetermined process for each of the second data equal to or less than the predetermined number that has not yet been targeted by the predetermined process for each of the target data for which it is determined that none of the second data targeted by the predetermined process satisfies the condition with the first data.
16. The parallel processing apparatus according to claim 15, wherein the execution control means excludes from the targets of the predetermined process each of the target data for which it is determined that any one or more of the second data targeted by the predetermined process satisfies the condition with the first data, and each of the target data for which the predetermined process has already been executed for all of the plurality of the second data.
17. The parallel processing apparatus according to claim 15 or 16, wherein the predetermined process is a process of calculating a value representing a degree of similarity between the first data and the second data, and the condition is a condition that the second data is similar to the first data.
18. The parallel processing apparatus according to claim 15 or 16, wherein the execution control means increases the predetermined number as the number of the target data to be subjected to the predetermined process decreases.
19. The parallel processing apparatus according to claim 18, wherein the execution control means determines the predetermined number based on a ratio between the number of target data to be subjected to the predetermined process and the number of target data acquired by the acquisition means.
20. The parallel processing apparatus according to claim 15 or 16, wherein the execution control means generates, for each of the target data, a thread in which the second data to be subjected to the predetermined process and the first data are used as operands and the process to be executed is the predetermined process, and causes each of the generated threads to be executed by a GPU (Graphics Processing Unit).
Citation Information
Patent Citations
Pattern-matching method and device for a multiprocessor environment
WO2011078108A1