Optical proximity correction method, electronic device, and storage medium

By implementing memory sharing and adjacent block information referencing between computational cores in optical proximity correction, the problems of low computational efficiency and resource waste in the OPC scheme are solved, and a more efficient optical proximity correction effect is achieved.

CN120539999BActive Publication Date: 2025-11-07QUANXIN INTELLIGENT MFG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511036765.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-07
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing optical proximity correction (OPC) schemes suffer from low computational efficiency and resource waste in distributed computing, mainly due to the lack of linkage between computing cores, resulting in a large amount of redundant computation and small edge problems.

Method used

By sharing memory among multiple computing cores, each computing core can reference information from adjacent blocks for OPC correction, generate correction blocks, and consider the influence of adjacent blocks during splicing, thus reducing the correction work for additional regions.

Benefits of technology

It significantly improves the running efficiency of the OPC program, reduces small-block edge problems, enhances computational performance, and increases the proportion of effective calculated area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120539999B_ABST
    Figure CN120539999B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to an optical proximity correction method, an electronic device and a storage medium. The method comprises: each operation core in a plurality of operation cores respectively referencing information of adjacent blocks of a to-be-corrected block in a layout stored in a shared memory to determine an influence of the adjacent blocks on the to-be-corrected block; performing OPC correction on the to-be-corrected block based on information of the to-be-corrected block and the influence to generate a corrected block; and splicing each corrected block according to a position of a corresponding original layout corresponding to each block to form a corrected layout. The technical solution of the embodiments of the present disclosure can significantly improve the running efficiency of the OPC program and reduce the small block edge problem.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure generally relate to integrated circuits, and more particularly, to an optical proximity correction method, an electronic device, and a storage medium. BACKGROUND

[0002] Computational lithography is an important driving force for the development of pattern miniaturization technology since 1990s, aiming to break through the hardware limit of the minimum exposure size by improving the resolution of software technology under the condition that the existing hardware environment of the lithography machine and other equipment is unchanged, which greatly promotes the development of advanced semiconductor processes.

[0003] The production of integrated circuit chips at advanced process nodes usually relies on patterning technology. The core of the patterning technology is optical proximity correction (OPC). The layout processed by the computational lithography OPC is actually very large, so a distributed operation method must be used for processing. At advanced process nodes, thousands or even tens of thousands of computing cores are often needed for distributed operation. The traditional OPC scheme has the problem of further improving the operation efficiency. SUMMARY

[0004] According to an example embodiment of the present disclosure, an optical proximity correction scheme is provided to at least partially overcome the above or other potential deficiencies.

[0005] According to one aspect of the present disclosure, an optical proximity correction method is provided. The method includes: each of a plurality of computing cores referencing information of adjacent blocks of a to-be-corrected block in a layout stored in a shared memory, to determine an influence of the adjacent blocks on the to-be-corrected block; performing OPC correction on the to-be-corrected block based on the information of the to-be-corrected block and the influence, to generate a corrected block; and splicing each corrected block according to a position of the corresponding original layout of each block to form a corrected layout.

[0006] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes a processor; and a memory coupled with the processor, the memory having stored therein instructions that, when executed by the processor, cause the device to perform actions comprising: each of a plurality of computing cores referencing information of adjacent blocks of a to-be-corrected block in a layout stored in a shared memory, to determine an influence of the adjacent blocks on the to-be-corrected block; performing OPC correction on the to-be-corrected block based on the information of the to-be-corrected block and the influence, to generate a corrected block; and splicing each corrected block according to a position of the corresponding original layout of each block to form a corrected layout.

[0007] In some embodiments, the information of the adjacent blocks of the to-be-corrected block in the layout stored in the shared memory is respectively referenced by each of the plurality of operation cores includes: each of the operation cores respectively determines a distribution condition of the adjacent blocks at any boundary of the to-be-corrected block; and information of the adjacent blocks having an influence on any boundary of the to-be-corrected block is read based on the distribution condition.

[0008] In some embodiments, the information of the adjacent blocks having an influence on any boundary of the to-be-corrected block is read based on the distribution condition includes: in response to determining that any boundary of the to-be-corrected block has an adjacent block, information of a portion of the adjacent block close to the boundary is read.

[0009] In some embodiments, the optical proximity correction on the to-be-corrected block based on the information of the to-be-corrected block and the influence includes: performing optical proximity correction on the to-be-corrected block based on the information of the to-be-corrected block and the information of the portion close to the boundary to generate a corrected block without an additional area at the boundary with the adjacent block.

[0010] In some embodiments, the optical proximity correction on the to-be-corrected block based on the information of the to-be-corrected block and the influence further includes: in response to determining that any boundary of the to-be-corrected block does not have an adjacent block, performing OPC correction on the to-be-corrected block and an additional area at the boundary without the adjacent block to generate a corrected block with the additional area at the boundary without the adjacent block.

[0011] In some embodiments, the corrected block does not include a corrected additional area at the boundary with the adjacent block in addition to an area corresponding to the to-be-corrected block.

[0012] In some embodiments, the layout is subjected to optical proximity correction on a plurality of servers in a distributed architecture, each server having a plurality of operation cores, and the method further includes: the corrected layout generated by the plurality of operation cores on each server is spliced with the corrected layout generated by the operation cores located on other servers to form a final corrected layout.

[0013] In some embodiments, the plurality of operation cores are distributed on a server in one of the following architectures: X86 architecture; ARM architecture; GPU architecture.

[0014] In some embodiments, the method further includes: in each correction cycle, scheduling is performed based on operation conditions of the respective operation cores, so that an operation core that has completed operation assists an operation core that has not completed operation to perform operation.

[0015] In some embodiments, the scheduling based on the operation status of each operation core to make the operation core that finishes operation earlier assist the operation core that does not finish operation to perform operation comprises: determining the number of split pieces in each block; sorting the blocks based on the number; in response to a first operation core of a first sorted block that has the least number of split pieces finishing the loop, making the first operation core assist an operation core of a last sorted block that has the largest number of split pieces to perform operation of the loop; in response to a second operation core of a second sorted block that has the second least number of split pieces finishing the loop, making the second operation core assist an operation core of a second last sorted block that has the second largest number of split pieces to perform operation of the loop; and so on until operation of all blocks of each sorted in the correction loop is finished.

[0016] In some embodiments, the layout is divided into a plurality of blocks, and each block is sent to different operation cores for OPC correction in a predetermined order or randomly.

[0017] In some embodiments, the sending each block to different operation cores for OPC correction in a predetermined order or randomly comprises: allocating blocks arranged in a matrix of M by N to a first server with M operation cores, where M and N are integers greater than or equal to 1; allocating blocks arranged in a matrix of M by N to a second server with M operation cores; and so on until all blocks arranged in a matrix of M by N are allocated to a corresponding server with M operation cores. In some embodiments, the sending each block to different operation cores for OPC correction in a predetermined order or randomly comprises: allocating blocks arranged in a matrix of M by N to a first server with M operation cores, where M and N are integers greater than or equal to 1; allocating blocks arranged in a matrix of M by N to a second server with M operation cores; and so on until all blocks arranged in a matrix of M by N are allocated to a corresponding server with M operation cores. In some embodiments, the sending each block to different operation cores for OPC correction in a predetermined order or randomly comprises: allocating blocks arranged in a matrix of M by N to a first server with M operation cores, where M and N are integers greater than or equal to 1; allocating blocks arranged in a matrix of M by N to a second server with M operation cores; and so on until all blocks arranged in a matrix of M by N are allocated to a corresponding server with M operation cores. In some embodiments, the sending each block to different operation cores for OPC correction in a predetermined order or randomly comprises: allocating blocks arranged in a matrix of M by N to a first server with M operation cores, where M and N are integers greater than or equal to 1; allocating blocks arranged in a matrix of M by N to a second server with M operation cores; and so on until all blocks arranged in a matrix of M by N are allocated to a corresponding server with M operation cores.

[0018] In a third aspect of the present disclosure, a computer readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the method according to the first aspect of the present disclosure.

[0019] It will be understood from the following description that the technical solution of the present disclosure can significantly improve the operation efficiency of the OPC program and reduce the small block edge problem.

[0020] The summary section is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary section is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A schematic diagram of a split block in a layout is shown;

[0022] ​​​Figure 2 A flowchart illustrating a process of performing OPC correction of a layout in a conventional scheme is shown;

[0023] Figure 3 A flowchart illustrating a process of performing OPC correction of a layout in a conventional scheme is shown; Figure 2 A flowchart illustrating a process of performing OPC correction of a layout in a conventional scheme is shown;

[0024] Figure 4 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented is shown;

[0025] Figure 5 A flowchart illustrating an optical proximity correction method according to some embodiments of the present disclosure is shown;

[0026] Figure 6 A flowchart illustrating a process of performing OPC correction of a layout in a conventional scheme is shown;

[0027] Figure 7 A flowchart illustrating a process of performing OPC correction of a layout in a conventional scheme is shown; Figure 6 A flowchart illustrating a process of performing OPC correction of a layout in a conventional scheme is shown;

[0028] Figure 8 A schematic diagram illustrating a size of a split block in a layout is shown;

[0029] Figure 9 A block diagram of a computing device capable of implementing various embodiments of the present disclosure is shown.

[0030] In the various drawings, like or corresponding elements are denoted by like or corresponding reference numerals. DETAILED DESCRIPTION

[0031] The principles of the present disclosure will now be described, by way of example only, with reference to various example embodiments and with the aid of the accompanying drawings. It is to be understood that the description of these embodiments is merely intended to illustrate the general principles of the present disclosure. Therefore, the scope of the present disclosure should not be limited to these example embodiments. It should be noted that like or corresponding elements are denoted by like or corresponding reference numbers in the various drawings. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein can be employed without departing from the principles of the application described herein.

[0032] The term "includes" and its variations are meant to cover a non-exclusive inclusion, such that a process, method, system, product, or apparatus that comprises a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, system, product, or apparatus. The term "or" means "and / or" unless otherwise specified. The term "based on" means "based, at least in part, on" unless otherwise specified. The terms "one example embodiment" and "an embodiment" mean "at least one example embodiment." The term "another embodiment" means "at least one additional example embodiment." The terms "a first," "a second," etc. do not require that there be only one of the indicated objects.

[0033] The logic of the existing computing lithography correction distributed operation is to cut the layout into small blocks (templates, also referred to as blocks), and then send them to each independent operation core unit for correction and other operation processing. The edges of the cut small blocks need additional OPC operation. At the same time, when the small blocks are spliced to form the whole post-OPC layout at the output, there is a certain probability of template boundary problem.

[0034] In a normal OPC program, millions or even hundreds of millions of small blocks are cut, and then sent to different CPU operation cores in a certain order (also randomly) for operation.

[0035] Referring to Figure 1 , Figure 1 A schematic diagram of a block cut in a layout is shown. The block 102 is the original block divided in the layout, which is the area that needs to be corrected by OPC. The additional area 104 outside the block 102 represents the area that also needs to be corrected in the OPC correction process. Because the additional area is equivalent to the peripheral environment of the area to be corrected (also referred to as the corrected area), it will have an impact on the area to be corrected. In other words, this area represents the peripheral environment of the block 102, which will have an impact on the central area. If the additional area 104 is not corrected, it will likely result in the final failure of the cooperation of the two adjacent blocks after correction. That is, appropriate correction of the additional area 104 is necessary to ensure that the correction of the block 102 is carried out in a correct environment.

[0036] The reference area 106 outside the additional area 104 is not corrected in the OPC correction process. It can be understood that this reference block is only cut out to contain the part mentioned above that needs to be corrected by OPC. In order to ensure accurate correction of the block 102 without losing peripheral environment information, it is necessary to correct the layout within a certain range around it. Figure 1The reference region 106 is sent to the operation core for correction. That is, all the blocks within the boundary of the reference region 106 are sent to the operation core for OPC correction. In short, in order to process a small region, a large block containing the small block is directly cut out and sent to the operation core for processing. Therefore, what is sent to the operation core is a block larger than the block 102. The reference block is loaded as a reference level to the operation core, and only the corrected block 102 part is taken for splicing at the final splicing.

[0037] Since the result of OPC correction is not a unique solution, the peripheral environment will have a certain influence on the inside, and the final correction result is determined by multiple (more than 20 advanced nodes in general) correction cycles, there is a probability of converging to different solutions. Therefore, when two corrected blocks (or small blocks) are spliced, there is a certain probability of small block boundary problems. This probability is basically proportional to the sum of the boundaries of each block 102.

[0038] Generally speaking, the current OPC operation will be carried out in a distributed operation environment, and in the current OPC operation, each individual operation core and the memory allocated to it will be treated as a separate machine, and there is basically no linkage between the calculation cores. Thus resulting in a large amount of redundant calculation, making the operation efficiency low and wasting resources.

[0039] The following will be described in conjunction with Figure 2 . Figure 2 The flowchart of the OPC correction of the layout in the conventional scheme is shown. As Figure 2 shown, the blocks before correction are shown, and a total of 64 blocks are shown. It should be understood that the number of blocks shown here is illustrative, and the actual number can be changed according to actual needs.

[0040] Suppose that a server participating in distributed operation has a total of 64 operation cores (2 CPUs, each CPU has 32 operation cores). There is an operation core corresponding to each block, as Figure 2 shown, operation core 00, operation core 01, operation core 02, and so on to operation core 63. Each operation core has a corresponding allocated memory, and the corresponding block can be stored in the corresponding memory.

[0041] As mentioned above, each operation core is independent of each other in the whole OPC process. In the current OPC operation, each single operation core and the memory allocated to it are treated as a single machine, and there is basically no linkage between the calculation cores. That is, each operation core cannot access each memory. In the case of not sharing the memory, the OPC program is relatively simple (directly occupies the memory when running, prohibits other processes from using it, and there is no need to judge whether a certain process can access the memory occupied by itself).

[0042] As shown in Figure 2 , each corrected block includes an additional part around the four sides, i.e., the additional area 104 shown in Figure 1 . The final splicing does not include the additional area 104. After each cut block is completely corrected, the final splicing action is performed. Therefore, a large amount of additional work needs to be done in the whole OPC correction process, i.e., the correction work for the additional area 104.

[0043] Figure 3 The flow interaction condition when correcting two blocks in the scheme shown in Figure 2 is shown. As mentioned above, each operation core is independent of each other in the whole OPC process, i.e., each core works independently, and there is basically no linkage. This leads to a large amount of redundant operation, waste of resources, and reduction of efficiency.

[0044] In view of this, the present disclosure provides an improved scheme.

[0045] Embodiments of the present disclosure provide an improved optical proximity correction method. The method comprises: each operation core in a plurality of operation cores respectively referencing information of adjacent blocks of a to-be-corrected block in a layout stored in a shared memory to determine an influence of the adjacent blocks on the to-be-corrected block; performing OPC correction on the to-be-corrected block based on information of the to-be-corrected block and the influence to generate a corrected block; and splicing each corrected block according to a position of a corresponding original layout of each block to form a corrected layout. By referring to the information of the adjacent blocks, the operation of OPC correction on the additional area is reduced, and the operation efficiency is significantly improved.

[0046] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0047] Figure 4 A schematic diagram of an example environment 400 in which embodiments according to the present disclosure can be implemented is shown. As shown in Figure 4 , the example environment 400 includes a computing device 410 and a client 420.

[0048] In some embodiments, the computing device 410 can interact with the client 420. For example, the computing device 410 can receive an input message from the client 420 and output a feedback message to the client 420. In some embodiments, the input message from the client 420 can be design layout data. The computing device 410 can perform corresponding mathematical operations on the design layout data and output corresponding operation results to the client 420.

[0049] In some embodiments, the computing device 410 can include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device such as a mobile phone, a personal digital assistant PDA, a media player, etc., a consumer electronic product, a mini computer, a mainframe computer, a cloud computing resource, etc.

[0050] It should be appreciated that the structure and function of the example environment 400 are described for illustrative purposes only and are not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions. The environment is merely illustrative and is not intended to limit the applicability of the embodiments of the present disclosure.

[0051] To more clearly explain the principles of the present disclosure scheme, the following will be described in more detail with reference to Figure 5 .

[0052] Figure 5 A flowchart of a method 500 for optical proximity correction is shown, according to some embodiments of the present disclosure.

[0053] At block 502, each of the plurality of operation cores respectively references information of neighboring blocks of a block to be corrected in a layout stored in a shared memory to determine the influence of the neighboring blocks on the block to be corrected.

[0054] In some embodiments, the various operation cores can share a memory in which layout blocks are stored, so that when performing OPC correction on a block to be corrected in the layout, the information of the neighboring blocks stored in the memory can be referenced (or referred to) so that the block to be corrected can be fully corrected by fully considering the surrounding environment. The block to be corrected refers to the block corresponding to each operation core that needs to be corrected.

[0055] The layout is divided into a plurality of blocks for OPC correction, and the allocation of each block to the operation core or the memory can be performed in a conventional manner. For example, each block can be sent to different operation cores for OPC correction in a predetermined order or randomly.

[0056] For example, the blocks arranged in a matrix of may be allocated to the first operation core with The first server has 1 computing core, where M and N are integers greater than or equal to 1; the second server is arranged as follows: The block of the matrix is ​​assigned to the second one with The second server with each computing core; and so on, until all are arranged into The matrix blocks are assigned to the corresponding blocks. A server with one computing core. For example, the first one... The matrix blocks run on server A with 64 cores, in parallel with a second... The matrix blocks run on server B with 64 cores, ... In other words, the first... The matrix blocks were allocated to the first server with 64 processing cores. The second... The matrix blocks were allocated to a second server with 64 processing cores. Server A completed the first... After the correction of the blocks in the matrix, it will take over the Nth one. The matrix blocks are computed in parallel to obtain the final result.

[0057] Normally, the small blocks are sent to the CPU processing core one by one. Only after the previous small block has been processed on that processing core can the next small block be received for processing. Furthermore, When the small blocks are sent to a 64-core machine, they all immediately enter their respective processing cores.

[0058] It should be noted that regardless of whether an optimization algorithm is used, the final generated correction region (corrected block 102) directly enters the final post-OPC layout stitching process (stitching is the last step in the OPC operation). That is, all operations involving the additional region 104 and the reference region 106 are ultimately discarded. A key aspect of some embodiments of this invention is minimizing these ultimately discarded operations, thereby achieving a significant improvement in efficiency.

[0059] Larger semiconductor manufacturers typically perform single OPC operations on servers purchased in the same batch. These servers generally have the same configuration of CPU, memory, operating system, etc. (which reduces many problems). Therefore, when running OPC programs (which usually have thousands of computing cores), they face the same computing environment, making it very suitable to use the method described in this invention to speed up the process.

[0060] In some embodiments, each of the plurality of operation cores respectively referencing information of adjacent blocks of the to-be-corrected block in the layout stored in the shared memory can include: each of the operation cores respectively determining a distribution condition of the adjacent blocks at any boundary of the to-be-corrected block; and reading information of the adjacent blocks having an influence on any boundary of the to-be-corrected block based on the distribution condition.

[0061] In some embodiments, reading information of the adjacent blocks having an influence on any boundary of the to-be-corrected block based on the distribution condition can include: in response to determining that any boundary of the to-be-corrected block has an adjacent block, reading information of a portion of the adjacent block close to the boundary. In the case that any boundary has no adjacent block, correction can be performed in a traditional manner.

[0062] At block 504, optical proximity correction is performed on the to-be-corrected block based on the information of the to-be-corrected block to generate a corrected block.

[0063] In some embodiments, performing optical proximity correction on the to-be-corrected block based on the information of the to-be-corrected block can include: performing optical proximity correction on the to-be-corrected block based on the information of the to-be-corrected block and the information of the portion close to the boundary to generate a corrected block without an additional area at the boundary having an adjacent block.

[0064] In some embodiments of the present application, each block is stored in the memory, and each operation core can access all the memory in addition to the corresponding memory, especially the memory on the same server. In this way, the information of the adjacent block can be known, so that the information of the adjacent block can be referred to or referenced in the OPC process. In some embodiments of the present application, the access of the operation core to all the memory can be realized by making some corresponding settings in the program.

[0065] In the traditional scheme, the correction process of the operation core is independent, and each operation core processes the to-be-corrected block through the provided algorithm. In order to speed up the process of data processing, the memory independent mode is selected, so that the prior art does not share the memory, and the OPC program is relatively simple. Industry insiders have not realized that the amount of redundant calculation can be reduced by sharing the memory.

[0066] The operation core can refer to (e.g., real-time refer to or refer to at a predetermined period) the data in the memory during the OPC correction process. In some embodiments of the present disclosure, real-time referring to the data in the memory can be real-time reading of the data in the memory. Basically, the operation core is a part of the CPU, while the memory is plugged on the motherboard and separated from the CPU (but still connected through the bus of the motherboard). During the operation process, the operation core (Core) on the CPU will refer to the data in the memory in real time and also change the data in the memory. In addition, during the OPC correction process of the operation core, the information of the corresponding block in the memory will be updated in real time, so that the operation core corresponding to the adjacent block can access the information of the adjacent block in real time. Therefore, the operation core can perform OPC correction on the to-be-corrected block based on the information of the corresponding block and the information of the adjacent block. In short, the operation core can refer to or understand the situation of the adjacent block during the correction process, so that the influence of the adjacent block is considered during the correction process. In this way, during the correction process, there is no need to correct the surrounding area of the to-be-corrected block, because the surrounding area is actually the adjacent to-be-corrected area, which will be precisely corrected during the correction process. In the traditional scheme, the surrounding area will be corrected with low precision. Now that the situation of the adjacent area has been understood, the to-be-corrected block can be corrected by referring to the information of the adjacent area, and there is no need to correct the surrounding area with low precision again.

[0067] In addition, it should be noted that the block information corrected by the CPU operation core is stored in the memory, so the block information refers to the information stored in the memory. The cache in the CPU is very small and cannot store the block information.

[0068] In some embodiments, performing the OPC correction on the to-be-corrected block based on the information of the to-be-corrected block and the influence further includes: in response to determining that any boundary of the to-be-corrected block does not have an adjacent block, performing the OPC correction on the to-be-corrected block and an additional area at the boundary without the adjacent block to generate a corrected block with the additional area at the boundary without the adjacent block.

[0069] In some embodiments, performing the OPC correction on the to-be-corrected block based on the information of the to-be-corrected block and the influence can include: in a case where any boundary of the to-be-corrected block has an adjacent block, performing the OPC correction on the to-be-corrected block without performing the OPC correction on an additional area at the boundary. In this case, the corrected block does not include a corrected additional area at the boundary with the adjacent block in addition to an area corresponding to the to-be-corrected block. In a case where any boundary of the to-be-corrected block does not have an adjacent block, performing the OPC correction on the to-be-corrected block and an additional area at the boundary without the adjacent block to generate a corrected block with the additional area. For example, in some embodiments, the adjacent block is a block adjacent to the to-be-corrected block in a direction perpendicular to the direction of the to-be-corrected block. (or some other combination) are sent to CPU cores installed on the same server (which can efficiently communicate through the memory on the same server) to perform operations according to an optimized algorithm, which is exactly the way described above "in a certain order".

[0070] In fact, the distributed computing environment commonly used is built by multiple servers. Taking the x86 architecture as an example, the memory on each server is shared by the computing cores on the server. Each server usually has 2 CPUs, and each CPU has dozens of computing cores. Assuming that there are 64 computing cores on the server participating in the distributed computing (2 CPUs, each with 32 computing cores), there are 1024 GB of shared memory (16 GB on average for each computing core).

[0071] The foregoing embodiments are described by way of example in which each computing core can access all the memory. As mentioned above, real-time access to the memory can be real-time reading. It should be understood that the present disclosure is not limited thereto, i.e., it is not limited to real-time reading of the memory, but as long as each computing core can learn the information of the adjacent block in real time, for example, each core can pass the information of the block in the relevant memory to each other in a predetermined manner, and each core can also learn the information of the adjacent block in time.

[0072] The modification of the additional area can be understood as affecting the modification of the basic area. Practice shows that the desired lithography result can be obtained by such modification. In addition, in the present application, at the boundary of the adjacent block, the information of the adjacent block is referred to (or said to be referenced) during the modification of each basic area, so that the modification of each basic area has actually taken into account the influence of the adjacent area (the additional area corresponding to the conventional scheme), so that the modification of the peripheral area (the additional area mentioned above) of the basic area is not required. The modification result obtained in this way can produce a desired lithography result.

[0073] At block 506, the correction blocks are spliced according to the positions of the original layout corresponding to each block to form a correction layout.

[0074] In some embodiments, the layout is optically proximity corrected on multiple servers in a distributed architecture, each server having multiple computing cores, and the method further comprises: the correction layout generated by the multiple computing cores on each server is spliced with the correction layout generated by the computing cores on other servers to form a final correction layout.

[0075] In some embodiments, the plurality of operation cores are distributed on a server of one of the following architectures: X86 architecture; Advanced RISC Machine (ARM) architecture; Graphics Processing Unit (GPU) architecture.

[0076] In some embodiments, in each correction cycle, the operation cores are scheduled based on their operation status so that the operation cores that have finished operation assist the operation cores that have not finished operation to perform operation.

[0077] The following refers to Figure 6 schemes of some embodiments of the present disclosure. Figure 6 A flowchart of OPC correction of a layout according to some embodiments of the present disclosure is shown. It is assumed that there are 64 operation cores (2 CPUs, each with 32 operation cores) on a server participating in distributed operation, and there are 1024 GB of shared memory (16 GB on average for each operation core). Figure 6 In the figure, 64 zones running on the same server are shown. Specifically, the zones in the figure are The zones in the figure are

[0078] As shown in Figure 6 , the same as in Figure 2 , 64 operation cores are shown, i.e., operation core 00, operation core 01, operation core 02, and so on, up to operation core 63. As mentioned before, in the traditional scheme, the operation cores are independent of each other during the entire OPC process, and each individual operation core and the memory allocated to it are treated as a separate machine, and there is basically no linkage between the operation cores.

[0079] In some embodiments of the present disclosure, unlike the scheme in Figure 2 , each operation core can access all the memory on the server. During the correction process, the operation core can refer to or be aware of the situation of the adjacent zones. Therefore, the influence of the adjacent zones is considered during the correction process. In this way, for the to-be-corrected zone with adjacent zones, there is no need to correct the surrounding area of the to-be-corrected zone, because the surrounding area is actually the adjacent to-be-corrected area, which will be precisely corrected during its corresponding correction process, and in the traditional scheme, the surrounding area will be corrected with lower precision. Now that the situation of the adjacent area is known, there is no need to correct the surrounding area with lower precision. As shown in Figure 6As shown, the modified individual patches only have the area corresponding to the additional area 104 at the edge locations where there is no adjacent patch, otherwise there is no area corresponding to the additional area 104, i.e. no OPC correction for the additional area 104 is necessary at the edge locations where there is an adjacent patch. In this way, a lot of extra work is avoided in the entire OPC correction process.

[0080] From Figure 6 It can be seen from FIG. 6, where the first circle 602, the second circle 604, the third circle 606, the fourth circle 608, and the fifth circle 610 are shown, each representing a portion of the patch. According to the scheme of some embodiments of the present disclosure, for the corresponding small patch (the patch to be corrected), the pattern in the third circle 606 and the fourth circle 608 does not need to be processed, because these two circles are in the additional area at the edge locations of the boundary with the adjacent patch mentioned above. Therefore, this part of information does not need to be output when outputting. This is the aforementioned "generating the corrected patch without the additional area at the boundary with the adjacent patch". In the actual operation, the corresponding information is actually replaced by the information in the second circle 604 and the fifth circle 610. Specifically, the information in the third circle 606 is replaced by the information in the fifth circle 610; the information in the fourth circle 608 is replaced by the information in the second circle 604. In short, the information in the circle pointed by the arrow is replaced by the information in the circle pointed by the arrow tail.

[0081] For the small patch at the "corner" (for example, the patch corresponding to the operation core 00), when outputting the corrected information, the information in the first circle 602 needs to be output (without the information of the small patch running on the same server), and the information in the third circle 606 does not need to be output (filled with the information of the small patch running on the same server).

[0082] As can be seen, by the scheme of the above-mentioned embodiments of the present disclosure, a lot of extra work can be saved at the boundary with the adjacent patch.

[0083] After each core completes the correction of the corresponding patch, the corrected patches are spliced to form the correction of the layout patch completed by all the cores on the server. This is referred to as preliminary splicing in Figure 6 After preliminary splicing, the corrected layout patch can be spliced with the corrected layout patch on other servers to form the final corrected layout.

[0084] Figure 7 It is shown in Figure 6The flow interaction condition when two blocks are corrected in the shown scheme is shown. As mentioned above, the memory is shared for each operation core in the whole OPC process, i.e. each operation core can access any memory on the server. Therefore, for example, the memory storing the blocks adjacent to the processed block can be referenced to obtain the information of the adjacent blocks, so that the surrounding environment information can be fully considered in the OPC correction process.

[0085] By uniformly accessing the layout information stored in the memory on the server, a large amount of additional OPC operations can be reduced (note that Figure 6 As shown in the figure, the output additional correction area is reduced, and only the edge part of the block has the additional area, thereby significantly reducing the operation amount. In addition, the correction patterns are integrally formed (the cores refer to each other in the OPC process), and the small edge problem, such as the alignment problem, etc., does not occur in the connected area.

[0086] In some embodiments of the present disclosure, during the running process, the basic correction area is replaced with the basic correction area and the additional correction area in the standard OPC flow by mutual reference, thereby significantly reducing the operation amount.

[0087] In some embodiments, the OPC flow can be further optimized. For example, in each correction cycle, the operation cores can be scheduled based on the operation status of each operation core, so that the operation core that completes the operation first assists the operation core that does not complete the operation to perform the operation.

[0088] In some embodiments, the operation cores are scheduled based on the operation status of each operation core to make the operation core that completes the operation first assist the operation core that does not complete the operation to perform the operation, which can include: first determining the number of segments in each block; sorting the blocks based on the determined number; in the case that the first operation core operating the first sorted block with the least number of segments completes the cycle, the first operation core can be made to assist the operation core operating the last sorted block with the largest number of segments to perform the operation of the cycle; in the case that the second operation core operating the second sorted block with the second least number of segments completes the cycle, the second operation core can be made to assist the operation core operating the second last sorted block with the second largest number of segments to perform the operation of the cycle; and so on, until the operation of all blocks of each sorted in the correction cycle is completed.

[0089] After each round of operation is completed, in general, the time when the program ends on each CPU is about the same. Because the mutual reference of the memory is involved here, the graphics are synchronized once in a period, so in this case, it is expected that the operation cores of each CPU are kept balanced, so as to further improve the efficiency.

[0090] Generally, the amount of correction needed for each of the small pieces is about the same, and can be considered as the same. However, in some extreme cases, there can be a large difference in the amount of correction needed for each of the small pieces. In this case, the efficiency can be lost when the optimization process is used.

[0091] For example, for 64 small pieces, the operation core for processing one of the small pieces can operate particularly fast. After it finishes operating, the CPU waits for the next cycle, and the efficiency is lost.

[0092] In some embodiments, when this situation occurs, the following method can be used: before the 64 small pieces start operating, the total number of the split pieces in each small piece is counted. The split piece refers to the segment of the edge of the pattern that is split. As known in the art, when performing OPC correction on a pattern, the edge of the pattern is usually segmented. There are sampling points on each segment. The total number of the split pieces in each small piece is basically proportional to the operating speed.

[0093] The small pieces are sorted according to the total number of the split pieces. For example, assume the following sequence:

[0094] ;

[0095] In this case, S1 represents the block with the least number of split pieces, S2 represents the block with the second least number of split pieces, S3 represents the block with the third least number of split pieces, and S64 represents the block with the most number of split pieces. It should be understood that the sorting method shown here is only illustrative, and other sorting methods can be used as needed, such as sorting in descending order. 64 It should be understood that the sorting method shown here is only illustrative, and other sorting methods can be used as needed, such as sorting in descending order. 64 It should be understood that the sorting method shown here is only illustrative, and other sorting methods can be used as needed, such as sorting in descending order.

[0096] After the above sorting is performed, the following can be set in the program:

[0097] When the operation core for processing S1 finishes operating for one cycle, the following operation (starting from the split piece in reverse order) in S64 that is not completed is performed; 64

[0098] When the operation core for processing S2 finishes operating for one cycle, the following operation (starting from the split piece in reverse order) in S63 that is not completed is performed; 63

[0099] …, and so on.

[0100] In this way, the efficiency of the OPC correction is further accelerated.

[0101] ​​The scheme of some embodiments of the present disclosure improves the operation efficiency by sharing memory, reference or citation, which can achieve more optimized efficiency on a single server, because the access speed to the memory is very fast. For multiple servers, access can be performed through a network cable or the like, depending on the network speed, which is generally relatively low compared with the above-mentioned embodiments executed on a single server.

[0102] Figure 8 A size diagram of a block cut in a layout is shown. The typical size of the small block cut by OPC is about Figure 8 a is about 20um (microns), and b is about 2um. Taking this as a typical example, the OPC layout area required for calculation of the above-mentioned 64-core machine in the above-mentioned running mode without optimization is (the actual output area is , so the effective calculation area is less than 70%); and the OPC layout area required for calculation in the optimized case is (the actual output area is , so the effective calculation area is more than 95%). The sum of the boundaries involved is 5120um in the unoptimized case and 640um in the optimized case. It can be seen that by reducing the operation of the additional area, the area of the layout block that needs to be corrected by OPC is significantly reduced, thereby greatly improving the operation efficiency. In addition, the probability of small block edge problems is significantly reduced, for example, it is 1 / 8 of the unoptimized case.

[0103] At present, in addition to running on the x86 architecture, the OPC program has begun to be popularized to other architectures, such as the ARM architecture and the GPU architecture. In these new architectures, the number of operation cores will be greatly improved compared with the x86 architecture (for example, a type of ARM architecture server has 384 operation cores in a single machine), but the operation capability of a single core will be significantly lower than that of the x86 architecture. In this case, the total area of the small block cut will be reduced, which is specifically shown in Figure 8 that the size of a needs to be reduced (b remains unchanged). If the standard OPC operation process is adopted, the proportion of the effective calculation area will be further reduced. Therefore, the scheme of the embodiments of the present disclosure will play a greater role on the new operation architecture.

[0104] In some embodiments of the present disclosure, the operation data in different operation cores on the same CPU are linked, and multiple operation cores uniformly access the layout information in the memory of the same machine (or server) to dynamically optimize the allocation of CPU operation resources.

[0105] Some embodiments of the present disclosure provide optical proximity correction methods. It should be noted that the examples in the above embodiments are only for illustrating the schemes of some embodiments of the present disclosure, and are not used to limit the schemes of the present disclosure.

[0106] The schemes of some embodiments of the present disclosure can significantly improve the running efficiency of the OPC procedure (effective calculation area is increased), and further reduce the problem of small block edges.

[0107] It should be understood that the embodiments shown in the drawings are only for illustratively showing the schemes of some embodiments of the present disclosure, and are not used to limit the present disclosure. The embodiments of the present disclosure can also have various other forms.

[0108] An electronic device is also disclosed in the embodiments of the present disclosure. The electronic device includes a processor, and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the device to perform acts comprising: each of a plurality of operation cores referencing information of adjacent blocks of a to-be-corrected block in a layout stored in a shared memory, to determine an influence of the adjacent blocks on the to-be-corrected block; performing OPC correction on the to-be-corrected block based on information of the to-be-corrected block and the influence, to generate a corrected block; and splicing each corrected block according to a position of a corresponding original layout of each block to form a corrected layout.

[0109] A computer readable storage medium is also disclosed in the embodiments of the present disclosure, and the computer program is stored on the computer readable storage medium, and the program is executed by a processor to implement the method according to the embodiments of the present disclosure.

[0110] Figure 9 A schematic block diagram of an electronic device in accordance with some example embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown in the electronic device, its connections and relationships, and its functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0111] As Figure 9As shown, the device 900 includes a CPU 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The CPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0112] A plurality of components in the device 900, including an input unit 906, such as a keyboard, a mouse, and the like, an output unit 907, such as various types of displays, speakers, and the like, a storage unit 908, such as a magnetic disk, a magneto-optical disk, and the like, and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, and the like, are connected to the I / O interface 905. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0113] The various processes and processes described above, such as the method 500, can be performed by the CPU 901. For example, in some embodiments, the method 500 can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the CPU 901, one or more steps of the method 500 described above can be performed.

[0114] The schemes according to the embodiments of the present disclosure can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for executing various aspects of the present disclosure. The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable program instructions can be downloaded to various computing / processing devices from a computer readable storage medium or to an external computer or an external storage device via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network.

[0115] Embodiments of the present disclosure have been described above, the above description is exemplary only, is only optional embodiments of the present disclosure, is not exhaustive, and is not used to limit the present disclosure. Although the claims in this application have been drafted in relation to a specific combination of features, it should be understood that the scope of the present disclosure also includes any novel feature or any novel combination of features disclosed herein, explicitly or implicitly, or any generalization thereof, regardless of whether it is involved in the same solution as any of the presently claimed claims. It should be understood that new claims can be drafted to these features and / or combinations of these features during the examination of this application or any further application derived therefrom.

[0116] The selection of the terms used herein is intended to best explain the principles of the embodiments, practical application, or technical improvement in the art, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The present disclosure can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method of optical proximity correction, comprising: each of a plurality of operation cores reading information of neighboring blocks of a block to be corrected in a layout stored in a shared memory in real time, respectively, to determine an impact of the neighboring blocks on the block to be corrected; performing optical proximity correction on the block to be corrected based on information of the block to be corrected and the impact to generate a corrected block; and stitching each of the corrected blocks according to a position of each of the corrected blocks in an original layout to form a corrected layout; wherein in each correction cycle, each of the operation cores is dispatched based on an operation status of each of the operation cores to allow an operation core that finishes operation earlier to assist an operation core that does not finish operation.

2. The method of claim 1, wherein each of the plurality of operation cores reading information of neighboring blocks of a block to be corrected in a layout stored in a shared memory in real time, respectively, comprises: each of the operation cores determining a distribution status of the neighboring blocks at any boundary of the block to be corrected, respectively; and each of the operation cores reading information of the neighboring blocks having an impact on any boundary of the block to be corrected based on the distribution status.

3. The method of claim 2, wherein each of the operation cores reading information of the neighboring blocks having an impact on any boundary of the block to be corrected based on the distribution status comprises: each of the operation cores reading information of a portion of the neighboring blocks close to the boundary in response to determining that any boundary of the block to be corrected has the neighboring blocks.

4. The method of claim 3, wherein each of the operation cores performing optical proximity correction on the block to be corrected based on information of the block to be corrected and the impact comprises: each of the operation cores performing optical proximity correction on the block to be corrected based on information of the block to be corrected and the portion of the neighboring blocks close to the boundary to generate a corrected block without additional areas at the boundary with the neighboring blocks.

5. The method of claim 3, wherein each of the operation cores performing optical proximity correction on the block to be corrected based on information of the block to be corrected and the impact further comprises: each of the operation cores performing optical proximity correction on the block to be corrected and additional areas at the boundary without the neighboring blocks in response to determining that any boundary of the block to be corrected does not have the neighboring blocks to generate a corrected block with the additional areas at the boundary without the neighboring blocks.

6. The method of claim 1, wherein the corrected block does not include corrected additional areas at the boundary with the neighboring blocks other than areas corresponding to the block to be corrected.

7. The method of claim 1, wherein the layout is subjected to the optical proximity correction on a plurality of servers in a distributed architecture, each of the servers having a plurality of operation cores, the method further comprising: the corrected layout generated by the plurality of operation cores on each of the servers being stitched with corrected layouts generated by operation cores on other servers to form a final corrected layout.

8. The method of claim 7, wherein the plurality of operation cores are distributed on servers in one of the following architectures: X86 architecture; graphics processor architecture; and advanced reduced instruction set computer machine architecture. ​ ​ 9. The method of claim 1, wherein the scheduling based on the operation status of each operation core to cause the operation core that finishes operation to assist the operation core that does not finish operation to perform operation comprises: determining the number of split pieces in each block; ranking the blocks based on the number; in response to a first operation core of a first ranked block that has the least number of split pieces finishing the loop, causing the first operation core to assist an operation core of a last ranked block that has the most number of split pieces to perform operation of the loop; in response to a second operation core of a second ranked block that has the second least number of split pieces finishing the loop, causing the second operation core to assist an operation core of a second last ranked block that has the second most number of split pieces to perform operation of the loop; and and so on until all blocks of each rank in the correction loop are finished.

10. The method of claim 1, wherein the layout is divided into a plurality of blocks, each block is sent to a different operation core to perform optical proximity correction in a predetermined order or randomly.

11. The method of claim 10, wherein the sending each block to a different operation core to perform optical proximity correction in a predetermined order or randomly comprises: allocating a first block arranged in a M*N matrix to a first server with M*N operation cores, where M and N are integers greater than or equal to 1; allocating a second block arranged in a M*N matrix to a second server with M*N operation cores; and so on until all blocks arranged in a M*N matrix are allocated to a corresponding server with M*N operation cores.

12. An electronic device, comprising: a processor; and a memory coupled with the processor, the memory having stored therein instructions that, when executed by the processor, cause the device to perform acts comprising: each operation core of a plurality of operation cores respectively reading information of adjacent blocks of a block to be corrected in a layout stored in a shared memory in real time to determine an effect of the adjacent blocks on the block to be corrected; performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and the effect to generate a corrected block; and stitching each of the corrected blocks according to a position of a corresponding original layout of each of the corrected blocks to form a corrected layout; wherein, in each correction loop, the scheduling is based on an operation status of each operation core to cause the operation core that finishes operation to assist the operation core that does not finish operation to perform operation.

13. A computer readable storage medium having stored thereon machine executable instructions that, when executed by a processor, cause the processor to implement the method of any one of claims 1 to 11. ​ ​ ​

Citation Information

Patent Citations

  • Integrated circuit optical proximity correction parallel processing method and system

    CN113777877A

  • Target layout graph forming method

    CN115600541A