Optical proximity correction method, electronic equipment and storage medium

By realizing memory sharing between computing cores and reference of adjacent block information between computing cores in optical proximity correction, the problems of low computing efficiency and small block edges are solved, and the operation efficiency and calculation area utilization of OPC programs are improved.

CN120539999AActive Publication Date: 2025-08-26QUANXIN INTELLIGENT MFG TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511036765.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-26
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

The existing optical proximity correction (OPC) schemes have problems with low computing efficiency and waste of resources in distributed computing, which is mainly due to the lack of linkage between the computing cores, resulting in redundant computing and small-block edge problems.

Method used

By sharing memory among multiple computing cores, each computing core allows reference to the information of adjacent blocks for OPC correction, generates correction blocks, and considers the impact of adjacent blocks during splicing, reducing correction operations on additional areas.

Benefits of technology

It significantly improves the operation efficiency of OPC programs, reduces the edge problems of small blocks, improves the computing efficiency, improves the effective calculation area, and reduces the computing volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120539999A_ABST
    Figure CN120539999A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an optical proximity correction method, electronic equipment and a storage medium. The method comprises the following steps: each operation core in a plurality of operation cores respectively references information of adjacent blocks of a to-be-corrected block in a layout stored in a shared memory so as to determine the influence of the adjacent blocks on the to-be-corrected block; performing OPC correction on the to-be-corrected block based on the information and the influence of the to-be-corrected block to generate a corrected block; and splicing each correction block according to the position of the original layout corresponding to each corresponding block to form a correction layout. According to the technical scheme, the operation efficiency of the OPC program can be remarkably improved, and the small block edge problem is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to integrated circuits, and more particularly, to optical proximity correction methods, electronic devices, and storage media. Background Art

[0002] Computational lithography technology has been an important driving force for the continued development of graphic miniaturization technology since the 1990s. It aims to break through the hardware limitations of the minimum exposure size by improving software technologies such as resolution while keeping the hardware environment of existing lithography equipment unchanged, greatly promoting the development of advanced semiconductor processes.

[0003] The production of integrated circuit chips at advanced process nodes typically relies on patterning technology. The core of patterning technology is optical proximity correction (OPC). The layouts processed by computational lithography OPC are extremely large, necessitating a distributed computing approach. At advanced process nodes, thousands or even tens of thousands of computing cores are often required for distributed computing. Traditional OPC solutions face challenges in terms of computational efficiency. Summary of the Invention

[0004] According to example embodiments of the present disclosure, an optical proximity correction scheme is provided to at least partially overcome the above or other potential drawbacks.

[0005] According to one aspect of the present disclosure, an optical proximity correction (OPC) method is provided. The method includes: each of a plurality of computing cores references information about adjacent blocks of a block to be corrected in a layout stored in a shared memory to determine the impact of the adjacent blocks on the block to be corrected; performing OPC correction on the block to be corrected based on the information and the impact of the block to be corrected to generate a correction block; and splicing the correction blocks according to the positions of the original layout corresponding to each block to form a correction layout.

[0006] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes a processor; and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the device to perform actions, the actions comprising: each of a plurality of computing cores referencing information of adjacent blocks of a block to be corrected in a layout stored in a shared memory to determine the impact of the adjacent blocks on the block to be corrected; performing OPC correction on the block to be corrected based on the information and the impact of the block to be corrected to generate a correction block; and splicing the correction blocks according to the positions of the original layout corresponding to the respective blocks to form a correction layout.

[0007] In some embodiments, each of the multiple computing cores respectively references information about neighboring blocks of the block to be corrected in a layout stored in a shared memory, including: each computing core respectively determines the distribution status of neighboring blocks at any boundary of the block to be corrected; and reads information about neighboring blocks that have an impact on any boundary of the block to be corrected based on the distribution status.

[0008] In some embodiments, reading information of adjacent blocks that affect any boundary of the block to be corrected based on the distribution condition includes: in response to determining that any boundary of the block to be corrected has an adjacent block, reading information of a portion of the adjacent block close to the boundary.

[0009] In some embodiments, performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and the influence includes: performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and information of a portion near a boundary, so as to generate a correction block without an additional area at a boundary with an adjacent block.

[0010] In some embodiments, performing optical proximity correction on the block to be corrected based on the information and the impact of the block to be corrected further includes: in response to determining that any boundary of the block to be corrected has no adjacent blocks, performing OPC correction on the block to be corrected and the additional area at the boundary without the adjacent blocks to generate a correction block with the additional area at the boundary without the adjacent blocks.

[0011] In some embodiments, the correction block does not include an additional corrected area at a boundary with an adjacent block except for an area corresponding to the block to be corrected.

[0012] In some embodiments, the layout is optically proximity corrected on multiple servers in a distributed architecture, each server having multiple computing cores, and the method further includes: splicing the corrected layout generated by the multiple computing cores on each server with the corrected layout generated by the computing cores located on other servers to form a final corrected layout.

[0013] In some embodiments, the multiple computing cores are distributed on a server of one of the following architectures: X86 architecture; ARM architecture; GPU architecture.

[0014] In some embodiments, the method further includes: in each calibration cycle, scheduling based on the computing status of each computing core, so that the computing core that completes the computing first assists the computing core that has not completed the computing to perform the computing.

[0015] In some embodiments, scheduling is performed based on the computing status of each computing core so that the computing core that completes the operation first assists the computing core that has not completed the operation to perform the operation, including: determining the number of sliced ​​fragments in each block; sorting each block based on the number; in response to the first computing core that operates the first sorted block with the least number of sliced ​​fragments completing the cycle, the first computing core assists the computing core that operates the last sorted block with the largest number of sliced ​​fragments to perform the operation in the cycle; in response to the second computing core that operates the second sorted block with the second least number of sliced ​​fragments completing the cycle, the second computing core assists the computing core that operates the second last sorted block with the second largest number of sliced ​​fragments to perform the operation in the cycle; and so on, until the operations of all blocks of each sort in the correction cycle are completed.

[0016] In some embodiments, the layout is divided into multiple blocks, and each block is sent to a different computing core for OPC correction in a predetermined order or randomly.

[0017] In some embodiments, each block is sent to different computing cores in a predetermined order or randomly for OPC correction, including: The block of the matrix is ​​assigned to the first one with The first server has 100 computing cores, where M and N are integers greater than or equal to 1; the second server is arranged into The blocks of the matrix are assigned to the second one with The second server with 100 computing cores; and so on, until all the servers are arranged into The blocks of the matrix are assigned to the corresponding A server with 100 computing cores.

[0018] In a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0019] It will be understood from the following description that the technical solution of the present disclosure can significantly improve the operating efficiency of the OPC program and reduce the small block edge problem.

[0020] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A schematic diagram of a block cut out from the layout is shown; Figure 2A schematic diagram of the process of performing OPC correction of a layout in a traditional solution is shown; Figure 3 Shown Figure 2 The process interaction when correcting two blocks in the solution shown; Figure 4 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented; Figure 5 A flowchart of an optical proximity correction method according to some embodiments of the present disclosure is shown; Figure 6 A schematic diagram of a process for performing OPC correction of a layout according to a method of some embodiments of the present disclosure is shown; Figure 7 Shown Figure 6 The process interaction when correcting two blocks in the scheme shown; Figure 8 A schematic diagram showing the size of a block cut out from the layout; Figure 9 A block diagram is shown of a computing device capable of implementing various embodiments of the present disclosure.

[0022] In the various drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION

[0023] The principles of the present disclosure will be described below with reference to the various exemplary embodiments shown in the accompanying drawings. It should be understood that the description of these embodiments is only to enable those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that similar or identical reference numerals can be used in the figures where possible, and similar or identical reference numerals can represent similar or identical functions. Those skilled in the art will readily recognize, from the description below, that alternative embodiments of the structures and methods described herein can be adopted without departing from the principles of the present invention described herein.

[0024] As used herein, the term "including" and its variations represent open inclusion, i.e., "including but not limited to." Unless otherwise stated, the term "or" means "and / or." The term "based on" means "based at least in part on." The terms "one example embodiment" and "an embodiment" mean "at least one example embodiment." The term "another embodiment" means "at least one additional embodiment." The terms "first," "second," etc. may refer to different or identical objects.

[0025] The existing distributed computational lithography correction logic divides the layout into small blocks (templates, also called blocks) and sends them to independent computing cores for correction and other processing. The edges of these blocks require additional OPC operations. Furthermore, when these blocks are spliced ​​together to form the overall post-OPC layout during output, there is a certain chance of template boundary problems.

[0026] During normal operation, a normal OPC program will be divided into millions or even hundreds of millions of small blocks, which are then sent to different CPU computing cores in a certain order (or randomly) for calculation.

[0027] See also Figure 1 , Figure 1 The diagram shows a block segmented from a layout. Block 102 is the original segmented block in the layout and is the area requiring OPC correction. The additional area 104 outside block 102 represents an area that also requires corresponding correction during the OPC correction process. Because the additional area is equivalent to the surrounding environment of the area to be corrected (also called the correction area), it will have an impact on the area to be corrected. In other words, this area represents the surrounding environment of block 102 and will have an impact on the central area. If additional area 104 is not corrected, it is likely that the two adjacent blocks will ultimately become incompatible after correction. In other words, proper correction of additional area 104 can ensure that the correction process for block 102 is performed in a correct environment.

[0028] The reference area 106 outside the additional area 104 will not be corrected during the OPC correction process. It can be understood that the reference block is only cut out to include the part that needs to be corrected. In order to ensure that the block 102 is accurately corrected without losing the surrounding environment information, it is necessary to Figure 1 The OPC correction is then sent to the computing core. Specifically, all blocks within the boundaries of reference area 106 are sent to the computing core for OPC correction. Simply put, to process a small area, a large block containing the small block is directly cut out and sent to the computing core for processing. Therefore, the block sent to the computing core is larger than block 102. The reference block is loaded into the computing core as a reference layer, and during the final splicing, only the corrected portion of block 102 is used for splicing.

[0029] Because the OPC correction result isn't a unique solution, and the external environment can influence internal conditions, coupled with the fact that the final correction result is determined through multiple correction cycles (typically over 20 on advanced nodes), there's a chance of convergence to different solutions. Consequently, when two corrected blocks (or small blocks) are spliced ​​together, there's a chance of small block boundary issues. This probability is roughly proportional to the sum of the boundaries of each block 102.

[0030] Generally speaking, current OPC operations are performed in a distributed computing environment. In these environments, each individual computing core and its allocated memory are treated as a separate machine, with little interaction between computing cores. This results in a large amount of redundant calculations, low computing efficiency, and wasted resources.

[0031] The following combination Figure 2 Provide a description. Figure 2 FIG. 1 shows a flow chart of OPC correction for a layout in a conventional solution. Figure 2 As shown, the blocks before correction are shown, wherein a total of 64 blocks are shown. It should be understood that the number of blocks shown here is schematic, and the actual number may vary according to actual needs.

[0032] Assume that there are 64 computing cores on a server participating in distributed computing (2 CPUs, each with 32 computing cores). Each block corresponds to a computing core, such as Figure 2 As shown in FIG, there are computing core 00, computing core 01, computing core 02, and so on to computing core 63. Each computing core has a corresponding allocated memory, and the corresponding blocks can be stored in the corresponding memory.

[0033] As mentioned earlier, during the entire OPC process, each computing core is independent of each other. Current OPC operations treat each individual computing core and its allocated memory as a single machine, with little interaction between them. This means that each computing core cannot access each other's memory. Without shared memory, OPC programs are relatively simple (they simply occupy memory during execution and prevent other processes from using it, eliminating the need to determine whether specific processes can access their allocated memory).

[0034] like Figure 2 As shown in , each modified block includes additional parts surrounding it, namely Figure 1The additional area 104 shown in FIG. However, the final splicing does not include the additional area 104. The final splicing is performed after each segmented block has been completely corrected. Therefore, the entire OPC calibration process requires a lot of additional work, namely the calibration work for the additional area 104.

[0035] Figure 3 Shown Figure 2 The solution shown shows the process interaction when correcting two blocks. As previously mentioned, throughout the OPC process, each computing core is independent of each other, meaning each core operates independently with little interaction. This leads to a large amount of redundant calculations, wasting resources and reducing efficiency.

[0036] In view of this, the present disclosure provides an improved solution.

[0037] Embodiments of the present disclosure provide an improved optical proximity correction (OPC) method. The method includes: each of multiple computing cores references information about adjacent blocks of a block to be corrected in a layout stored in shared memory to determine the impact of the adjacent blocks on the block to be corrected; performing OPC correction on the block to be corrected based on the information and impact of the block to be corrected to generate a correction block; and splicing each correction block according to the position of the original layout corresponding to each block to form a corrected layout. By referencing the information of adjacent blocks, the number of operations required to perform OPC correction on additional areas is reduced, significantly improving computational efficiency.

[0038] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0039] Figure 4 Schematic diagram of an example environment 400 in which embodiments according to the present disclosure can be implemented. Figure 4 As shown, the example environment 400 includes a computing device 410 and a client 420 .

[0040] In some embodiments, computing device 410 can interact with client 420. For example, computing device 410 can receive input messages from client 420 and output feedback messages to client 420. In some embodiments, the input messages from client 420 can be design layout data. Computing device 410 can perform corresponding mathematical operations on the design layout data and output the corresponding operation results to client 420.

[0041] In some embodiments, computing device 410 may include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant (PDA), a media player, etc.), consumer electronics, a minicomputer, a mainframe computer, cloud computing resources, etc.

[0042] It should be understood that the structure and functionality of the example environment 400 is described for illustrative purposes only and is not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions. This environment is merely illustrative and is not intended to limit the application environment of the embodiments of the present disclosure.

[0043] In order to explain the principle of the present disclosure more clearly, the following will refer to Figure 5 Let's describe it in more detail.

[0044] Figure 5 A flow chart of a method 500 for optical proximity correction according to some embodiments of the present disclosure is shown.

[0045] At block 502 , each of the plurality of computing cores refers to information of neighboring blocks of a block to be corrected in a layout stored in a shared memory to determine the impact of the neighboring blocks on the block to be corrected.

[0046] In some embodiments, each computing core can share a memory storing layout blocks. Therefore, when performing OPC correction on a block in the layout to be corrected, information about adjacent blocks stored in the memory can be referenced. This allows for comprehensive correction of the block to be corrected, taking into account the surrounding environment. The block to be corrected refers to the block corresponding to each computing core that requires correction.

[0047] The layout for OPC calibration is typically divided into multiple blocks. Traditional methods can be used to assign each block to a processing core or memory. For example, blocks can be assigned to different processing cores for OPC calibration in a predetermined order or randomly.

[0048] For example, the first one can be arranged as The block of the matrix is ​​assigned to the first one with The first server has 100 computing cores, where M and N are integers greater than or equal to 1; the second server is arranged into The blocks of the matrix are assigned to the second one with The second server with 100 computing cores; and so on, until all the servers are arranged into The blocks of the matrix are assigned to the corresponding For example, the first The matrix blocks are run on server A with 64 cores, and the second The matrix blocks are run on server B with 64 cores, .... In other words, the first The blocks of the matrix are assigned to the first server with 64 computing cores. The second The blocks of the matrix are assigned to the second server with 64 computing cores. After the correction of the matrix block, it will take over the Nth The blocks of the matrix are calculated in parallel to obtain the final result.

[0049] Normally, the split small blocks are sent to the CPU computing core one by one. Only after the previous small block is completed on the computing core can the next small block be received for processing. In addition, When small blocks of data are sent to a 64-core machine, they all go into their own different computing cores at once.

[0050] It's important to note that, regardless of whether or not the optimization algorithm is used, the resulting corrected region (corrected block 102) directly enters the final post-OPC layout stitching process (stitching is the final step in the OPC calculation process). In other words, all calculations involving the additional region 104 and reference region 106 are ultimately discarded. A key aspect of some embodiments of the present invention is minimizing these discarded calculations, thereby significantly improving efficiency.

[0051] Larger semiconductor manufacturers typically perform single OPC operations on servers purchased from the same batch. These servers generally have the same CPU, memory, operating system, and other configurations (this can reduce many problems). Therefore, when running OPC programs (usually thousands of computing cores), the computing environment is the same, which is very suitable for using the method described in the invention to speed up the process.

[0052] In some embodiments, each of the multiple computing cores separately references information about neighboring blocks of the block to be corrected in a layout stored in a shared memory, which may include: each computing core separately determining the distribution status of neighboring blocks at any boundary of the block to be corrected; and reading information about neighboring blocks that have an impact on any boundary of the block to be corrected based on the distribution status.

[0053] In some embodiments, reading information about neighboring blocks that affect any boundary of the block to be corrected based on the distribution may include: upon determining that any boundary of the block to be corrected has a neighboring block, reading information about a portion of the neighboring block near the boundary. If no neighboring block exists at any boundary, correction may be performed using conventional methods.

[0054] At block 504 , optical proximity correction is performed on the block to be corrected based on the information of the block to be corrected and the impact to generate a correction block.

[0055] In some embodiments, performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and the influence thereof may include performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and information of a portion near a boundary, so as to generate a correction block without an additional area at a boundary with an adjacent block.

[0056] In some embodiments of the present invention, each block is stored in memory, and each computing core can not only reference the corresponding memory but also access all memories, especially the memory on the same server. In this way, the information of adjacent blocks can be known, so during the OPC process, the information of adjacent blocks can be referenced or quoted. In some embodiments of the present disclosure, the computing core can access all memories by making some corresponding settings in the program.

[0057] In traditional solutions, the correction process for each computing core is independent. Each computing core processes the block to be corrected using a specified algorithm. To speed up data processing, independent memory is chosen. Therefore, with existing technology, no shared memory is used, and the OPC process is relatively simple. However, industry insiders have not yet realized that shared memory can reduce redundant computation.

[0058] During the OPC correction process, the computing core can reference (e.g., in real time or at a predetermined interval) data in memory. In some embodiments of the present disclosure, referencing data in memory in real time can mean reading data in memory in real time. Essentially, the computing core is part of the CPU, while the memory is plugged into the motherboard and is separate from the CPU (but still connected via the motherboard's bus). During the operation process, the computing core on the CPU references and modifies data in memory in real time. Furthermore, during the OPC correction process, the computing core updates the information of the corresponding block in memory in real time. Therefore, the computing core corresponding to the adjacent block can access the information of the adjacent block in real time. Therefore, the computing core can perform OPC correction on the block to be corrected based on the information of the corresponding block and the information of the adjacent blocks. In short, the computing core can refer to or understand the information of the adjacent blocks during the correction process, thus taking the influence of the adjacent blocks into account. This eliminates the need to correct the surrounding area of ​​the block to be corrected during the correction process, as the surrounding area is actually adjacent to the area to be corrected and therefore undergoes precise correction during the correction process. Conventional solutions, however, perform less precise correction on the surrounding area. Now that the situation of the adjacent areas is known, it is only necessary to refer to the information of the adjacent areas to correct the block to be corrected, without having to perform lower-precision corrections on the surrounding areas.

[0059] It should also be noted that the block information for CPU core calibration is stored in memory, so block information refers to the information stored in memory. The CPU's internal cache is very small and cannot hold block information.

[0060] In some embodiments, performing optical proximity correction on the block to be corrected based on the information and influence of the block to be corrected further includes: in response to determining that any boundary of the block to be corrected has no adjacent blocks, performing optical proximity correction on the block to be corrected and an additional area at the boundary without the adjacent blocks to generate a correction block with the additional area at the boundary without the adjacent blocks.

[0061] In some embodiments, performing OPC correction on the block to be corrected based on the information and influence of the block to be corrected may include: when any boundary of the block to be corrected has an adjacent block, performing OPC correction on the block to be corrected, and not performing OPC correction on the additional area at any boundary. In this case, the correction block does not include the corrected additional area at the boundary with the adjacent block except for the area corresponding to the block to be corrected. When any boundary of the block to be corrected does not have an adjacent block, performing OPC correction on the block to be corrected and the additional area at the boundary without the adjacent block to generate a correction block with the additional area. For example, in some embodiments, the adjacent The blocks (or some other combinations) are sent to the CPU cores installed on the same server (these CPU cores can communicate efficiently through the memory on the same server) to be calculated according to the optimized algorithm, using the method mentioned above "in a certain order".

[0062] In practice, commonly used distributed computing environments are built from multiple servers. Taking the x86 architecture as an example, the memory on each server is shared by its computing cores. Each server typically has two CPUs, each with dozens of computing cores. Assume that the servers participating in the distributed computing have a total of 64 computing cores (two CPUs, each with 32 cores), sharing 1024GB of memory (an average of 16GB per core).

[0063] In the aforementioned embodiments, each computing core is described as being able to access all memory. As previously mentioned, memory can be accessed in real time, and real-time access can be real-time reading. It should be understood that the present disclosure is not limited to this, that is, it is not limited to real-time reading of memory. As long as each computing core can obtain information about adjacent blocks in real time, for example, each core can communicate information about related blocks in memory to each other in a predetermined manner, which can also enable each core to obtain information about adjacent blocks in a timely manner.

[0064] Corrections to the additional regions can be understood as influencing corrections to the basic regions. Practice has demonstrated that such corrections can achieve ideal lithography results. Furthermore, in the present invention, at the boundaries of adjacent blocks, the correction process for each basic region simultaneously references (or, in other words, references) the information of the adjacent blocks. Therefore, the corrections to each basic block already account for the impact of the adjacent regions (corresponding to the additional regions in conventional solutions), eliminating the need for additional corrections to the peripheral regions of the basic region (the aforementioned additional regions). The resulting corrections can produce ideal lithography results.

[0065] At block 506 , the correction blocks are spliced ​​according to the positions of the original layout corresponding to the corresponding blocks to form a correction layout.

[0066] In some embodiments, the layout is optically proximity corrected on multiple servers in a distributed architecture, each server having multiple computing cores. The method further includes: splicing the corrected layout generated by the multiple computing cores on each server with the corrected layout generated by the computing cores located on other servers to form a final corrected layout.

[0067] In some embodiments, the plurality of computing cores are distributed on a server of one of the following architectures: X86 architecture; Advanced RISC Machine (ARM) architecture; Graphics Processing Unit (GPU) architecture.

[0068] In some embodiments, in each calibration cycle, scheduling is performed based on the computing status of each computing core, so that the computing core that has completed the computing operation first assists the computing core that has not completed the computing operation to perform the computing operation.

[0069] Refer to the following Figure 6 Aspects of some embodiments of the present disclosure are described. Figure 6 A schematic diagram illustrates a process for performing OPC correction on a layout according to some embodiments of the present disclosure. Assume that a server participating in distributed computing has 64 computing cores (two CPUs, each with 32 computing cores) and a shared 1024GB memory (an average of 16GB per computing core). Figure 6 64 blocks running on the same server are shown in FIG. For example, the blocks are put into the same 64-core server for calculation.

[0070] like Figure 6 As shown in Figure 2, which shows 64 computing cores, namely, computing core 00, computing core 01, computing core 02, and so on to computing core 63. As mentioned earlier, in traditional solutions, each computing core is independent of each other during the entire OPC process. Each individual computing core and its allocated memory are treated as a separate machine, and there is basically no linkage between computing cores.

[0071] In some embodiments of the present invention, Figure 2 The difference between the solution in is that each computing core can access all the memory on the server. The computing core can refer to or understand the situation of the adjacent blocks during the correction process. Therefore, the influence of the adjacent blocks is taken into account during the correction process. In this way, during the correction process, for the block to be corrected with adjacent blocks, there is no need to correct the surrounding area of ​​the block to be corrected, because the surrounding area is actually the adjacent area to be corrected, and precise correction will be performed in its corresponding correction process, while the traditional solution will perform lower-precision correction on the surrounding area. Now that the situation of the adjacent area is understood, there is no need to perform lower-precision correction on the surrounding area. Figure 6 As shown, each corrected block has an area corresponding to the additional area 104 only at an edge position with no adjacent blocks. Otherwise, there is no area corresponding to the additional area 104. That is, OPC correction does not need to be performed on the additional area 104 at an edge position with an adjacent block. In this way, a lot of extra work is avoided during the entire OPC correction process.

[0072] from Figure 6 As can be seen in the diagram, the first circle 602, second circle 604, third circle 606, fourth circle 608, and fifth circle 610 are shown, each representing a portion of the block. According to some embodiments of the present disclosure, for the corresponding small block (block to be corrected), there is no need to process the graphics in the third and fourth circles 606 and 608, for example, because these two circles are located in the aforementioned additional area at the boundary with adjacent blocks. Therefore, this information does not need to be output during the output. This is what is referred to as "generating a corrected block without an additional area at the boundary with adjacent blocks." During the specific calculation process, the corresponding information is actually replaced by the information in the second and fifth circles 604 and 610. Specifically, the information in the third circle 606 is replaced by the information in the fifth circle 610, and the information in the fourth circle 608 is replaced by the information in the second circle 604. In short, the information in the circle indicated by the arrow is replaced by the information in the circle indicated by the tail of the arrow.

[0073] For the small blocks in the "corner" (for example, the block corresponding to computing core 00), when outputting the corrected information, it is necessary to output the information in the first circle 602 (without the small block information running on the same server), and there is no need to output the information in the third circle 606 (filled with the information running on the same server).

[0074] It can be seen that, through the solution of the above embodiments of the present disclosure, a lot of extra work can be saved at the boundaries between adjacent blocks.

[0075] After each core completes the correction of the corresponding block, the corrected blocks are spliced ​​together to form the correction of the layout blocks completed by all cores on the server. Figure 6 After the initial stitching, the tiles can be stitched together with the corrected tiles on other servers to form the final corrected tile.

[0076] Figure 7 Shown Figure 6 The illustrated scheme illustrates the interaction between processes when correcting two blocks. As mentioned earlier, throughout the OPC process, memory is shared between all computing cores, meaning each core can access any memory on the server. Therefore, for example, a core can reference memory containing adjacent blocks to obtain information about the adjacent blocks, effectively accounting for the surrounding environment during the OPC correction process.

[0077] By uniformly accessing the layout information stored in the server memory, a large number of additional OPC operations can be avoided (note Figure 6 As shown in the figure, the output of the additional correction area is reduced, with only the edge blocks containing the additional area. This significantly reduces the computational effort. Furthermore, the correction pattern is formed as a single piece (the cores reference each other during the OPC process), so the connected areas do not suffer from small edge issues, such as alignment problems.

[0078] In some embodiments of the present disclosure, during operation, the basic correction region replaces both the basic correction region and the additional correction region in the standard OPC process by referencing each other, thereby significantly reducing the amount of computation.

[0079] In some embodiments, the OPC process can be further optimized. For example, in each calibration cycle, scheduling can be performed based on the computing status of each computing core, so that the computing core that completes the operation first assists the computing core that has not completed the operation to perform the operation.

[0080] In some embodiments, scheduling based on the computing status of each computing core so that the computing core that completes the operation first assists the computing core that has not completed the operation to perform the operation may include: first determining the number of segments in each block; sorting each block based on the determined number; when the first computing core that operates the first sorted block with the least number of sliced ​​fragments completes the cycle, the first computing core may assist the computing core that operates the last sorted block with the largest number of sliced ​​fragments to perform the operation in the cycle; when the second computing core that operates the second sorted block with the second least number of sliced ​​fragments completes the cycle, the second computing core may assist the computing core that operates the second last sorted block with the second largest number of sliced ​​fragments to perform the operation in the cycle; and so on, until the operations of all blocks of each sort in the correction cycle are completed.

[0081] After each round of computation, the program typically completes in about the same amount of time on each CPU. This is because memory references are involved, and these graphics must be synchronized during each cycle. Therefore, in this case, it is desirable to maintain a balanced computing core load on each CPU to further improve performance.

[0082] Typically, the amount of work required to correct each small block is similar and can be considered essentially the same. However, in some extreme cases, the amount of work required to correct each small block may differ significantly, which may result in efficiency loss when running the optimization process.

[0083] For example, for 64 small blocks, the computing core that processes one of the small blocks calculates very quickly. After it completes the calculation, the CPU waits for the next cycle, and there will be a loss of efficiency at this time.

[0084] In some embodiments, when encountering this situation, the following method can be used: before starting the operation on the entire 64 small blocks, the total number of fragments in each small block is counted. Fragments here refer to the segments formed by the edges of the graphics. As is known in the industry, when performing OPC correction on graphics, it is generally necessary to segment the graphics boundaries. Each segment has sampling points. The total number of these fragments in each small block is generally proportional to the operating speed.

[0085] Sort by the total number of shards, for example, assuming the following sequence: ; Among them, S1 represents the block with the least fragments, S 64 represents the block with the most shards, i.e., from S1 to S 64It should be understood that the sorting method shown here is only illustrative and other sorting methods can be used as needed, such as sorting in descending order of number.

[0086] After the above sorting, you can set it in the program: When the core operation of operation S1 completes one cycle, the subsequent operation S 64 Unfinished operations in (starting from sharding in reverse order); When the core operation of operation S2 completes one cycle, the subsequent operation S 63 Unfinished operations in (starting from sharding in reverse order); ...and so on.

[0087] In this way, the efficiency of OPC correction is further accelerated.

[0088] Some embodiments of the present disclosure utilize shared memory, references, or invocations to improve computational efficiency. This is particularly effective on a single server, where access to memory is very fast. For multiple servers, access can be achieved via network cables, etc., which, depending on network speed, is generally less efficient than the aforementioned embodiments executed on a single server.

[0089] Figure 8 The following diagram shows the size of a block cut out from the layout. The typical size of the small block cut out by OPC is roughly as follows: Figure 8 As shown. Among them, a is about 20um (micrometers), and b is about 2um. Taking this as a typical example, the OPC layout area that needs to be calculated for the above 64-core machine in this operating mode without optimization is (The actual output area is , so the effective calculation area is less than 70%); Under the optimized condition, the OPC layout area that needs to be calculated is (The actual output area is , thus effectively calculating over 95% of the area. The total length of the boundaries involved is 5120 μm without optimization and 640 μm with optimization. This shows that by reducing the computation of the additional area, the area of ​​the layout blocks requiring OPC correction is significantly reduced, significantly improving computational efficiency. Furthermore, the probability of small edge issues is significantly reduced, for example, to 1 / 8 of the original rate after optimization.

[0090] Currently, OPC programs are not only running on the x86 architecture, but have also begun to be promoted to other architectures, such as ARM architecture and GPU architecture. In these new architectures, the number of computing cores will be greatly improved compared to the x86 architecture (for example, a type of ARM architecture server has 384 computing cores per machine), but the computing power of a single core will be significantly lower than that of the x86 architecture. In this case, the total area of ​​the divided blocks will often be reduced, specifically in Figure 8 The size of a in [ 1 ] needs to be reduced (b remains unchanged). If the standard OPC operation process is used, the effective computing area ratio will be further reduced. Therefore, the solution of the embodiment of the present invention will play a greater role in the new operation architecture.

[0091] In some embodiments of the present disclosure, the computing data in different computing cores on the same CPU are linked, and multiple computing cores uniformly access the layout information in the memory of the same machine (or server), thereby dynamically optimizing the allocation of CPU computing resources.

[0092] Some embodiments of the present disclosure provide an optical proximity correction method. It should be noted that the examples given in the above embodiments are only for illustrating the solutions of the embodiments of the present disclosure and are not intended to limit the solutions of the present disclosure.

[0093] The solutions of some embodiments of the present disclosure can significantly improve the operating efficiency of the OPC program (increase the effective computing area), thereby reducing the problem of small block edges.

[0094] It should be understood that the embodiments shown in the drawings are only for schematically illustrating some embodiments of the present disclosure and are not intended to limit the present disclosure. The embodiments of the present disclosure may also have various other forms.

[0095] The present disclosure also discloses an electronic device. The electronic device includes: a processor; and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the device to perform actions, including: each of a plurality of computing cores respectively referencing information of adjacent blocks of a block to be corrected in a layout stored in a shared memory to determine the impact of the adjacent blocks on the block to be corrected; performing OPC correction on the block to be corrected based on the information and the impact of the block to be corrected to generate a correction block; and splicing each correction block according to the position of the original layout corresponding to each block to form a correction layout.

[0096] An embodiment of the present disclosure further discloses a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method according to the embodiment of the present disclosure is implemented.

[0097] Figure 9Schematic block diagrams of electronic devices according to some exemplary embodiments of the present disclosure are shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0098] like Figure 9 As shown, the device 900 includes a CPU 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the device 900. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0099] Multiple components in device 900 are connected to I / O interface 905, including an input unit 906, such as a keyboard, mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, optical disk, etc.; and a communication unit 909, such as a network card, modem, wireless communication transceiver, etc. The communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0100] The various processes and processing described above, such as method 500, can be executed by CPU 901. For example, in some embodiments, method 500 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by CPU 901, one or more steps in method 500 described above can be performed.

[0101] The solutions according to the embodiments of the present disclosure may be methods, devices, systems, and / or computer program products. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing various aspects of the present disclosure are loaded. The computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable program instructions may be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network.

[0102] Various embodiments of the present disclosure have been described above. The above descriptions are exemplary and are only optional embodiments of the present disclosure. They are not exhaustive and are not intended to limit the present disclosure. Although the claims in this application have been formulated for specific combinations of features, it should be understood that the scope of the present disclosure also includes any novel feature or any novel combination of features disclosed herein, whether explicitly or implicitly or in any generalization thereof, regardless of whether it relates to the same scheme in any claim currently claimed. It should be understood that new claims may be formulated to these features and / or combinations of these features during the examination of this application or in any further application derived therefrom.

[0103] The terminology used herein is selected to best explain the principles of the various embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the various embodiments disclosed herein. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this disclosure are intended to be included within the scope of protection of this disclosure.

Claims

1. An optical proximity correction method, comprising: Each of the plurality of computing cores refers to information of adjacent blocks of the block to be corrected in the layout stored in the shared memory to determine the influence of the adjacent blocks on the block to be corrected; performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and the impact to generate a correction block; as well as The correction blocks are spliced ​​according to the positions of the original layout corresponding to the corresponding blocks to form a correction layout.

2. The method according to claim 1, wherein each of the plurality of computing cores references information of adjacent blocks of the block to be corrected in the layout stored in the shared memory, comprising: Each of the computing cores determines the distribution status of adjacent blocks at any boundary of the block to be corrected; as well as Information of adjacent blocks that have an impact on any boundary of the block to be corrected is read based on the distribution status.

3. The method according to claim 2, wherein reading information of adjacent blocks that have an impact on any boundary of the block to be corrected based on the distribution condition comprises: In response to determining that any boundary of the block to be corrected has a neighboring block, information of a portion of the neighboring block close to the boundary is read.

4. The method according to claim 3, wherein performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and the impact comprises: Optical proximity correction is performed on the block to be corrected based on information of the block to be corrected and information of a portion close to the boundary, so as to generate a correction block without an additional area at a boundary with an adjacent block.

5. The method according to claim 3, wherein performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and the impact further comprises: In response to determining that any boundary of the block to be corrected has no adjacent blocks, optical proximity correction is performed on the block to be corrected and the additional area at the boundary without the adjacent blocks to generate a correction block with the additional area at the boundary without the adjacent blocks. 6 . The method according to claim 1 , wherein the correction block does not include a corrected additional area other than an area corresponding to the block to be corrected at a boundary with an adjacent block.

7. The method of claim 1 , wherein the optical proximity correction is performed on a plurality of servers in a distributed architecture, each server having a plurality of computing cores, the method further comprising: The correction layout generated by the multiple computing cores on each server is spliced ​​with the correction layout generated by the computing cores on other servers to form a final correction layout.

8. The method according to claim 7, wherein the plurality of computing cores are distributed on a server of one of the following architectures: X86 architecture; Graphics processor architecture; and Advanced RISC Machine Architecture.

9. The method according to claim 1, further comprising: In each calibration cycle, scheduling is performed based on the computing status of each computing core, so that the computing core that has completed the computing operation first assists the computing core that has not completed the computing operation to perform the computing operation.

10. The method according to claim 9, wherein scheduling based on the computing status of each computing core so that a computing core that has completed computing first assists a computing core that has not completed computing to perform computing comprises: Determine the number of shards in each block; sorting the blocks based on the quantity; In response to a first computing core operating on a first sorted block having a smallest number of shards completing the loop, causing the first computing core to assist a computing core operating on a last sorted block having a largest number of shards in performing the loop; In response to the second computing core operating on the second sorted block having the next smallest number of shards completing the loop, causing the second computing core to assist the computing core operating on the next last sorted block having the next largest number of shards in performing the loop; as well as And so on, until the calculation of all blocks of each sort in the correction cycle is completed.

11. The method according to claim 1, wherein the layout is divided into a plurality of blocks, and each block is sent to a different computing core for optical proximity correction in a predetermined order or randomly.

12. The method according to claim 11, wherein the blocks are sent to different computing cores in a predetermined order or randomly for optical proximity correction, comprising: Arrange the first one into The block of the matrix is ​​assigned to the first one with A first server with 100 computing cores, wherein M and N are integers greater than or equal to 1; Arrange the second one into The blocks of the matrix are assigned to the second one with A second server with 100 computing cores; as well as And so on, until all the The blocks of the matrix are assigned to the corresponding A server with 100 computing cores.

13. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the device to perform actions, the actions comprising: Each of the plurality of computing cores refers to information of adjacent blocks of the block to be corrected in the layout stored in the shared memory to determine the influence of the adjacent blocks on the block to be corrected; performing optical proximity correction on the block to be corrected based on the information of the block to be corrected and the impact to generate a correction block; and The correction blocks are spliced ​​according to the positions of the original layout corresponding to the corresponding blocks to form a correction layout.

14. A computer-readable storage medium having machine-executable instructions stored thereon, wherein when the machine-executable instructions are executed by a processor, the processor is caused to implement the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Load equilibration scheduling method and device

    CN101458634A

  • Load balancing method based on task distribution of multicore system

    CN102681902A

  • Integrated circuit optical proximity correction parallel processing method and system

    CN113777877A

  • Target layout graph forming method

    CN115600541A

  • Shared memory processing method and device, electronic equipment and storage medium

    CN116662037A