Instance segmentation mask generation method, device, equipment, storage medium and product
Patent Information
- Application Number
- CN202611048897.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-15
AI Technical Summary
[0004]本申请的主要目的在于提供一种实例分割掩码生成方法、装置、设备、存储介质及产品,旨在解决整体计算效率低下,难以满足普通工控机的硬件性能限制的技术问题
与相关技术中,通过深度学习技术对海量的SMT/SPI相关图像进行实例分割标注,但是图像处理算法存在内存资源消耗随并发进程数呈倍数增长(易引发内存溢出),且全图掩码遍历编码方式时间复杂度极高,导致整体计算效率低下,难以满足普通工控机的硬件性能限制相比,本申请响应于实例分割掩码生成指令,通过主进程获取待处理的图像信息,并将所述图像信息写入共享内存块;通过所述子进程读取所述共享内存块中对应的局部图像区域,得到局部掩码数据,并返回所述局部掩码数据至所述主进程,以使所述主进程基于所述局部掩码数据,生成全局掩码数据。本申请在收到实例分割掩码生成指令时,会通过主进程将获取的待处理图像信息写入共享内存块,通过子进程读取共享内存块中对应的局部图像区域,得到局部掩码数据,避免内存资源消耗随并发进程数呈倍数增长,且,仅读取对应的局部图像区域,避免局部图像区域,可以提升整体计算效率。
Smart Images

Figure CN122550606B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine vision technology, and in particular to an instance segmentation mask generation method, apparatus, device, storage medium and product. Background Technology
[0002] In the electronics manufacturing (EMS) industry, SMT (Surface Mount Technology) and SPI (Solder Plasma Inspection) processes are key steps to ensure the soldering quality of printed circuit boards (PCBs). Among them, the component and pad separation task of SMT / SPI is the core prerequisite for ensuring the accuracy of mounting and soldering.
[0003] In related technologies, deep learning technology is used to segment and label massive SMT / SPI related images. However, the image processing algorithm suffers from memory resource consumption that increases exponentially with the number of concurrent processes (which can easily lead to memory overflow), and the full-image mask traversal encoding method has extremely high time complexity, resulting in low overall computational efficiency and making it difficult to meet the hardware performance limitations of ordinary industrial control computers. Summary of the Invention
[0004] The main objective of this application is to provide a method, apparatus, device, storage medium, and product for generating instance segmentation masks, aiming to solve the technical problem of low overall computing efficiency and difficulty in meeting the hardware performance limitations of ordinary industrial control computers.
[0005] To achieve the above objectives, this application proposes an instance segmentation mask generation method, which includes: In response to the instance segmentation mask generation instruction, the image information to be processed is obtained through the main process and written into a shared memory block; The subprocess reads the corresponding local image region in the shared memory block to obtain local mask data, and returns the local mask data to the main process so that the main process can generate global mask data based on the local mask data.
[0006] In one embodiment, the step of reading the corresponding local image region in the shared memory block through a subprocess to obtain local mask data includes: Determine the subtask corresponding to the child process, wherein the subtask includes external rectangle information and access handle of the shared memory block; The child process reads the shared memory block corresponding to the access handle to obtain the local image region corresponding to the external rectangular box information. The image column containing the local image region is scanned column by column to obtain the initial mask data; The initial mask data is corrected to obtain local mask data.
[0007] In one embodiment, the step of correcting the initial mask data to obtain local mask data includes: Based on the image information, the overall image height and width are determined, and based on the outer rectangular frame information, the local width of the local image region and the horizontal coordinate of the preset position are determined. Based on the horizontal coordinate and the overall image height, calculate the number of blank pixels on the left side of the image column; based on the overall image height, the overall image width, the local width, and the horizontal coordinate, calculate the number of blank pixels on the right side of the image column. Based on the number of blank pixels on the left and the number of blank pixels on the right, the initial mask data is corrected to obtain local mask data.
[0008] In one embodiment, the initial mask data encoding includes a count value array and a pixel value array, and the step of correcting the initial mask data based on the number of blank pixels on the left and the number of blank pixels on the right to obtain local mask data includes: Based on the number of blank pixels on the left, the first count value of the count value array is corrected to obtain the initial count value array; Determine the pixel value corresponding to the last element in the pixel value array, and based on the pixel value and the number of blank pixels on the right, modify the initial count value array to obtain the target count value array; Based on the number of pixels of the left blank pixels and the right blank pixels respectively, add corresponding blank pixel values to the beginning and end of the pixel value array to obtain the target pixel value array; The local mask data is determined based on the target count value array and the target pixel value array.
[0009] In one embodiment, the target count value array includes a first count value array and a second count value array, and the step of correcting the initial count value array based on the pixel value and the number of blank pixels on the right to obtain the target count value array includes: If the pixel value is a preset first value, then based on the number of blank pixels on the right, the last count value of the initial count value array is corrected to obtain a first count value array, wherein the preset first value represents the background pixel; If the pixel value is a preset second value, then the number of blank pixels on the right is added as a count value to the initial count value array to obtain a second count value array, wherein the preset second value represents the foreground pixel.
[0010] In one embodiment, the image information includes an instance target and the bounding box information corresponding to the instance target, and the step of determining the subtask corresponding to the subprocess includes: Through the main process, a corresponding subtask is created for each instance target in the image information; Create a subprocess pool and distribute the subtasks sequentially to the idle subprocesses in the subprocess pool.
[0011] Furthermore, to achieve the above objectives, this application also proposes an instance segmentation mask generation apparatus, which includes: The writing module is used to respond to the instance segmentation mask generation instruction, obtain the image information to be processed through the main process, and write the image information into a shared memory block; The generation module is used to read the corresponding local image region in the shared memory block through the subprocess, obtain local mask data, and return the local mask data to the main process, so that the main process can generate global mask data based on the local mask data.
[0012] In addition, to achieve the above objectives, this application also proposes an instance segmentation mask generation apparatus, the apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the instance segmentation mask generation apparatus method as described above.
[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the instance segmentation mask generation apparatus method described above.
[0014] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the instance segmentation mask generation apparatus method described above.
[0015] One or more technical solutions proposed in this application have at least the following technical effects: Compared to related technologies that use deep learning to segment and label massive amounts of SMT / SPI related images, image processing algorithms suffer from memory resource consumption that increases exponentially with the number of concurrent processes (easily leading to memory overflow), and the extremely high time complexity of full-image mask traversal encoding, resulting in low overall computational efficiency and difficulty meeting the hardware performance limitations of ordinary industrial control computers, this application, in response to an instance segmentation mask generation instruction, obtains the image information to be processed through the main process and writes the image information into a shared memory block; the subprocess reads the corresponding local image region in the shared memory block to obtain local mask data, and returns the local mask data to the main process, so that the main process can generate global mask data based on the local mask data. When this application receives an instance segmentation mask generation instruction, it writes the obtained image information to be processed into a shared memory block through the main process, and the subprocess reads the corresponding local image region in the shared memory block to obtain local mask data. This avoids the exponential increase in memory resource consumption with the number of concurrent processes, and by only reading the corresponding local image region, avoiding local image regions, it can improve overall computational efficiency. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating an embodiment of the segmentation mask generation method of this application. Figure 2 A global coordinate mapping diagram based on a columnar mask view for the segmentation mask generation method of this application; Figure 3 This is a flowchart illustrating Embodiment 2 of the segmentation mask generation method in this application. Figure 4 This is a schematic diagram of the module structure of the segmentation mask generation device in an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the instance segmentation mask generation method in this application embodiment.
[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0022] The main solution of this application embodiment is: in response to the instance segmentation mask generation instruction, the main process obtains the image information to be processed and writes the image information into a shared memory block; the subprocess reads the corresponding local image region in the shared memory block to obtain local mask data and returns the local mask data to the main process, so that the main process generates global mask data based on the local mask data.
[0023] In related technologies, deep learning technology is used to segment and label massive SMT / SPI related images. However, the image processing algorithm suffers from memory resource consumption that increases exponentially with the number of concurrent processes (which can easily lead to memory overflow), and the full-image mask traversal encoding method has extremely high time complexity, resulting in low overall computational efficiency and making it difficult to meet the hardware performance limitations of ordinary industrial control computers.
[0024] When this application receives an instance segmentation mask generation instruction, it writes the acquired image information to be processed into a shared memory block through the main process, and reads the corresponding local image region in the shared memory block through the child process to obtain local mask data. This avoids memory resource consumption from increasing exponentially with the number of concurrent processes. Furthermore, by only reading the corresponding local image region, avoiding local image regions, the overall computational efficiency can be improved.
[0025] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or instance segmentation mask generation device capable of performing the above functions. The following description uses an instance segmentation mask generation device as an example to illustrate this embodiment and the subsequent embodiments.
[0026] Based on this, embodiments of this application provide a method for generating an instance segmentation mask, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the segmentation mask generation method of this application.
[0027] In this embodiment, the instance segmentation mask generation method includes steps S10 to S20: Step S10: In response to the instance segmentation mask generation instruction, the image information to be processed is obtained through the main process and written into the shared memory block; It should be noted that the execution entity in this embodiment is the instance segmentation mask generation device. The instance segmentation mask generation instruction refers to the external trigger signal or control command that initiates the entire instance segmentation mask generation process. The main process refers to the primary control process in a multi-process computer program, responsible for initializing the runtime environment, allocating system resources, loading data, and managing the working state of other child processes. Image information refers to ultra-high resolution images (such as 100-megapixel-level SPI images), including but not limited to the grayscale or color values of all pixels, as well as the width and height dimensions of the image. A shared memory block refers to a specific physical memory region allocated in the operating system kernel space, which allows multiple different processes to directly access it through memory address mapping, thereby enabling cross-process reuse of the same physical data.
[0028] Understandably, after receiving the instance segmentation mask generation instruction, the instance segmentation mask generation device retrieves the original ultra-high resolution image to be processed and its related label matrix and other image information from external storage or devices through the main process. This image data is then completely loaded and written into a shared memory block in the operating system kernel space. This provides a unified, single data access source for all subsequent child processes that need to process the image, eliminating the need to copy physical data. This avoids the situation where multiprocessing accelerates the processing of ultra-high resolution images (such as megapixel-level SPI images), which can cause memory usage to increase linearly with the number of processes, easily exceeding the hardware physical memory limit and triggering an OutOfMemory (OOM) crash.
[0029] Step S20: The subprocess reads the corresponding local image region in the shared memory block to obtain local mask data, and returns the local mask data to the main process so that the main process can generate global mask data based on the local mask data.
[0030] It is understood that a subprocess refers to an independent working process created and derived from the main process, used to concurrently execute specific image processing tasks assigned to it. A local image region refers to a subset of pixels from the original complete image corresponding to the image information, allocated to a specific instance target for processing by the subprocess. Local mask data refers to the mask-encoded data generated by the subprocess after performing instance segmentation processing on the local image region it is responsible for, covering only that local area. Global mask data refers to the final mask result that completely covers the entire original image after stitching and mapping all the local mask data according to their absolute positions in the original complete image.
[0031] It should be noted that the subprocesses of the instance segmentation mask generation device directly access the shared memory block through memory address mapping, read the pixel data of the local image region allocated to them for processing, run the segmentation algorithm on the local image region to obtain the corresponding local mask data, and then return the local mask data to the main process. After receiving the local mask data returned by all the subprocesses, the main process splices and integrates them according to the relative position of each local image region in the original complete image, thereby generating the global mask data covering the entire original complete image.
[0032] Since the subprocesses directly read the corresponding local image regions from the shared memory block, the overhead of copying the complete image data from the main process is completely avoided. Furthermore, each subprocess only needs to return the processed local mask data, and the main process only needs to perform simple spatial concatenation of these local results to generate the global mask data. This eliminates the need to generate a full-size mask for each target in the full image data, thus making the computational load of the data aggregation stage independent of the total number of pixels in the original image and only related to the amount of effective local mask data returned by each subprocess. This further ensures the overall computational efficiency during large-scale concurrent processing.
[0033] In one feasible implementation, step S20 includes: Determine the subtask corresponding to the subprocess, wherein the subtask includes external rectangle information and access handle of the shared memory block; Understandably, a subtask refers to the smallest unit of work to be completed within a parallel processing framework, allocated to a specific subprocess for independent processing. The outer bounding box information refers to the set of bounding rectangle parameters (x, y, w, h) in image space used to describe the spatial extent of a target to be processed. It typically includes the horizontal and vertical coordinates of the top-left vertex of the rectangle, as well as the rectangle's width and height. The access handle refers to a reference identifier provided by the operating system kernel, used to uniquely identify and locate the shared memory block. A process can establish a memory address mapping relationship with the shared memory block through this handle. The instance segmentation mask generation device determines the subtask that each initiated subprocess needs to handle. This subtask explicitly includes the bounding box information of the target to be processed in the original complete image, as well as the access handle for accessing the shared memory block of the globally unified data source. After obtaining the subtask, the subprocess can determine the spatial range it needs to process based on the bounding box information, and establish a mapping connection with the shared memory block based on the access handle. This provides complete context information for the subprocess to independently and accurately start subsequent local image region reading operations.
[0034] The child process reads the shared memory block corresponding to the access handle to obtain the local image region corresponding to the external rectangular box information. It should be noted that after obtaining the subtask, the subprocess of the instance segmentation mask generation device uses the access handle carried in the subtask to initiate a shared memory mount request to the operating system kernel, establishing an address mapping relationship with the shared memory block, thereby obtaining read permission for the original complete image data stored in the shared memory block. Subsequently, based on the outer bounding box information carried in the subtask, the subprocess accurately locates and reads the pixel data within the spatial range defined by the outer bounding box from the complete image data corresponding to the mapped shared memory block, obtaining the local image region corresponding to the outer bounding box information. Thus, the subprocess does not need to access or read any other irrelevant pixels in the original complete image, only obtaining the minimum data set required for its own task.
[0035] The image column containing the local image region is scanned column by column to obtain the initial mask data; It is understood that an image column refers to a set of pixels distributed vertically in the local image region, consisting of all pixels at the same horizontal coordinate position. The initial mask data (RLE sequence) refers to the original mask encoding sequence obtained directly through the column-by-column scan, arranged in column-priority order, which has not yet undergone any global coordinate offset correction or boundary padding processing. After obtaining the local image region corresponding to the outer rectangular box information, the subprocess of the instance segmentation mask generation device performs a column-by-column scan operation on the local image region. That is, following the column order from left to right, it sequentially reads the mask values of pixels from top to bottom for each image column, and compresses consecutively repeating foreground or background pixel values into run-length count form, thereby generating the initial mask data corresponding to the local image region, arranged in column-priority order.
[0036] Specifically, when processing local image regions, the subprocesses of the instance segmentation mask generation device logically treat them as regions with a height equal to the original image. Figure 1 A vertical stripe with a width (w) consistent with the target. Due to the column-first scanning method, the data order within this stripe is naturally consistent with the target. Figure 1 Therefore, the intermediate data segment does not require any modification. The initial mask data is then corrected to obtain the local mask data.
[0037] Optionally, the step of scanning the image column containing the local image region column by column to obtain the initial mask data includes: Based on the distribution of continuous pixel values in each image column in the local image region, the start and end row positions of the foreground pixels in the image column are determined; based on the start and end row positions, the foreground pixel column segment of the image column is determined; the foreground pixel column segment is converted into the corresponding run count to obtain the initial mask data.
[0038] It should be noted that the starting row position refers to the row index number of the first pixel to display the preset second value (i.e., foreground pixel) when scanning a column of images from top to bottom. The ending row position refers to the row index number of the last pixel to display the preset second value (i.e., foreground pixel) when scanning a column of images from top to bottom. A foreground pixel column segment refers to a continuous pixel segment in a column of images from the starting row position to the ending row position, where all pixel values within this segment are the preset second value (i.e., foreground pixel). When the subprocess of the instance segmentation mask generation device performs a column-by-column scan of the local image region, for each image column, it reads the pixel values of all pixels in that column in a top-to-bottom order, and records the row position where the preset second value first appears as the starting row position and the row position where the preset second value last appears as the ending row position. Subsequently, the subprocess identifies the continuous pixel segments between the starting and ending row positions as the foreground pixel column segments. Finally, the subprocess uses the length of the foreground pixel column segment (i.e., the ending row position minus the starting row position plus one) as the foreground count, and the lengths of the background pixels above and below the foreground segment in that column as the background counts, respectively. These run-length counts are written sequentially into the RLE sequence in column-major order to obtain the initial mask data. Therefore, the subprocess only needs to record the start and end positions of the foreground pixels in each column to generate a complete run-length encoding, without needing to record the position of each foreground pixel in the column pixel by pixel.
[0039] In SMT / SPI industrial inspection scenarios, each device or pad target to be inspected typically appears as a continuous foreground pixel range within a single column of pixels, with background pixels above and below this range. The height of the local image region is usually on the order of several thousand pixels, and each column contains a large number of continuous background pixels. If the run count of each pixel were recorded one by one, the amount of data required would be equal to the image height. However, since this step directly aggregates the foreground pixels of each column into a single continuous segment by recording the start and end row positions, only the start and end position information of the foreground segment in the column needs to be extracted to generate the corresponding run count. Compared with the pixel-by-pixel scanning and recording method, the processing complexity of a single column is reduced from being related to the total number of pixels in the column to being related to the number of foreground segments. Furthermore, since the width of the local image region is usually greater than its height, and the target distribution in the width direction is much larger than in the height direction, this difference is further amplified. In industrial scenarios where the target area accounts for a relatively small proportion, this results in a significant order-of-magnitude improvement in the generation speed of the initial mask data.
[0040] The step of determining the start and end row positions of foreground pixels in the image column based on the continuous pixel value distribution of each image column in the local image region includes: Obtain the vertical coordinates of the top-left corner of the outer rectangle of the local image region as the reference row offset; determine the vertical scanning interval of each image column based on the reference row offset and the height of the local image region; within the vertical scanning interval, read pixel values sequentially in ascending order of row index, and record the row index of the first pixel equal to the preset second value as the candidate start row position; within the vertical scanning interval, read pixel values sequentially in descending order of row index, and record the row index of the first pixel equal to the preset second value as the candidate end row position; if no pixel value equal to the preset second value is read within the vertical scanning interval, determine the image column as a pure background column, mark both the candidate start row position and the candidate end row position as null values, and skip the generation of run count for this column.
[0041] It is understood that the reference row offset refers to the vertical starting position of the local image region in the original complete image, that is, the vertical coordinate value of the top left corner vertex in the outer rectangular box information, used to convert the row index in the local coordinate system to the row index in the global coordinate system. The vertical scan interval refers to the range of row indices of pixel values that actually need to be read when scanning each column of the image. The lower limit of this range is the top boundary of the local image region (i.e., row index 0), and the upper limit is the bottom boundary of the local image region (i.e., row index is the local height minus 1). Before the subprocess of the instance segmentation mask generation device begins scanning the local image region column by column, it first extracts the vertical coordinates of the top-left vertex of the outer rectangular frame from the bounding box information and records them as the reference row offset. Then, using this reference row offset as the starting point and the height of the local image region as the range, the subprocess determines the vertical scanning interval for each image column to be read, with the row index ranging from 0 to the local height minus 1. Next, within the vertical scanning interval, the subprocess reads pixel values row by row from the first pixel of the column in ascending order of row index and compares them. If the currently read pixel value equals the preset second value, the scanning in that direction is immediately stopped. The process scans the image and records the current row index as the candidate start row position. Simultaneously, within the vertical scanning interval, the subprocess reads pixel values row by row upwards from the last pixel of the column, following the order of row indices from largest to smallest, and compares them. If the currently read pixel value equals the preset second value, the scanning in that direction is immediately stopped, and the current row index is recorded as the candidate end row position. If no pixel value equal to the preset second value is read during the scanning process in both the top-to-bottom and bottom-to-top directions within the vertical scanning interval, the subprocess determines that the column is the pure background column, marks the corresponding candidate start row position and candidate end row position as null, and skips any run-length counting operation for that column. Therefore, each image column only needs to be scanned from its top downwards to the position of the first foreground pixel and from its bottom upwards to the position of the last foreground pixel to complete the extraction of the column's start and end positions, without needing to traverse the intermediate pixels of the column.Because this step uses a two-way squeezing approach, starting from the top of the column and ending from the bottom, to determine the candidate start row position and the candidate end row position respectively, only a limited number of background pixels before the first foreground pixel and after the last foreground pixel in the column need to be read before the scanning can stop. This completely skips the traversal of all intermediate pixels inside the foreground segment and between the upper and lower boundaries. Furthermore, since the height of the local image region is usually on the order of several thousand pixels, while the height of the foreground segment of a single target in a single column is usually only tens to hundreds of pixels, this two-way squeezing strategy reduces the amount of data read per column scan from linearly related to the total height of the column to only related to the distance from the top of the target to the top of the column and the distance from the bottom of the target to the bottom of the column. In industrial scenarios where the target is centrally distributed and occupies a relatively low column height, the number of pixels accessed in the overall column-by-column scanning can be significantly reduced by orders of magnitude.
[0042] The initial mask data is corrected to obtain local mask data.
[0043] It should be noted that after obtaining the initial mask data by scanning column by column, the subprocess of the instance segmentation mask generation device performs a correction operation on the initial mask data.
[0044] Optionally, after the step of correcting the initial mask data to obtain local mask data, the method includes: acquiring processed neighborhood mask data adjacent to the local image region; determining the edge continuity features between the local image region and the neighborhood based on the processed neighborhood mask data; and smoothing the edge contours of the local mask data based on the edge continuity features to obtain corrected local mask data.
[0045] It is understood that the processed neighborhood mask data refers to the local mask data corresponding to other local image regions spatially adjacent to the local image region currently being processed by the sub-process, which has been corrected by other sub-processes and returned to the main process. Edge continuity feature refers to a consistency metric indicating whether the mask pixel values on both sides of the common boundary between the local image region and the region covered by the processed neighborhood mask data exhibit a continuous transition. After the subprocess of the instance segmentation mask generation device completes the correction of the initial mask data to obtain the local mask data, before returning the data to the main process, it first obtains the processed neighborhood mask data that is adjacent to the current local image region and has been processed and submitted by other subprocesses from the main process. Then, the subprocess extracts the boundary pixel sequence on the side of the common boundary in its own local mask data and extracts the boundary pixel sequence on the side of the same common boundary in the processed neighborhood mask data. The two are compared, and the pixel value difference rate on both sides of the boundary is calculated as the edge continuity feature. If the edge continuity feature indicates that the difference rate on both sides of the boundary is lower than a preset threshold (i.e., the contours on both sides are smooth and consistent), no correction is performed. If the indication difference rate is higher than the preset threshold (i.e., there is a break or misalignment in the contours on both sides), the subprocess performs a pixel-by-pixel comparison of its own side boundary pixel sequence according to the neighborhood side boundary pixel sequence, and adjusts the pixel values at the difference positions to be consistent with the neighborhood side, thereby completing the smoothing correction and obtaining the corrected local mask data. This ensures that the local regions processed by adjacent subprocesses will not experience mask breaks or misalignments at the boundaries.
[0046] In one feasible implementation, the step of correcting the initial mask data to obtain local mask data includes: Based on the image information, the overall image height and width are determined, and based on the outer rectangular frame information, the local width of the local image region and the horizontal coordinate of the preset position are determined. It is understood that the overall image height Hfull refers to the total number of rows of pixels in the vertical direction of the original complete image corresponding to the image information. The overall image width Wfull refers to the total number of columns of pixels in the horizontal direction of the original complete image corresponding to the image information. The local width w refers to the number of columns of pixels contained in the local image region in the horizontal direction, that is, the width value of the outer rectangular frame information. The horizontal coordinate x of the preset position refers to the horizontal starting position of the outer rectangular frame information in the original complete image, which can be set as the horizontal coordinate value of the upper left corner vertex of the outer rectangular frame. The subprocess of the instance segmentation mask generation device reads the size metadata of the original complete image from the image information to determine the overall image height and the overall image width; at the same time, it extracts the width value of the rectangular frame from the outer rectangular frame information carried in the subtask as the local width, and extracts the horizontal coordinate value of the upper left corner vertex of the rectangular frame as the horizontal coordinate of the preset position. Thus, it provides the necessary calculation parameters for the subsequent head offset correction and tail filling loop closure of the initial mask data.
[0047] Based on the horizontal coordinate and the overall image height, calculate the number of blank pixels on the left side of the image column; based on the overall image height, the overall image width, the local width, and the horizontal coordinate, calculate the number of blank pixels on the right side of the image column. It should be noted that the left blank pixel count refers to the total number of background pixels contained between the left boundary of the image and the left boundary of the local image region in the original complete image. The right blank pixel count refers to the total number of background pixels contained between the right boundary (x+w) of the local image region (width w, height h) and the right boundary (Wfull) of the image in the original complete image. The subprocess of the instance segmentation mask generation device multiplies the horizontal coordinate of the preset position by the overall image height, and the product is the number of blank pixels on the left side of the image column; simultaneously, the subprocess subtracts the sum of the horizontal coordinate of the preset position and the local width from the overall image width, and the difference is the number of blank columns on the right side of the local image region. This blank column number is then multiplied by the overall image height, and the product is the number of blank pixels on the right side of the image column. Thus, the head offset correction amount required for subsequent head offset correction of the initial mask data and the tail padding amount required for tail padding loop closure are obtained respectively.
[0048] Specifically, the number of blank pixels on the left side = x * Hfull, and the number of blank pixels on the right side = (Wfull) * Hfull. x w)*Hfull.
[0049] Based on the number of blank pixels on the left and the number of blank pixels on the right, the initial mask data is corrected to obtain local mask data.
[0050] Understandably, the subprocess of the instance segmentation mask generation device modifies the initial mask data based on the number of blank pixels on the left and the number of blank pixels on the right.
[0051] Specifically, the subprocess adds the number of left-side blank pixels to the first count element of the RLE encoding sequence corresponding to the initial mask data, completing the head offset correction. This expands the head count from covering only the left-side background of a local area to covering the entire left-side blank area from the left boundary of the original complete image to the left boundary of the local image region. Subsequently, the subprocess obtains the last count element of the RLE sequence of the initial mask data and determines whether the last count element corresponds to a foreground count or a background count. If it corresponds to a background count, the number of right-side blank pixels is merged into the last count element. If it corresponds to a foreground count, a new count element is appended to the end of the RLE sequence, with a value equal to the number of right-side blank pixels and corresponding to the background count, completing the tail filling loop closure. After the head and tail corrections are completed, the local mask data is obtained. Thus, by modifying only the head and tail count positions of the RLE sequence, a complete mapping from local relative coordinates to global absolute coordinates is achieved.
[0052] In one feasible implementation, the initial mask data encoding includes a count value array and a pixel value array, and the step of correcting the initial mask data based on the number of blank pixels on the left and the number of blank pixels on the right to obtain local mask data includes: Based on the number of blank pixels on the left, the first count value of the count value array is corrected to obtain the initial count value array; It should be noted that the count value array refers to the numerical sequence used in the RLE encoding format to store the number of times consecutive identical pixel values appear repeatedly. For example, if pixel values [0, 0, 0, 1, 1, 1, 1, 0, 0] are scanned sequentially (3 background, 4 foreground, 2 background), the count value array would be [3, 4, 2], indicating that the first segment appears 3 times consecutively, the second segment appears 4 times consecutively, and the third segment appears 2 times consecutively. The instance segmentation mask generation device adds the number of blank pixels on the left to the first count value in the count value array corresponding to the initial mask data, and replaces the original first count value with the result of the addition, thus completing the correction of the beginning of the count value array and obtaining the initial count value array. Therefore, the first count value of the count value array is corrected from originally only representing the number of background pixels on the left side of the local image region to the total number of blank pixels on the left side from the left boundary of the original complete image to the left boundary of the local image region, realizing global coordinate alignment of RLE encoding in the horizontal direction.
[0053] Determine the pixel value corresponding to the last element in the pixel value array, and based on the pixel value and the number of blank pixels on the right, modify the initial count value array to obtain the target count value array; It is understood that the pixel value array refers to a marker sequence in the RLE encoding structure that corresponds one-to-one with the count value array and is used to identify the pixel category (foreground or background) represented by each count value. For example, 0 represents background pixels and 1 represents foreground pixels. This array and the count value array together constitute the complete RLE encoded data. After obtaining the initial count value array, the subprocess of the instance segmentation mask generation device determines the pixel value corresponding to the last element in the pixel value array corresponding to the initial count value array, that is, it determines whether the end of the current RLE sequence ends with a foreground pixel value or a background pixel value. Subsequently, the subprocess performs tail correction on the initial count value array based on the pixel value and the number of blank pixels on the right. If the pixel value corresponding to the last element is a background pixel value, the number of blank pixels on the right is added to the last count value of the initial count value array. If the pixel value corresponding to the last element is a foreground pixel value, a new count value is appended to the end of the initial count value array. This count value is equal to the number of blank pixels on the right and the corresponding pixel value is a background pixel value. This completes the tail correction and obtains the target count value array.
[0054] Based on the number of pixels of the left blank pixels and the right blank pixels respectively, add corresponding blank pixel values to the beginning and end of the pixel value array to obtain the target pixel value array; It should be noted that blank pixel values refer to pixel category markers used to identify background areas in an image. In a binary mask, they are typically represented by the value 0, contrasting with pixel values representing foreground target areas (usually 1). After obtaining the initial count array, the subprocess of the instance segmentation mask generation device synchronously corrects the pixel value array corresponding to the initial count array. Specifically, the subprocess adds a corresponding number of blank pixel values to the beginning of the pixel value array based on the number of blank pixels on the left, so that the newly added blank pixel values at the beginning correspond to the increment of the count value at the beginning. Simultaneously, the subprocess obtains the pixel value corresponding to the last element in the pixel value array. If the last element corresponds to a blank pixel value, the existing blank pixel value at the end is merged with the newly added blank area on the right, without needing to add a new element at the end. If the last element corresponds to a foreground pixel value, a blank pixel value is appended to the end of the pixel value array, so that the newly added blank pixel value at the end corresponds to the increment or appended item of the count value at the end. After the beginning and end corrections are completed, the target pixel value array is obtained. Thus, the target pixel value array and the target count value array are structurally perfectly aligned, together forming a complete global RLE encoded data that can be directly used for serialization and storage.
[0055] The local mask data is determined based on the target count value array and the target pixel value array.
[0056] It is understandable that after obtaining the target count value array and the target pixel value array, the subprocess of the instance segmentation mask generation device combines these two arrays according to the RLE encoded data structure, so that each count value in the target count value array is paired one-to-one with the pixel value at the corresponding position in the target pixel value array, together forming a complete RLE encoded data that has been corrected by global coordinate mapping, and this complete RLE encoded data is determined as the local mask data.
[0057] In one feasible implementation, the target count value array includes a first count value array and a second count value array, and the step of correcting the initial count value array based on the pixel value and the number of blank pixels on the right to obtain the target count value array includes: If the pixel value is a preset first value, then based on the number of blank pixels on the right, the last count value of the initial count value array is corrected to obtain a first count value array, wherein the preset first value represents the background pixel; It should be noted that the preset first value refers to a predefined specific numerical marker used to identify background pixels in the pixel value array, such as the value 0. This value is compared with the preset second value (such as the value 1) representing foreground pixels. After determining the pixel value corresponding to the last element in the pixel value array, the subprocess of the instance segmentation mask generation device determines whether the pixel value is the preset first value. If the pixel value is determined to be the preset first value, it indicates that the end of the current RLE sequence is a background pixel. Then, the subprocess obtains the number of blank pixels on the right and adds this number of blank pixels to the last count value of the initial count value array. The accumulated result replaces the original last count value to obtain the first count value array. Thus, the existing background count at the end and the count of the newly added blank area on the right are merged into a single count value, avoiding the addition of new elements to the end of the array.
[0058] If the pixel value is a preset second value, then the number of blank pixels on the right is added as a count value to the initial count value array to obtain a second count value array, wherein the preset second value represents the foreground pixel.
[0059] It is understood that the preset second value refers to a predefined specific numerical marker used to identify foreground target pixels in the pixel value array, such as the value 1. This value is compared with the preset first value (such as the value 0) representing background pixels. After determining the pixel value corresponding to the last element in the pixel value array, the subprocess of the instance segmentation mask generation device determines whether the pixel value is the preset second value. If the pixel value is determined to be the preset second value, it indicates that the end of the current RLE sequence ends with a foreground pixel. Then, the subprocess obtains the number of right-side blank pixels and adds this number as a new count value to the end of the initial count value array to form the second count value array. Thus, an independent background count item is added after the original foreground count to represent all right-side blank areas from the right boundary of the local image region to the right boundary of the original complete image.
[0060] Specifically, within the current child process's view, the target is located within the entire image with a width of Wfull, with x columns of blank space to its left, its own width of w, and (Wfull - x - w) columns of blank space to its right. The total number of pixels in these three parts constitutes the area of the entire image. The total number of pixels in these three parts constitutes the area of the entire image. Based on the above data, a three-step correction is performed: Head correction: The total number of blank pixels on the left (x * Hfull) is added to the first count value of the local RLE to complete the global start-point alignment; Body preservation: All count values in the middle are directly copied without any traversal or modification; Tail loop closure: The total number of blank pixels on the right Pad = (Wfull - x - w) * Hfull is calculated. If the end of the RLE is the background, the counts are merged; if it is the foreground, a new count segment is added. Finally, the global RLE mask is output, whose total number of pixels is strictly equal to Wfull * Hfull, and the origin of the coordinates is aligned with the top left corner of the entire image. Figure 2 , Figure 2 A global coordinate mapping graph based on a columnar mask view is provided.
[0061] In this embodiment, the operation of appending the number of blank pixels on the right as a new count value to the end of the initial count value array is only performed when the last element of the pixel value array corresponds to the preset second value. This operation only involves adding a single element to the end of the array and does not require traversing or modifying any middle part of the initial count value array. Furthermore, since this operation makes the newly added blank area on the right attached to the original foreground count as an independent count value item, the strict correctness of the pixel value alternation rule (background, foreground, background, etc.) in RLE encoding is maintained while ensuring that the total number of pixels is strictly equal to the product of the height and width of the entire image. This avoids encoding errors caused by merging counts of different pixel types and further ensures the data integrity and format accuracy of the local mask data in the global coordinate system.
[0062] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 The image information includes the instance target and the outer bounding box information corresponding to the instance target. Before the step of determining the subtask corresponding to the subprocess, the instance segmentation mask generation method further includes steps S01~S02: Step S01: Through the main process, create a corresponding subtask for each instance target in the image information; It should be noted that an instance target refers to each individual object to be detected in the image information that needs to be segmented and labeled, such as each device or each pad. After performing target recognition operations such as connected component analysis on the image information, the main process of the instance segmentation mask generation device creates a corresponding subtask for each identified instance target. Each subtask contains the outer bounding box information corresponding to the instance target and the access handle of the shared memory block.
[0063] Specifically, the instance segmentation mask generation device identifies independent targets in an image and extracts their geometric information (x, y, w, h) through conventional connected component analysis and other methods.
[0064] Step S02: Create a subprocess pool and distribute the subtasks sequentially to the idle subprocesses in the subprocess pool.
[0065] Understandably, the main process of the instance segmentation mask generation device creates a subprocess pool containing multiple subprocesses. Then, the created subtasks for each instance target are sequentially assigned to currently idle subprocesses according to the working status of each subprocess in the subprocess pool. This achieves dynamic scheduling and concurrent execution of multiple subtasks among multiple subprocesses.
[0066] In this embodiment, since the main process centrally manages multiple subprocesses through the subprocess pool and dynamically distributes the subtasks to the idle subprocesses, it is not necessary to wait for a subprocess to finish processing all tasks serially before processing the next one, nor is it necessary to stop the overall process due to the blocking of a subprocess. This achieves efficient reuse of task execution units and dynamic balancing of computational load, effectively squeezing the parallel computing power of multi-core CPUs. It ensures that even with a large number of instance targets, the overall processing time is significantly shortened as the number of parallel subprocesses increases, greatly improving the annotation throughput in scenarios with massive numbers of small targets.
[0067] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the segmentation mask generation method of this application. Any simple transformations based on this technical concept are within the protection scope of this application.
[0068] This application also provides an instance segmentation mask generation apparatus, please refer to... Figure 4 The instance segmentation mask generation device includes: The writing module 10 is used to respond to the instance segmentation mask generation instruction, obtain the image information to be processed through the main process, and write the image information into the shared memory block; The generation module 20 is used to read the corresponding local image region in the shared memory block through the subprocess, obtain local mask data, and return the local mask data to the main process, so that the main process can generate global mask data based on the local mask data.
[0069] Optionally, the generation module includes: The scanning submodule is used to determine the subtask corresponding to the subprocess, wherein the subtask includes external bounding box information and access handle of the shared memory block; the subprocess reads the shared memory block corresponding to the access handle to obtain the local image region corresponding to the external bounding box information; the image column containing the local image region is scanned column by column to obtain initial mask data; the initial mask data is corrected to obtain local mask data.
[0070] Optionally, the scanning submodule includes: The correction unit is used to determine the overall image height and width based on the image information; determine the local width and the horizontal coordinate of a preset position of the local image region based on the outer rectangular frame information; calculate the number of blank pixels on the left side of the image column based on the horizontal coordinate and the overall image height; calculate the number of blank pixels on the right side of the image column based on the overall image height, the overall image width, the local width, and the horizontal coordinate; and correct the initial mask data based on the number of blank pixels on the left and the number of blank pixels on the right to obtain local mask data.
[0071] A creation unit is used to create a corresponding subtask for each instance target in the image information through the main process; and to create a subprocess pool to distribute the subtasks to idle subprocesses in the subprocess pool in sequence.
[0072] Optionally, the correction unit includes: A sub-unit is added to correct the first count value of the count value array based on the number of blank pixels on the left, to obtain an initial count value array; the pixel value corresponding to the last element in the pixel value array is determined, and the initial count value array is corrected based on the pixel value and the number of blank pixels on the right, to obtain a target count value array; corresponding blank pixel values are added to the beginning and end of the pixel value array based on the number of pixels of the number of blank pixels on the left and the number of blank pixels on the right, respectively, to obtain a target pixel value array; the local mask data is determined based on the target count value array and the target pixel value array.
[0073] The determination subunit is used to, if the pixel value is a preset first value, modify the last count value of the initial count value array based on the number of blank pixels on the right to obtain a first count value array, wherein the preset first value represents the background pixel; if the pixel value is a preset second value, add the number of blank pixels on the right as a count value to the initial count value array to obtain a second count value array, wherein the preset second value represents the foreground pixel.
[0074] The instance segmentation mask generation apparatus provided in this application, employing the instance segmentation mask generation method in the above embodiments, can solve the technical problem of instance segmentation mask generation. Compared with the prior art, the beneficial effects of the instance segmentation mask generation apparatus provided in this application are the same as those of the instance segmentation mask generation method provided in the above embodiments, and other technical features in the instance segmentation mask generation apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0075] This application provides an instance segmentation mask generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the instance segmentation mask generation method in Embodiment 1 above.
[0076] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an instance segmentation mask generation device suitable for implementing embodiments of this application. The instance segmentation mask generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablets, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The instance segmentation mask generation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0077] like Figure 5As shown, the instance segmentation mask generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the instance segmentation mask generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the instance segmentation mask generation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show instance segmentation mask generation devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0078] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0079] The instance segmentation mask generation device provided in this application, employing the instance segmentation mask generation method in the above embodiments, can solve the technical problem of instance segmentation mask generation. Compared with the prior art, the beneficial effects of the instance segmentation mask generation device provided in this application are the same as those of the instance segmentation mask generation method provided in the above embodiments, and other technical features in this instance segmentation mask generation device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0080] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0081] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0082] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the instance segmentation mask generation method in the above embodiments.
[0083] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0084] The aforementioned computer-readable storage medium may be included in the instance segmentation mask generation device; or it may exist independently and not assembled into the instance segmentation mask generation device.
[0085] The aforementioned computer-readable storage medium carries one or more programs. When the one or more programs are executed by the instance segmentation mask generation device, the instance segmentation mask generation device: in response to an instance segmentation mask generation instruction, obtains image information to be processed through the main process and writes the image information into a shared memory block; reads the corresponding local image region in the shared memory block through a subprocess, obtains local mask data, and returns the local mask data to the main process, so that the main process generates global mask data based on the local mask data.
[0086] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0088] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0089] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described instance segmentation mask generation method, thereby solving the technical problem of instance segmentation mask generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the instance segmentation mask generation method provided in the above embodiments, and will not be repeated here.
[0090] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the instance segmentation mask generation method described above.
[0091] The computer program product provided in this application can solve the technical problem of instance segmentation mask generation. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the instance segmentation mask generation method provided in the above embodiments, and will not be repeated here.
[0092] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
Claims
1. A method for generating an instance segmentation mask, characterized in that, The instance segmentation mask generation method includes: In response to the instance segmentation mask generation instruction, the image information to be processed is obtained through the main process and written into a shared memory block; The subprocess reads the corresponding local image region in the shared memory block to obtain local mask data, and returns the local mask data to the main process so that the main process can generate global mask data based on the local mask data. The step of reading the corresponding local image region in the shared memory block through a subprocess to obtain local mask data includes: Determine the subtask corresponding to the child process, wherein the subtask includes external rectangle information and access handle of the shared memory block; The child process reads the shared memory block corresponding to the access handle to obtain the local image region corresponding to the external rectangular box information. The image column containing the local image region is scanned column by column to obtain the initial mask data; The initial mask data is corrected to obtain local mask data; The step of correcting the initial mask data to obtain local mask data includes: Based on the image information, the overall image height and width are determined, and based on the outer rectangular frame information, the local width of the local image region and the horizontal coordinate of the preset position are determined. Based on the horizontal coordinate and the overall image height, calculate the number of blank pixels on the left side of the image column; based on the overall image height, the overall image width, the local width, and the horizontal coordinate, calculate the number of blank pixels on the right side of the image column. Based on the number of blank pixels on the left and the number of blank pixels on the right, the initial mask data is corrected to obtain local mask data.
2. The instance segmentation mask generation method as described in claim 1, characterized in that, The initial mask data includes a count value array and a pixel value array. The step of correcting the initial mask data based on the number of blank pixels on the left and the number of blank pixels on the right to obtain the local mask data includes: Based on the number of blank pixels on the left, the first count value of the count value array is corrected to obtain the initial count value array; Determine the pixel value corresponding to the last element in the pixel value array, and based on the pixel value and the number of blank pixels on the right, correct the initial count value array to obtain the target count value array; Based on the number of pixels of the left blank pixels and the right blank pixels respectively, add corresponding blank pixel values to the beginning and end of the pixel value array to obtain the target pixel value array; The local mask data is determined based on the target count value array and the target pixel value array.
3. The instance segmentation mask generation method as described in claim 2, characterized in that, The target count value array includes a first count value array and a second count value array. The step of correcting the initial count value array based on the pixel value and the number of blank pixels on the right to obtain the target count value array includes: If the pixel value is a preset first value, then based on the number of blank pixels on the right, the last count value of the initial count value array is corrected to obtain a first count value array, wherein the preset first value represents the background pixel; If the pixel value is a preset second value, then the number of blank pixels on the right is added as a count value to the initial count value array to obtain a second count value array, wherein the preset second value represents the foreground pixel.
4. The instance segmentation mask generation method as described in claim 1, characterized in that, The image information includes the instance target and the bounding box information corresponding to the instance target. Before the step of determining the subtask corresponding to the subprocess, the following steps are included: Through the main process, a corresponding subtask is created for each instance target in the image information; Create a subprocess pool and distribute the subtasks sequentially to the idle subprocesses in the subprocess pool.
5. An instance segmentation mask generation apparatus, characterized in that, The device includes: The writing module is used to respond to the instance segmentation mask generation instruction, obtain the image information to be processed through the main process, and write the image information into a shared memory block; The generation module is used to read the corresponding local image region in the shared memory block through a subprocess, obtain local mask data, and return the local mask data to the main process, so that the main process can generate global mask data based on the local mask data; The generation module includes: The scanning submodule is used to determine the subtask corresponding to the subprocess, wherein the subtask includes external bounding box information and access handle of the shared memory block; the subprocess reads the shared memory block corresponding to the access handle to obtain the local image region corresponding to the external bounding box information; the image column containing the local image region is scanned column by column to obtain initial mask data; the initial mask data is corrected to obtain local mask data; The scanning submodule includes: The correction unit is used to determine the overall image height and width based on the image information; determine the local width and the horizontal coordinate of a preset position of the local image region based on the outer rectangular frame information; calculate the number of blank pixels on the left side of the image column based on the horizontal coordinate and the overall image height; calculate the number of blank pixels on the right side of the image column based on the overall image height, the overall image width, the local width, and the horizontal coordinate; and correct the initial mask data based on the number of blank pixels on the left and the number of blank pixels on the right to obtain local mask data.
6. An instance segmentation mask generation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the instance segmentation mask generation method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the instance segmentation mask generation method as described in any one of claims 1 to 4.
8. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the instance segmentation mask generation method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Remote sensing image cloud detection method based on DeepLabV3 semantic segmentation framework
CN121353900A
Wide remote sensing image segmentation and sample automatic generation method based on map matching and forward and reverse projection
CN122089797A