Image segmentation method, object classification method, device, equipment, medium, product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NUCTECH CO LTD
- Filing Date
- 2026-07-03
- Publication Date
- 2026-08-07
AI Technical Summary
在一个示例中,矿石分选可以包括跳汰、重介、浮选等湿选方法,但是,该些方法普遍存在着能耗高、淡水消耗大、尾矿处理困难、排污量大等缺点
Smart Images

Figure CN122530593A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of computer, image processing, and ore sorting technologies, and more specifically, to an image segmentation method, an object classification method, an apparatus, a device, a medium, and a product. Background Technology
[0002] Ore sorting is a crucial step in mineral processing. In one example, ore sorting may include wet methods such as jigging, heavy media, and flotation. However, these methods generally suffer from drawbacks such as high energy consumption, large freshwater consumption, difficulty in tailings treatment, and significant pollution. In another example, with increasing environmental protection requirements and the growing scarcity of water resources, dry sorting methods, represented by photoelectric methods, are gaining importance, enabling ore sorting in areas with limited land and water resources.
[0003] With the development of pattern recognition and deep learning technologies, intelligent ore sorting methods based on image sources such as visible light and X-rays have been widely applied. These methods can acquire ore images through imaging equipment and then use computer vision models for automatic identification and sorting, thereby effectively reducing grinding energy consumption and tailings emissions.
[0004] However, in actual imaging, since the ores are usually stuck or overlapping on the conveyor belt, how to achieve accurate segmentation of the ores image is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the present disclosure provides an image segmentation method, an object classification method, an apparatus, a device, a medium, and a product.
[0006] According to one aspect of this disclosure, an image segmentation method is provided, comprising: acquiring an initial mask image of an image to be segmented, wherein the initial mask image includes initial masks for a plurality of objects to be segmented; performing morphological adjustments on the plurality of initial masks to form boundary indication information between the adjusted masks that satisfy preset position conditions, thereby obtaining a reference mask image; and segmenting the plurality of objects to be segmented in the image to be segmented based on the reference mask image to obtain an image segmentation result.
[0007] According to one aspect of this disclosure, an object classification method is provided, comprising: acquiring images of an object to be classified in at least two candidate modalities, wherein images of different candidate modalities reflect different physical properties of the object to be classified; performing image segmentation processing on images of reference modalities in at least two of the images using the image segmentation method to obtain image segmentation results; and inputting the images of at least two candidate modalities and the image segmentation results into a trained object classification model to obtain the category of each of the objects to be classified.
[0008] According to another aspect of this disclosure, an image segmentation apparatus is provided, comprising: a first acquisition module for acquiring an initial mask image of an image to be segmented, wherein the initial mask image includes initial masks for a plurality of objects to be segmented; a morphological adjustment module for morphologically adjusting the plurality of initial masks respectively, so that boundary indication information is formed between the adjusted masks that satisfy preset position conditions, thereby obtaining a reference mask image; and a first image segmentation module for segmenting the plurality of objects to be segmented in the image to be segmented according to the reference mask image, thereby obtaining an image segmentation result.
[0009] According to another aspect of this disclosure, an object classification apparatus is provided, comprising: a second acquisition module for acquiring images of an object to be classified in at least two candidate modalities, wherein images of different candidate modalities reflect different physical properties of the object to be classified; a second image segmentation module for performing image segmentation processing on images of reference modalities in at least two of the images using the image segmentation apparatus to obtain image segmentation results; and an object classification module for inputting the images of at least two candidate modalities and the image segmentation results into a trained object classification model to obtain the categories of each of the objects to be classified.
[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more instructions, wherein, when executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in this disclosure.
[0011] According to another aspect of this disclosure, a computer-readable storage medium is provided having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the methods described in this disclosure.
[0012] According to another aspect of this disclosure, a computer program product is provided, which includes computer-executable instructions that, when executed, are used to perform the methods described in this disclosure. Attached Figure Description
[0013] The above and other objects, features, and advantages of this disclosure will become clearer from the following description of embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0014] Figure 1 This illustration schematically shows a system architecture to which image segmentation methods and object classification methods can be applied according to embodiments of the present disclosure;
[0015] Figure 2A flowchart illustrating an image segmentation method according to an embodiment of the present disclosure is shown schematically.
[0016] Figure 3 This illustration schematically shows an example of a process of morphologically adjusting multiple initial masks to obtain a reference mask image according to an embodiment of the present disclosure;
[0017] Figure 4 This illustration schematically shows an example of a process for segmenting multiple objects to be segmented in an image to be segmented according to a reference mask image, based on an embodiment of the present disclosure, to obtain an image segmentation result.
[0018] Figure 5 An example schematic diagram of an image segmentation process according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A flowchart illustrating an object classification method according to an embodiment of the present disclosure is shown schematically;
[0020] Figure 7 The illustration shows an example diagram of the training process of an object classification model according to an embodiment of the present disclosure;
[0021] Figure 8 The illustration shows an example diagram of the training process of an object classification model according to another embodiment of the present disclosure;
[0022] Figure 9 This illustration schematically shows an example diagram of an object classification process according to an embodiment of the present disclosure;
[0023] Figure 10 A block diagram of an image segmentation apparatus according to an embodiment of the present disclosure is shown schematically;
[0024] Figure 11 A block diagram of an object classification apparatus according to an embodiment of the present disclosure is schematically shown; and
[0025] Figure 12 A block diagram schematically illustrates an electronic device suitable for implementing an image segmentation method and an object classification method according to embodiments of the present disclosure. Detailed Implementation
[0026] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0029] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0030] In the technical solution of this invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0031] In one example, to perform ore identification and sorting, the ore targets in the image first need to be segmented using a mask. However, in actual imaging, because the ores often stick together or overlap on the conveyor belt, a single segmentation model is difficult to effectively separate the sticky ores while ensuring the accuracy of the outer edges.
[0032] While some methods exist that can determine ore location through detection models or obtain accurate edges through threshold segmentation, there is still a lack of effective means to organically unify location information, segmentation clues, and edge accuracy. This is especially true in drop-type sorting equipment, where the time from image recognition to air-jet separation is extremely short, placing high demands on the computational speed of the segmentation algorithm.
[0033] In another example, due to the differences in the separability of different types and origins of ores in visible light and X-ray images, for instance, some ore concentrates and tailings look similar and are difficult to distinguish in visible light images; the composition of some ores is similar in the X-ray energy spectrum and cannot be distinguished by X-ray images.
[0034] Currently, for each new ore, technicians need to train and experiment with various image sources, such as visible light, high-energy X-rays, and low-energy X-rays, as well as the combinations thereof, to determine which data source can achieve the best sorting effect. This process is time-consuming and inefficient, which restricts the rapid iteration of ore sorting models and their deployment across different ore types.
[0035] Therefore, this disclosure provides an image segmentation method, object classification method, apparatus, device, medium, and product, which can be applied to the fields of computer, image processing, and ore sorting technology. The image segmentation method includes: acquiring an initial mask image of the image to be segmented, wherein the initial mask image includes initial masks for multiple objects to be segmented; performing morphological adjustments on the multiple initial masks to form boundary indication information between the adjusted masks that meet preset position conditions, thereby obtaining a reference mask image; and segmenting the multiple objects to be segmented in the image to be segmented based on the reference mask image to obtain an image segmentation result.
[0036] Figure 1 The illustration schematically depicts a system architecture to which image segmentation methods and object classification methods can be applied according to embodiments of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0037] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between different devices.
[0038] It should be noted that the image segmentation method and object classification method provided in the embodiments of this disclosure can generally be executed by the server 105. Accordingly, the image segmentation device and object classification device provided in the embodiments of this disclosure can generally be set in the server 105.
[0039] Alternatively, the image segmentation method and object classification method provided in the embodiments of this disclosure can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the image segmentation apparatus and object classification apparatus provided in the embodiments of this disclosure can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0040] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0041] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.
[0042] The foregoing section described the system architecture provided in this disclosure, which allows for the application of image segmentation and object classification methods. The following section will use... Figure 2 The image segmentation process of this disclosure will be further illustrated as an example.
[0043] Figure 2 A flowchart illustrating an image segmentation method according to an embodiment of the present disclosure is shown schematically.
[0044] like Figure 2 As shown, the image segmentation method 200 may include operations S210 to S230.
[0045] In operation S210, an initial mask image of the image to be segmented is obtained, wherein the initial mask image includes the initial masks of each of the multiple objects to be segmented.
[0046] In operation S220, the shapes of multiple initial masks are adjusted respectively so that boundary indication information is formed between the adjusted masks of the initial masks that meet the preset position conditions, and a reference mask image is obtained.
[0047] In operation S230, multiple objects to be segmented in the image to be segmented are segmented based on the reference mask image to obtain the image segmentation result.
[0048] The image to be segmented is the original input image from which object segmentation is to be performed. For example, taking ore as the object to be segmented, the image to be segmented can include high-energy X-ray images, low-energy X-ray images, visible light images, infrared images, etc. The object to be segmented is the smallest target unit in the image to be segmented that needs to be separated and identified. For example, the image to be segmented can include ore, metal, parts, fruit, cells, etc.
[0049] The initial mask image can be a binary or category-labeled image directly output by the initial segmentation model, representing the approximate region of each object to be segmented. Because the edges of each initial mask in the initial mask image are relatively coarse, it is difficult to separate clustered objects or provide precise edges for each object. An initial mask is the region corresponding to each independent connected component in the initial mask image; each initial mask can represent the approximate shape and location of an object to be segmented.
[0050] The method for obtaining the initial mask image can be configured according to actual business needs and is not limited here. For example, a deep learning model dedicated to single-class segmentation (such as UNet or DeepLab) can be used to predict the image to be segmented to obtain the initial mask image. Alternatively, an object detection model (such as YOLO or Faster R-CNN) can be used to predict the bounding box of each object to be segmented, and then a classic segmentation algorithm can be applied to each bounding box to segment the foreground within the box as the initial mask.
[0051] Morphological adjustment refers to mathematical morphological operations that modify the shape and size of the initial mask, expanding the object region in a controlled manner to create spatial interaction between adjacent but independent mask regions, thereby explicitly forming boundary cues. The specific method of morphological adjustment can be configured according to actual business needs and is not limited here. For example, a fixed size such as 3×3 or an adaptive structuring element based on the object to be segmented can be used to expand each initial mask; alternatively, an active contour model can be used to evolve each initial mask outwards until it meets the evolution edge of an adjacent initial mask.
[0052] Preset location conditions are rules used to filter which initial masks require shape adjustments. They can be used to filter out ores that are spatially close, adjacent, or overlapping. The specific content of the preset location conditions can be configured according to actual business needs and is not limited here. For example, a preset location condition can be that the minimum distance between initial masks is less than a preset distance threshold; alternatively, a preset location condition can also be that the intersection-over-union (IoU) ratio between the bounding boxes of two initial masks is greater than a preset IoU threshold, etc.
[0053] An adjusted mask is a new mask region obtained by adjusting the shape of an initial mask. For example, an adjusted mask can be a mask region with a slightly larger area after the initial mask of a piece of ore is expanded.
[0054] Boundary indication information is information formed by the interaction of multiple adjusted masks to indicate potential dividing lines between adhered objects. For example, the dividing lines formed between adjusted masks are obtained after morphological adjustment of an initial mask that satisfies preset positional conditions.
[0055] The reference mask image is an image that integrates all the adjusted masks and the boundary indication information they form. For example, the reference mask image can be an image of the same size as the original image to be segmented, where each adjusted mask is assigned a different number. The image segmentation result is the final output mask map that accurately separates each object to be segmented.
[0056] After obtaining the reference mask image, thresholding can be performed on the original image to be segmented to obtain a binary image, which provides accurate outer edges of the ore. Then, using the boundary indication information in the reference mask image, the contiguous regions in the binary image are cut apart to obtain an optimized binary image. Connected components are calculated on the optimized binary image to give each object to be segmented an independent number. Based on this, the connected component results are matched with the object positions and numbers obtained by the detection model. Using the detection results as the standard, incorrectly segmented fragments are removed to obtain the final accurate image segmentation result.
[0057] In the embodiments of this disclosure, by actively adjusting the shape of an initial mask that meets preset positional conditions, clear boundary indication information is generated at the boundaries of the objects to be segmented through artificial collision. This helps to solve the problem of difficult separation of objects in complex adhesion scenarios, transforming passive adhesion recognition into active boundary creation. Based on this, multiple objects to be segmented in the image to be segmented are segmented according to the reference mask image, ensuring that the outer edge of each object matches the true edge of the original image to be segmented. This enables a coarse-to-fine segmentation process with lower computational complexity and higher robustness, improving the accuracy of the image segmentation results.
[0058] According to embodiments of this disclosure, the object to be segmented may include ore, the preset position conditions may include the ore corresponding to the initial mask being stuck together or overlapping in the image to be segmented, and the boundary indication information may be used to separate the stuck or overlapping ore.
[0059] After performing shape adjustments, the initial mask image can be analyzed to identify which connected components represent ore and determine whether there are spatial relationships of adhesion or overlap between these ore components, thereby determining which initial masks need shape adjustments. The specific determination method can be configured according to actual business needs and is not limited here.
[0060] For example, the results of the detection model can be used to make a judgment. If an initial mask connected region contains two or more ore bounding boxes predicted by the detection model, then the ore within that connected region is determined to be connected or overlapping. Alternatively, shape analysis can be performed on each initial mask connected region, such as calculating its area, perimeter, convex hull area, and aspect ratio of the minimum bounding rectangle.
[0061] "Adhesion" refers to two or more pieces of ore being in physical contact with each other, but due to insufficient imaging dimension or segmentation accuracy, forming a connected region in the initial mask, lacking clear boundaries. For example, if two pieces of ore are placed side by side on a conveyor belt without any gaps, their initial mask will show a figure-eight-shaped connected region, thus satisfying "adhesion".
[0062] "Phase overlap" refers to the stacking of ores in the projection direction, causing one ore in the image to be segmented to partially or completely obscure another. For example, if a small ore is stacked on the upper edge of a large ore, and viewed from above in an X-ray image, the area of the small ore partially overlaps with the area of the large ore, then "phase overlap" is satisfied.
[0063] In the embodiments of this disclosure, by introducing the adhesion or overlap of the ore corresponding to the initial mask as a necessary preset position condition to trigger subsequent operations, instead of blindly adjusting all masks, the most difficult ore to solve is accurately identified, making the generation of boundary indication information highly purposeful. It can obtain an accurate mask for a single ore with extremely high robustness, reducing misselection or omission caused by failure of segmentation adhesion, and improving resource recovery rate and sorting purity.
[0064] According to embodiments of this disclosure, the adjusted mask is obtained by: determining shape adjustment parameters based on attribute information obtained from multiple initial masks, wherein the shape adjustment parameters indicate the shape adjustment range required to form boundary indication information between the adjusted masks; and adjusting the shape of the initial masks according to the shape adjustment parameters to obtain the adjusted mask.
[0065] Attribute information can be quantifiable feature data extracted by analyzing the spatial relationship, geometry, or association with the original image of two or more initial masks. For example, attribute information can be the average area, equivalent diameter, maximum length, shortest distance between the edges of two initial masks, overall compactness or concavity of adhering ore clusters, etc.
[0066] Shape adjustment parameters are numerical values that specifically control the strength, range, or degree of shape adjustment operations. Shape adjustment parameters can be used to determine how much the adjusted mask can expand, and whether and how they meet to form a boundary. Shape adjustment parameters can be configured according to actual business needs and are not limited here. For example, shape adjustment parameters can include the size of the structuring element, such as deciding whether to use a 3×3 or 7×7 convolution kernel for dilation; shape adjustment parameters can also include the number of dilation iterations, such as dilating once or three times; shape adjustment parameters can also include the expanded pixel distance, such as expanding each mask outward by 2 pixels or 5 pixels.
[0067] The morphological adjustment range must be neither too small, preventing the masks of the adhered ores from making contact and forming a boundary line, nor too large, leading to incorrect mask merging. For example, if the shortest distance between the edges of two initial ores masks is 4 pixels, and the morphological adjustment uses a disk structure element with a radius of 2 pixels that expands once, with each mask expanding outward by 2 pixels, then the two will just touch, forming a boundary indicator line 1 pixel wide. In this case, "radius 2 pixels, expansion once" can be used as the morphological adjustment range for the target.
[0068] The method for determining the shape adjustment parameters can be configured according to actual business needs and is not limited here. For example, the area or equivalent diameter of each of multiple initial masks stuck together can be calculated, and according to preset rules, such as "when the average equivalent diameter is greater than 50 pixels, a 7×7 expansion kernel is selected; otherwise, a 3×3 expansion kernel is selected", the size attributes are mapped to specific structural element size parameters.
[0069] Alternatively, the shortest Euclidean distance between the two closest initial mask edges can be calculated, and the shape adjustment parameters can be determined directly using a preset formula or the number of pixels for morphological expansion. Another option is to set preset initial adjustment parameters, perform shape adjustment, and then check if the adjusted masks have made contact to form boundary indication information; if not, the parameters are incremented and adjusted again until contact is detected for the first time, and the parameters at this point are recorded as the final selected shape adjustment parameters.
[0070] In one embodiment, the reference mask image can be obtained as shown in the following formula (1).
[0071] (1)
[0072] in, Characterizes the reference mask image. The k-th adjusted mask is represented by k = 1, 2, ..., q, where q represents the number of the initial masks.
[0073] In the embodiments of this disclosure, attribute information obtained from multiple initial masks is introduced as a decision-making basis, and morphological adjustment parameters are purposefully determined based on this attribute information, transforming the fuzzy target into calculable and executable steps. On this basis, the initial masks are morphologically adjusted based on these morphological adjustment parameters in a highly adaptive manner, ensuring the consistency and reliability of boundary indication information generation under different adhesion conditions. This enhances the robustness and generalization ability of the entire segmentation process while maintaining segmentation accuracy.
[0074] According to embodiments of this disclosure, determining shape adjustment parameters based on attribute information obtained from multiple initial masks may include at least one of the following: determining shape adjustment parameters based on first statistical information obtained based on the area ratio of each initial mask relative to the initial mask image; determining shape adjustment parameters based on second statistical information obtained based on the relative distance between each initial mask.
[0075] Because large ore blocks have a relatively large image area, slight dilation has little effect on their relative size, and it's easier to make the mask contact them through moderate dilation. Smaller ore blocks, however, require more careful adjustment to avoid excessive morphological adjustment relative to their size, which could lead to mask distortion or the absorption of neighboring objects. Therefore, the overall or local block size of the ore can be mapped to the adjustment range.
[0076] The first statistical information is a statistical measure reflecting the size characteristics of the ore, calculated by statistically analyzing the number of pixels occupied by each independent initial ore mask region in the initial mask image and comparing it with the total number of pixels in the entire image or a specific reference region. For example, the first statistical information can be determined based on average area proportion, average equivalent diameter, and maximum area mask proportion.
[0077] The average area percentage is calculated by taking the average of the proportions of the areas of all connected components of the initial masks to the total image area. For example, if an image contains 20 mineral masks with respective percentages of 0.3%, 0.5%, 0.2%, etc., the calculated average of 0.35% is the first statistical information. The average equivalent diameter is calculated by converting the area of each initial mask into the diameter of a circle with equal area, and then finding the average or median of all diameters. For example, the equivalent diameters of three bonded minerals are 45, 52, and 48 pixels, respectively, and the average of 48 pixels is the first statistical information. The maximum area mask percentage is calculated by taking only the initial mask with the largest area and calculating its proportion of the image area.
[0078] After obtaining the initial statistical information, multiple size threshold ranges can be set, each corresponding to a set of preset shape adjustment parameters. For example, when the mean diameter is <30 pixels, a 2-pixel expansion radius is selected; when it is 30-60 pixels, a 3-pixel expansion radius is selected; and when it is >60 pixels, a 5-pixel expansion radius is selected. Alternatively, the initial statistical information can be used as the independent variable to design a continuous function to calculate the shape adjustment parameters, thereby producing smooth, stepless parameter changes.
[0079] Since the ores are close together, only slight expansion is needed to create a boundary; if the ores have some spacing, a greater force is required to adjust their shape. Based on this, the amount of space that needs to be filled to bring them into contact can be determined directly based on the current degree of spatial separation of the ores.
[0080] The second statistical information is a statistical measure reflecting the spacing characteristics of the ore particles, obtained by calculating the pixel distances between multiple initial mask edges that meet preset position conditions and performing statistical processing on these distance values. For example, the second statistical information may include the nearest edge distance, the average edge distance, and a distance distribution histogram.
[0081] Nearest edge distance is calculated by taking the shortest Euclidean distance between boundary pixels of multiple adhered initial masks. For example, if two minerals are 3 pixels apart at their edges in an image, then 3 represents this information. Average edge distance is calculated as the average edge distance between all neighboring masks for complex adhered objects. Distance distribution histograms are used to count the number of mask pairs within different distance intervals, using the most frequent distance values or the overall distribution as secondary statistical information.
[0082] After obtaining the second statistical information, taking the nearest edge distance as an example, half of the nearest edge distance can be determined as the shape adjustment parameter to ensure that after each mask expands `r` towards the other, they can at least meet at the midpoint. Alternatively, taking the distance distribution histogram as an example, if the histogram is unimodal and the peak width is narrow, the peak distance can be used to calculate the parameter; if the distance distribution is dispersed, the larger distance value with the highest frequency can be selected to calculate the parameter, to ensure that those ore pairs that are farthest apart in the entire cluster but still need to be separated can be separated.
[0083] In the embodiments of this disclosure, by introducing first statistical information to determine parameters based on area ratio, the inherent dimensional characteristics of the ore can be captured. This allows for intelligent matching of the adjustment range with the ore size, effectively avoiding the risk of boundary distortion in small ore due to over-adjustment, and ensuring that large ore clusters have sufficient force to form boundaries. By introducing second statistical information to determine parameters based on relative distance, the severity of adhesion can be captured, enabling shape adjustment to accurately compensate for different gaps between ore fragments. This improves the automation and precision of shape adjustment, allowing it to handle complex ore adhesion scenarios with varying sizes and spacing, and enhancing the accuracy and consistency of the entire ore segmentation process.
[0084] According to embodiments of this disclosure, morphological adjustments may include dilation. Dilation is a mathematical morphological operation whose basic effect is to expand the foreground region in a binary image outwards. In this embodiment, dilation is used to enlarge the initial mask of each ore, so that masks that were originally close to each other but not touching meet through expansion, thereby constructing a dividing line.
[0085] Shape adjustment parameters can include the size of the convolution kernel used for dilation. The size of the convolution kernel determines the number of pixels by which the foreground region expands outward during each dilation. The size can be in pixels, and the convolution kernel can be the side length of a square, the diameter of a circle, or the span of a custom shape, etc., without limitation.
[0086] Based on the shape adjustment parameters, the initial mask is shaped to obtain an adjusted mask. This can include the following operations: sliding through the initial mask with a convolution kernel of a certain size to obtain the adjusted mask.
[0087] Sliding traversal refers to aligning the origin of the convolution kernel sequentially with each pixel in the initial mask image according to a determined kernel size, checking whether there are foreground pixels in the neighborhood covered by the kernel. If there are foreground pixels, the position is set as the foreground, so that the mask boundary expands outward. The expansion magnitude is positively correlated with the kernel size, thereby achieving dilation of the initial mask.
[0088] The specific method of sliding traversal can be configured according to actual business needs and is not limited here. For example, an image processing library can be called, passing in the initial mask image, a rectangular or elliptical structuring element of a specified size, and the number of iterations to complete the sliding traversal calculation. Alternatively, for larger rectangular convolution kernels, a row-column separation acceleration strategy can be adopted. First, the initial mask is dilated one-dimensionally in the horizontal direction, and then the result is dilated one-dimensionally in the vertical direction to reduce computational complexity.
[0089] In one embodiment, the initial mask is traversed by a convolution kernel of this size to obtain an adjusted mask as shown in the following formula (2).
[0090] (2)
[0091] in, Representing the k-th initial mask, Characterize the convolution kernel, Characterizes morphological expansion operators.
[0092] In the embodiments of this disclosure, dilation is used as the method for morphological adjustment, leveraging its parallelizable and vectorizable characteristics to meet the real-time requirements of image segmentation. By using a convolution kernel of this size to slide through the initial mask, this pixel-level traversal process ensures uniform expansion in all boundary directions of the mask, achieving robust isotropic expansion.
[0093] The following will utilize Figure 3 The process of obtaining an adjusted mask by morphologically adjusting the initial mask is further explained.
[0094] Figure 3 The illustration shows an example of a process of morphologically adjusting multiple initial masks to obtain a reference mask image according to an embodiment of the present disclosure.
[0095] like Figure 3As shown, in embodiment 300 where the initial mask is morphologically adjusted to obtain the adjusted mask, the initial mask image 310 including initial mask 311, initial mask 312 and initial mask 313 will be used as an example for explanation.
[0096] In one embodiment, the area proportions 321 of initial masks 311, 312, and 313 relative to the initial mask image 310 can be determined respectively, and first statistical information 331 can be determined based on the area proportions 321 of each initial mask. Based on this, shape adjustment parameters 340 can be determined according to the first statistical information 331.
[0097] In another embodiment, relative distances 322 between initial masks 311 and 312, between initial masks 311 and 313, and between initial masks 312 and 313 can be determined respectively, and second statistical information 332 can be determined based on each relative distance 322. Based on this, shape adjustment parameters 340 can be determined according to the second statistical information 332.
[0098] After obtaining the shape adjustment parameters 340, the initial masks can be shaped according to the shape adjustment parameters 340 to obtain a reference mask image 350. For example, the initial mask 311 can be shaped according to the shape adjustment parameters 340 to obtain an adjusted mask 351; the initial mask 312 can be shaped according to the shape adjustment parameters 340 to obtain an adjusted mask 352; and the initial mask 313 can be shaped according to the shape adjustment parameters 340 to obtain an adjusted mask 353.
[0099] Before the shape adjustment, there was no clear dividing line between the initial mask 311 corresponding to the adjusted mask 351 and the initial mask 313 corresponding to the adjusted mask 353. However, after the shape adjustment, boundary indication information 360 is formed between the adjusted mask 351 and the adjusted mask 353 to indicate the potential dividing line between the adhered objects.
[0100] According to an embodiment of this disclosure, operation S230 may include the following operations: using boundary indication information, adjusting the binary image obtained by binarization of the image to be segmented to obtain an adjusted image; using the position indication information obtained by target detection of the image to be segmented as a reference, correcting the adjusted image to obtain an image segmentation result, wherein the position indication information represents the position of the object to be segmented in the image to be segmented.
[0101] For an image to be segmented, a thresholding segmentation algorithm can be used to convert it into a binary image containing only the foreground (such as ore) and the background. For example, the image to be segmented can be processed based on an empirical threshold or Otsu's thresholding method to obtain a binary image. For instance, using Otsu's thresholding method, pixels with gray values greater than the threshold (background) in the image to be segmented can be set to 0, and pixels with gray values less than the threshold can be set to 1, thus obtaining a binary image.
[0102] In one embodiment, the specific method for obtaining a binary image can be shown in the following formula (3).
[0103] (3)
[0104] in, Representing a binary image, Characterizing threshold, Characterizes grayscale values.
[0105] After obtaining the binary image, pixels located at boundary indicator information can be modified to background values to achieve the effect of cutting apart the adhering blocks along the segmentation line. This allows the adjusted image to retain both the precise outer edge of the threshold segmentation and the accurate separation line from the morphological adjustment. The specific adjustment method can be configured according to actual business needs and is not limited here. For example, a pixel-by-pixel check can be performed on the binary image; if a pixel is marked as a segmentation line pixel in the reference mask image, then that position can be forcibly set to 0 in the image.
[0106] For the image to be segmented, a location indication information representing the spatial location of each object can be obtained from the image using an object detection algorithm. For example, the location indication information can be in the form of bounding box coordinates, center point coordinates plus dimensions, or instance masks, etc.
[0107] After obtaining the location indication information, the connected components in the adjusted image can be matched with the number and location of the location indication information as a reference to correct any oversegmentation or undersegmentation problems that may exist in the adjusted image, and to remove or merge redundant regions that are inconsistent with the location indication information.
[0108] For example, if the image has 23 connected components due to noise in the segmentation lines and the detection model only outputs 20 bounding boxes, the correction process can be to match these 23 connected components with the 20 bounding boxes, discard the 3 unmatched and extra connected components, and finally output a mask of exactly 20 ore instances.
[0109] The specific method of correction can be configured according to actual business needs and is not limited here. For example, the connected components of the adjusted image can be calculated first to obtain several candidate ore regions; then the centroid or bounding box of each candidate connected component can be calculated and matched with all bounding boxes of the location indication information based on the intersection-union ratio or center distance; those connected components that can be successfully matched with detection boxes are retained, and the number of matched detection boxes is the final number; connected components that fail to match any detection boxes are discarded.
[0110] Alternatively, in addition to using location indication information, confidence scores of each location indication information output by the detection model can also be utilized. Each candidate connected region in the adjusted image is assigned a comprehensive score based on its area, shape regularity, and degree of matching with the detection box; then, non-maximum suppression is performed on all candidate regions, with connected regions that have high confidence and a high degree of matching with the detection box being preferentially retained, while low-scoring redundant connected regions are suppressed.
[0111] In the embodiments of this disclosure, by utilizing boundary indication information, the binary image obtained after binarization of the image to be segmented is adjusted, cleverly fusing the artificially created segmentation lines with the threshold segmentation results that reflect the real physical boundaries, so that the segmentation results are separated at the points of adhesion. Based on this, by introducing the position indication information obtained from target detection as a benchmark to correct the adjusted image, the process errors that may be introduced during the aforementioned dilation process are effectively suppressed, ensuring the accuracy of the number of ore instances and the reliability of their locations in the results.
[0112] According to embodiments of this disclosure, the object to be segmented in a binary image can be displayed using a first grayscale value, and other regions in the binary image besides the object to be segmented can be displayed using a second grayscale value. The first grayscale value is the pixel brightness value used to represent the foreground object after binarization. The second grayscale value is the pixel brightness value used to represent the background and other non-target regions in the binary image, excluding the foreground object. For example, the first grayscale value can be 255, displayed as pure white, and the second grayscale value can be 0, displayed as pure black.
[0113] Using boundary indication information, the binary image obtained by binarization of the image to be segmented is adjusted to obtain an adjusted image. This can include the following operations: according to the boundary indication information, the value of the pixel position in the binary image corresponding to the boundary indication information is set to the second gray value, so as to form a segmentation line in the overlapping area between the adjusted masks, thereby obtaining the adjusted image.
[0114] The pixel positions corresponding to the boundary indication information are the pixel coordinates that are explicitly marked as dividing lines in the reference mask image generated in the previous step. These pixel positions are not the actual edges of the ore, but rather ideal cutting paths created through morphological adjustments to separate adhering ore.
[0115] An overlapping region is a portion of the pixels where two or more morphologically adjusted masks spatially overlap. For example, at the boundary between an expanded mask for ore A and an expanded mask for ore B, there is an area approximately several pixels wide that is simultaneously covered by both masks; this portion of pixels is the overlapping region.
[0116] A dividing line is a background-colored slit cut into the foreground region of minerals that was originally connected in the image by setting the pixel position corresponding to the boundary indicator information to the background gray value. This divides the minerals that were originally connected as a single connected region in the binary image into two independent connected regions.
[0117] In one embodiment, the specific method for forming the dividing line can be shown in the following formula (4).
[0118] (4)
[0119] in, The representation is the pixel position in a binary image that corresponds to the boundary indication information. It represents the pixel position in a non-binary image that corresponds to the boundary indication information.
[0120] In the embodiments of this disclosure, by setting the value of the pixel position corresponding to the boundary indication information to the second gray value, the artificial dividing line created in the previous step is cleverly introduced into the binary image as a background, thereby making the overlapping part of the adjusted mask form a dividing line, and converting the boundary indication information into an executable pixel-level cutting operation, thus achieving accurate separation of the binary image of the adhered ore.
[0121] According to embodiments of this disclosure, the image segmentation result is obtained by correcting an adjusted image based on the location indication information obtained by object detection in the image to be segmented. This can include the following operations: performing connected component analysis on the adjusted image to distinguish the boundaries of each object to be segmented, thereby obtaining an object region image, wherein the object region image includes candidate regions for each of the multiple objects to be segmented; and matching each candidate region based on the location indication information to obtain the image segmentation result.
[0122] For the adjusted image, all independent foreground connected regions in the adjusted image can be identified by pixel-by-pixel scanning or contour tracing, and each connected region can be assigned a unique number, thereby transforming the original image with the same foreground gray value into an object region image with an independent digital identifier for each ore.
[0123] Connected component analysis refers to the process of detecting the spatial connectivity of foreground pixels in an adjusted image, grouping connected foreground pixels into a single region, and assigning a unique identifier to each region. For example, taking an adjusted image with three white regions, connected component analysis scans all pixels in the image, labeling all pixels of the first white patch as 1, the second as 2, the third as 3, and the background as 0. These three patches constitute three connected components.
[0124] It should be noted that, due to different definitions of connectivity, connected component analysis includes 4-connectivity (only adjacent elements in the top, bottom, left, and right directions) and 8-connectivity (adjacent elements in the top, bottom, left, right, and diagonal directions). Taking ore segmentation as an example, the appropriate method can be selected based on the shape of the ore to adapt to irregular ore contours.
[0125] The object region image is the output image after connected component analysis, where each connected component is assigned a unique numerical number, representing a candidate region. Each uniquely identified independent connected component in the object region image constitutes a candidate ore instance region.
[0126] After obtaining the object region image, the candidate regions generated by connected component analysis can be aligned and matched with the position information output by the detection model. Candidate regions that can successfully match the detection box can be confirmed as real objects and retained. Candidate regions that cannot match any detection box can be determined as invalid and removed.
[0127] The specific method of correction can be configured according to actual business needs and is not limited here. For example, the bounding box of each candidate region can be calculated, and the intersection-union ratio (IoU) can be calculated one by one with each location indication information. If an IoU threshold is set, the candidate region is only retained when the highest IoU of the candidate connected component exceeds this threshold and the detection box also successfully matches. Alternatively, in addition to matching the location of the detection box, a confidence scoring mechanism can be introduced. For example, a comprehensive quality score can be calculated for each candidate connected component, which is composed of weights such as its IoU with the detection box, the consistency between the area of the candidate connected component and the area of the detection box, and the shape regularity of the candidate connected component; then, non-maximum suppression is applied to all candidates, and only instances with high confidence and high matching with the detection box are retained.
[0128] In one embodiment, the image segmentation result can be obtained as shown in the following formula (5).
[0129] (5)
[0130] in, Characterize the image segmentation results, Represent the k-th candidate region.
[0131] In the embodiments of this disclosure, by performing connected component analysis on the adjusted image, the previously scattered foreground pixels are transformed into candidate regions with independent identifiers, enabling each potential object to acquire an operable individual identity. Based on this, by matching each candidate region using location indication information as a benchmark, the location information output by the detection model, which has high reliability in determining presence and location, can be used as a reference to filter candidate regions generated solely through image binarization and morphological adjustments. This fundamentally solves the problems of over-segmentation noise and redundant fragmentation that may be introduced by threshold segmentation and morphological adjustments, resulting in image segmentation results that combine the edge accuracy of threshold segmentation, the separation effectiveness of morphological adjustments, and the high reliability of the detection model in terms of instance number and spatial location.
[0132] According to embodiments of this disclosure, matching each candidate region with location indication information to obtain an image segmentation result may include the following operations: matching the candidate region with the location indication information to obtain a matching result; deleting redundant candidate regions in response to the matching result indicating that there are candidate regions that do not match any location indication information; and obtaining an image segmentation result based on the candidate regions corresponding to each location indication information.
[0133] After obtaining candidate regions, spatial correspondence can be matched one by one with location indication information. That is, by quantifying the degree of similarity between the two in terms of spatial location, coverage area, etc., it can be determined which candidate region belongs to which object to be segmented by which detection box. For example, all location indication information can be traversed one by one for each candidate region, and spatial similarity measures such as intersection-union ratio and center distance can be calculated between them. The pairing relationship can be determined according to the preset matching criteria.
[0134] The specific matching method can be configured according to actual business needs and is not limited here. For example, the minimum bounding rectangle of each candidate region can be calculated, and the intersection-union ratio (IoU) can be calculated with all location indication information one by one; starting from the highest IoU pair, the pairing is confirmed sequentially, and each detection box and candidate region participate in the pairing at most once. Alternatively, instead of relying solely on IoU or center distance, multiple similarity measures can be calculated comprehensively, such as area similarity, shape similarity, positional similarity, and overlap based on the ratio between the area of the candidate region and the area of the detection box. These measures are weighted and fused into a comprehensive matching score, and the pairings that exceed the threshold are confirmed as matches.
[0135] For example, taking a candidate region A1 as an example, the centroid coordinates of the candidate region A1 can be calculated, and the Euclidean distance between the candidate region A1 and the center point of each location indication information can be calculated one by one. If it is found that the candidate region A1 is closest to the center of location indication information B1 and has the highest intersection-union ratio, then it can be determined that the candidate region A1 and location indication information B1 are successfully matched.
[0136] The matching results can include successfully matched candidate regions and their corresponding location indicators, a list of candidate regions that failed to match any location indicators, and a list of location indicators that may not have matched at all. Redundant candidate regions are those that cannot establish a valid correspondence with any location indicator in the matching results. These regions are usually spurious fragments introduced by oversegmentation due to thresholding, image noise, or morphological adjustments, and do not represent the real object. If a candidate region is marked as not matching any location indicator, a deletion action can be triggered, removing these redundant candidate regions from the object region image.
[0137] After redundant candidate regions are cleared, a set of candidate regions that clearly correspond to the location indication information in the matching results can be selected. These verified candidate regions are then reorganized according to the numbering order of the detection boxes or the spatial order to generate the final image segmentation result. Each retained candidate region is then confirmed as a real object instance to be segmented.
[0138] In the embodiments of this disclosure, a matching mechanism based on spatial correspondence is established by matching candidate regions with location indication information. This enables cross-validation of candidate results that rely solely on grayscale and morphology with high-level semantic locations from the detection network. Furthermore, in response to the matching results indicating that a candidate region does not match any location indication information, redundant candidate regions are deleted. This effectively filters out false fragments and over-segmentation noise introduced by threshold fluctuations and morphological adjustments, while completely preserving all real instances corresponding to the location indication information. Since the detection model ensures the reliability of the number and location of object instances, and segmentation and morphological adjustments ensure the accuracy of the mask edges for each object, the accuracy of the image segmentation results is improved.
[0139] The following will utilize Figure 4 The process of segmenting multiple objects in an image to be segmented based on an obtained reference mask image to obtain image segmentation results is further explained.
[0140] Figure 4 The illustration shows an example of a process for segmenting multiple objects in an image to be segmented according to a reference mask image, based on an embodiment of the present disclosure, to obtain an image segmentation result.
[0141] like Figure 4 As shown, in embodiment 400 for obtaining image segmentation results, the image 410 to be segmented can be binarized to obtain a binary image 420. The object to be segmented in the binary image 420 is displayed with a first grayscale value, and other regions in the binary image other than the object to be segmented are displayed with a second grayscale value.
[0142] Based on the boundary indication information 434 in the reference mask image 430, the value of the pixel position corresponding to the boundary indication information 434 in the binary image 420 is set to the second gray value, so as to form a dividing line 444 in the overlapping area between the adjusted mask 431 and the adjusted mask 433, and thus obtain the adjusted image.
[0143] Connectivity analysis is performed on the adjusted image to distinguish the boundaries of each object to be segmented, resulting in object region image 440. For example, object region image 440 may include candidate region 441 for object 1 to be segmented, candidate region 442 for object 2 to be segmented, and candidate region 443 for object 3 to be segmented. Candidate regions 441 and 443 can be distinguished by a dividing line 444.
[0144] Target detection is performed on the image to be segmented to obtain position indication information 450. Based on the position indication information 450, each candidate region is matched to obtain the image segmentation result 460.
[0145] The above sections have provided illustrative examples of obtaining an adjusted mask by morphologically adjusting an initial mask, and segmenting multiple objects in the image to be segmented based on the obtained reference mask image to obtain the image segmentation result. The following section will utilize... Figure 5 The overall image segmentation process is explained.
[0146] Figure 5 The illustration shows an example schematic diagram of an image segmentation process according to an embodiment of the present disclosure.
[0147] like Figure 5 As shown, in the image segmentation embodiment 500, the detection model M1 can be used to perform target detection on the image 510 to be segmented, and obtain the position indication information 520 of each object to be segmented. The segmentation model M2 is used to perform target segmentation on the image 510 to be segmented, and an initial mask image is obtained. The shapes of multiple initial masks in the initial mask image are adjusted so that boundary indication information is formed between the adjusted masks that meet the preset position conditions, and a reference mask image 530 is obtained.
[0148] Based on this, the reference mask image 530 can be thresholded using the position indication information 520 as a reference to obtain the image segmentation result 540.
[0149] The above are merely exemplary embodiments, but are not limited thereto. Other image segmentation methods known in the art may also be included, as long as they can improve the accuracy of the image segmentation results.
[0150] The image segmentation process provided in this disclosure has been described above. The following will use... Figure 6As an example, the object classification process of this disclosure is further explained.
[0151] Figure 6 A flowchart illustrating an object classification method according to an embodiment of the present disclosure is shown schematically.
[0152] like Figure 6 As shown, the object classification method 600 may include operations S610 to S630.
[0153] In operation S610, images of the object to be classified in at least two candidate modalities are acquired, wherein the images of different candidate modalities reflect different physical properties of the object to be classified.
[0154] In operation S620, using the image segmentation method, image segmentation processing is performed on the reference modality image in at least two images to obtain the image segmentation result.
[0155] In operation S630, images and image segmentation results from at least two candidate modalities are input into a trained object classification model to obtain the category of each object to be classified.
[0156] The object to be classified is the smallest target unit that needs to be classified after image segmentation. For example, in a mineral classification scenario, the object to be classified can be a single piece of ore whose precise mask has been extracted using segmentation methods. For the object to be classified, an ore sorting device equipped with multiple imaging sensors can be used to simultaneously or nearly simultaneously acquire images of the same batch of ore in at least two candidate modalities. It should be noted that the images in at least two candidate modalities need to be spatially registered and aligned so that the position of the same piece of ore corresponds in the images of different modalities.
[0157] Candidate modalities are imaging methods available for selection, based on different physical principles or energy ranges. Each modality's image reflects a specific physical property of the object to be classified. Physical properties refer to the different physical characteristics of the object captured by images from different candidate modalities. Due to differences in imaging principles, different modalities emphasize different features of the same ore.
[0158] A reference mode is a specific mode selected from multiple candidate modes for performing an image segmentation task. In one embodiment, taking an ore as an example, the candidate modes may include a visible light imaging mode and an X-ray imaging mode. Since the contrast between the ore and the background is high in X-ray transmission images, which is beneficial for obtaining accurate outer edges of the ore, the reference mode can be the X-ray imaging mode.
[0159] After obtaining the image segmentation results, the images and segmentation results from at least two candidate modalities can be input into a trained object classification model to obtain the category of each object to be classified. The trained object classification model is a neural network model that has completed the training process and is able to output the category label of each object to be classified based on the input images from at least two candidate modalities.
[0160] In the embodiments of this disclosure, by acquiring images of the object to be classified in at least two candidate modalities, the dimensionality of information is expanded at the physical level. This allows for the capture of both the surface optical features and the internal density features of the object, fundamentally solving the problem of insufficient object discrimination power of a single modality. By utilizing image segmentation methods to perform image segmentation processing on the reference modal images in at least two images, image segmentation results that can accurately indicate the boundaries of each object to be classified are obtained. Based on this, by inputting the images in at least two candidate modalities and the image segmentation results into a trained object classification model, each independent object to be classified obtains a comprehensive representation that integrates multimodal physical attributes, improving the efficiency and accuracy of object classification.
[0161] According to embodiments of this disclosure, the candidate mode may include at least two of the following: a visible light imaging mode, an X-ray imaging mode of a first energy level, and an X-ray imaging mode of a second energy level, wherein the energy range corresponding to the first energy level is different from the energy range corresponding to the second energy level.
[0162] Visible light imaging mode refers to the method of imaging ores using electromagnetic radiation in the visible light band (such as wavelengths of approximately 380 nm to 780 nm). Visible light imaging mode can capture the reflectivity of the ore surface to ambient light or active illumination, and can reflect the physical properties of the ore, such as surface color, texture, luster, and visible gangue distribution.
[0163] For example, an industrial RGB camera can be mounted above the sorting machine conveyor belt to capture high-resolution color images of the ore under white LED lighting. In the image, white quartz veins, black magnetite, and copper-green oxide minerals all exhibit different color and texture characteristics.
[0164] The first-energy-level X-ray imaging mode and the second-energy-level X-ray imaging mode refer to the methods of transmission imaging of ores using X-rays in different specific energy ranges. Because substances of different densities and atomic numbers absorb X-rays to varying degrees when they penetrate the ore, grayscale contrast images are formed. It should be noted that the energy range corresponding to the first energy level is different from that corresponding to the second energy level, so that the two modes can provide complementary physical information for ore classification.
[0165] For example, the first energy level could be high-energy X-rays with a tube voltage of 160 kVp and an average photon energy of approximately 100 keV. Due to its higher average photon energy, it has stronger penetrating power. Therefore, the absorption difference of this high-energy X-ray for heavy elements is different compared to that of the low-energy X-ray, making it more able to penetrate thicker ore blocks and provide information on their internal structure. The second energy level could be low-energy X-rays with a tube voltage of 60 kVp and an average photon energy of approximately 40 keV. Due to its lower average photon energy, when this low-energy X-ray penetrates the ore, the absorption difference for light elements such as silicon and aluminum and heavy elements such as iron and tungsten is significant. After imaging, the grayscale of the iron ore area is extremely low, while the grayscale of the gangue area is relatively high.
[0166] In the embodiments of this disclosure, since the candidate modes include at least two of the visible light imaging mode, the first energy level X-ray imaging mode, and the second energy level X-ray imaging mode, it ensures comprehensive coverage of two types of difficult ores that are similar in appearance but different in internal composition and those with similar compositional spectra but different surface features from a physical principle perspective. This provides a sufficiently rich selection space for the subsequent automatic mode selection mechanism, ensuring that the most discriminative mode can be found when facing ores with any physical properties, thereby improving the accuracy of object classification.
[0167] According to embodiments of this disclosure, the trained object classification model can be trained by: acquiring sample images of the sample object in at least two candidate modalities; inputting the sample image of each candidate modality into the feature extraction module for candidate modalities in the model to be trained, respectively, to obtain candidate sample features for each of the at least two candidate modalities; inputting the candidate sample features for each of the at least two candidate modalities into the classification module in the model to be trained, to obtain classification outputs corresponding to each of the at least two candidate modalities; and training the model to be trained based on the differences between the at least two classification outputs to obtain the object classification model.
[0168] Sample objects are objects to be classified that have been labeled with true class labels and used during the training phase. Sample images are image data of sample objects collected under various candidate modalities. For example, 5,000 pieces of ore collected from a mining area, whose categories have been identified by professional engineers, each piece of ore is accompanied by images in three modalities: visible light, high-energy X-ray, and low-energy X-ray. These images and their labels together constitute sample images of the sample objects under at least two candidate modalities.
[0169] The model to be trained is a neural network model that has not yet been trained and whose parameters need to be optimized with data. The model to be trained may include a feature extraction module and a classification module for each candidate modality. For the sample images under each candidate modality, a forward propagation process for feature extraction can be performed separately, that is, the sample images of the candidate modality are fed into the corresponding feature extraction module, and each feature extraction module independently calculates and outputs the candidate sample features under that candidate modality.
[0170] The feature extraction module is a sub-network used to convert the original images of the corresponding candidate modalities into high-dimensional feature representations. The candidate sample features are the high-dimensional feature representation tensors obtained after the sample images of each candidate modality have been processed by their respective feature extraction modules. They are abstract representations that convert the original pixels into discriminative expressions for classification tasks.
[0171] For example, taking a candidate mode including a visible light imaging mode, a first energy level X-ray imaging mode, and a second energy level X-ray imaging mode as an example, the model to be trained may include a feature extraction module 1, a feature extraction module 2, and a feature extraction module 3. Feature extraction module 1 is used to input a 3-channel visible light image and output a 128-dimensional feature map; feature extraction module 2 is used to input a 1-channel high-energy X-ray image and output a 128-dimensional feature map; feature extraction module 3 is used to input a 1-channel low-energy X-ray image and output a 128-dimensional feature map.
[0172] The specific structure of the feature extraction module for candidate modalities can be configured according to actual business needs and is not limited here. For example, n structurally similar but parameter-independent convolutional neural network branches can be designed, each receiving an image input from one modality and outputting a feature map of uniform dimension. Alternatively, for the feature extraction modules of each modality, the weights can be initialized using a general backbone network pre-trained on a large dataset. Single-channel X-ray images can be copied to three channels to adapt to the RGB pre-trained network. Alternatively, a partially shared feature extraction architecture can be designed, such as shallow convolutional layers shared across all modalities to extract common low-level features such as edges and textures; while deep convolutional layers are independent for each modality to extract modality-specific high-level semantic features.
[0173] For candidate sample features of at least two candidate modalities, the independent candidate sample features of each candidate modality can be fed into the classification module. The cross-layer connection design of the classification module maintains their independence in the channel, enabling the network to produce classification prediction results for different single modalities or combinations of modalities.
[0174] The classification module is a sub-network in the model to be trained that receives feature representations and outputs category predictions. For example, the classification module can be a classification head network consisting of several convolutional layers, pooling layers, and fully connected layers. The classification output is the category prediction produced by the classification module for a specific modality or combination of modalities based on the input features.
[0175] For example, candidate sample features from each modality can be concatenated along the channel dimension and fed into the classification module. The classification module employs a cross-layer connection design to ensure that the independence of features across channels is not compromised by fully connected operations. Alternatively, a multi-head classification structure can be designed, passing the concatenated mixed features through a multi-head attention layer, with each attention head focusing on a feature subspace of one modality. Each head is then connected to an independent lightweight classifier, outputting the classification prediction for that modality.
[0176] After obtaining the classification outputs of at least two candidate modalities, the differences between the classification outputs of each candidate modality can be calculated, and the model to be trained can be trained based on these differences. For example, the cosine distance between the classification outputs of each candidate modality can be calculated. A smaller cosine distance means that the class features are closer, i.e., the weaker the discriminative ability; a larger cosine distance means that the class features are farther apart, i.e., the stronger the discriminative ability. Alternatively, the L2 distance or Mahalanobis distance between the feature centers under the classification outputs of each candidate modality can be calculated as an inter-class separation index.
[0177] In the embodiments of this disclosure, by inputting sample images of each candidate modality into the feature extraction module separately, the physical attributes of each candidate modality are independently encoded, avoiding confusion of information from different physical sources at the feature level. By inputting the candidate sample features into the classification module, corresponding classification outputs are obtained. Through the design of maintaining independence across layers, the network can simultaneously produce predictions from different modal perspectives in a single forward propagation. Based on this, by training the model to be trained according to the differences between the classification outputs, automatic value judgment of which is superior is achieved, and it naturally converges to an object classification model that retains only the optimal modality, improving deployment efficiency and classification accuracy in application.
[0178] The following will utilize Figure 7 The first-stage training process of the object classification model provided in one embodiment of this disclosure will be further explained.
[0179] Figure 7 The illustration shows an example diagram of the training process of an object classification model according to an embodiment of the present disclosure.
[0180] like Figure 7As shown, in the embodiment 700 of training the object classification model, the sample object includes ore 710, and the model to be trained 730 includes a feature extraction module 7311 for candidate mode 1, a feature extraction module 7312 for candidate mode 2, ..., a feature extraction module 731M for candidate mode M and a classification module 732, which is a positive integer.
[0181] Obtain sample images of ore 710 under at least two candidate modes: sample image 721 under candidate mode 1, sample image 722 under candidate mode 2, ..., sample image 72M under candidate mode M.
[0182] For sample image 721, it can be input to feature extraction module 7311 for candidate mode 1 to obtain candidate sample features 741 corresponding to candidate mode 1; for sample image 722, it can be input to feature extraction module 7312 for candidate mode 2 to obtain candidate sample features 742 corresponding to candidate mode 2; and so on, for sample image 72M, it can be input to feature extraction module 731M for candidate mode M to obtain candidate sample features 74M corresponding to candidate mode M.
[0183] In one embodiment, the candidate sample features are obtained as shown in formula (6).
[0184] (6)
[0185] in, Characterize the features of the candidate sample for the k-th candidate modality. The sample image representing the k-th candidate modality. The convolutional layer operation performed by the feature extraction module for the k-th candidate modality in the model to be trained represents the operation performed by the module.
[0186] After obtaining the candidate sample features of at least two candidate modes, the mixed mode feature 750 can be determined based on the candidate sample feature 741, candidate sample feature 742, ..., candidate sample feature 74M.
[0187] In one embodiment, the mixed modal features are obtained as shown in formula (7).
[0188] (7)
[0189] in, Characterizing mixed modal features, Characterize the features of the first candidate modality candidate sample. Characterize the features of the candidate sample for the nth candidate modality.
[0190] The mixed modality feature 750 is input into the classification module 732 to obtain the mixed modality output 760. The mixed modality output 760 includes the classification output 771 corresponding to candidate modality 1, the classification output 772 corresponding to candidate modality 2, ..., and the classification output 77M corresponding to candidate modality M.
[0191] In one embodiment, the mixed modal output and the classification output of each candidate modality are obtained as shown in formulas (8) and (9).
[0192] (8)
[0193] (9)
[0194] in, Characterizing mixed-mode output, Representation classification module, The classification output characterizes the first candidate mode. The classification output characterizes the nth candidate mode.
[0195] Based on this, the model to be trained 730 can be trained according to the differences 780 between classification outputs 771, 772, ..., 77M to obtain the object classification model.
[0196] According to an embodiment of this disclosure, operation S630 may include the following operations: determining at least one target modality from at least two candidate modalities based on the difference between at least two classification outputs; adjusting the model parameters of the feature extraction module and the classification module corresponding to the target modality in the training model using sample images until a preset convergence condition is met to obtain an object classification model; the object classification model is configured to perform object classification based on images of the target modality.
[0197] For the classification outputs of at least two candidate modalities, the ability of each candidate modality to separate different categories can be quantified in the feature space to determine the target modality. The target modality is a single modality or a combination of multiple modalities that is automatically determined from multiple candidate modalities based on the differences in the classification outputs and is the most suitable for the current object classification task.
[0198] The method for determining the difference can be configured according to actual business needs and is not limited here. For example, the cosine distance between the feature centers of the classification output corresponding to each candidate modality can be calculated. If the cosine distance is larger, it indicates that the feature direction difference between the two classes of samples under that modality is greater, that is, the discriminative power is stronger. In this case, a single modality or multimodal combination that maximizes the cosine distance between classes can be selected as the target modality.
[0199] Alternatively, after several rounds of training in the first phase, the current model parameters can be frozen, and the classification accuracy can be evaluated on a validation set using the feature inputs of each candidate modality or their combination. In this case, the modality combination with the highest accuracy on the validation set can be selected as the target modality.
[0200] After determining the target modality through the first stage of training, the second stage of training can retain the feature extraction modules corresponding to the target modality in the model to be trained, while removing the feature extraction modules corresponding to other candidate modalities. The training is then iterated through a standard deep learning training loop: forward propagation to calculate the loss, backpropagation to calculate the gradient, and the optimizer to update the parameters, until the preset convergence condition is met.
[0201] A preset convergence criterion is a predetermined criterion for determining when model training can stop iterating. For example, a preset convergence criterion may include at least one of the following: the classification accuracy on the validation set reaches a threshold, the loss function value decreases by less than a threshold over several consecutive iterations, or the number of training iterations reaches a preset upper limit.
[0202] The resulting object classification model, after training, retains both the trained classification module and the trained feature extraction module corresponding to the target modality in its structure. During inference deployment, only the image of the target modality needs to be input to complete object classification.
[0203] In the embodiments of this disclosure, by determining the target modality based on the differences between classification outputs, the separability of each modality to the object can be automatically evaluated in the feature space, ensuring the objectivity and adaptability of modality selection. Based on this, by adjusting the model parameters of the feature extraction and classification modules corresponding to the target modality using sample images, the model capacity and training are focused on the most discriminative physical attributes, improving parameter adjustment efficiency. Furthermore, the final object classification model is configured to classify objects based on images of the target modality. Because its structure is adapted to the target modality, it does not require processing redundant modality data during inference, thus improving classification accuracy and efficiency.
[0204] According to embodiments of this disclosure, determining at least one target mode from at least two candidate modes based on the difference between at least two classification outputs may include the following operations: determining the difference between each classification output and other classification outputs based on the relative distance between any two classification outputs; and determining the candidate mode corresponding to the classification output with the largest numerical difference as the target mode.
[0205] Relative distance is a mathematical measure used to quantify the similarity or difference between two classification outputs. For example, the relative distance between two classification outputs can be obtained by calculating the cosine of the angle between them, reflecting their directional differences in the feature space. A larger relative distance indicates that the two classification outputs point in different directions in high-dimensional space.
[0206] For a specific candidate modality, the statistical value of the relative distance between the classification output corresponding to that candidate modality and all other classification outputs can be comprehensively measured. This statistical value reflects the uniqueness of the classification output among all classification outputs. For example, taking candidate modalities 1 to 3, if the relative distance between classification output 1 of candidate modality 1 and classification output 2 of candidate modality 2 is distance 1, and the relative distance between classification output 1 of candidate modality 1 and classification output 3 of candidate modality 3 is distance 2, then the difference between classification output 1 and classification output 2 and classification output 3 can be determined based on distance 1 and distance 2.
[0207] After obtaining the differences between each classification output and the other classification outputs, the target mode can be determined by identifying the largest difference. For example, the maximum value can be found directly among the differences of all classification outputs, and the candidate mode or mode combination associated with the classification output corresponding to this maximum value can be identified as the target mode. Alternatively, the maximum value can be found by sorting or comparing all differences one by one, and the candidate mode corresponding to this maximum value can be identified as the target mode.
[0208] In addition, it is possible not only to require the difference to be the largest, but also to require that the maximum and the second largest differences meet a preset difference to ensure that the difference is significant enough; if the maximum difference and the second largest difference do not meet the preset difference, then two candidate modalities with similar differences can be retained to further verify the accuracy on the validation set.
[0209] In one embodiment, the difference can be determined as shown in the following formula (10).
[0210] (10)
[0211] in, Representing the i-th classification output, Representing the output of the j-th classification, The cosine distance between the i-th and j-th classification outputs is represented by... The representation takes the maximum value. Characterize differences.
[0212] In the embodiments of this disclosure, the difference between each classification output and other classification outputs is determined based on the relative distance between any two classification outputs, fundamentally ensuring that the selected modality has sufficient discriminative power for different object types. Furthermore, by determining the candidate modality corresponding to the classification output associated with the largest numerical difference as the target modality, the target modality possesses both outstanding discriminative power and relative uniqueness, thereby contributing to improved accuracy in subsequent object classification.
[0213] The following will utilize Figure 8 The second-stage training process of the object classification model provided in one embodiment of this disclosure will be further explained.
[0214] Figure 8 The illustration shows an example diagram of the training process of an object classification model according to another embodiment of the present disclosure.
[0215] like Figure 8 As shown, in Example 800 of the training object classification model, the sample object includes ore 810, and the determined target mode is candidate mode 2. That is, the classification module 832 and the feature extraction module 8312 for candidate mode 2 in the model to be trained 830 are retained as an example.
[0216] The sample image 822 of ore 810 in candidate mode 2 is input to the feature extraction module 8312 for candidate mode 2 to obtain candidate sample features 840 corresponding to candidate mode 2. The candidate sample features 840 are then input to the classification module 832 to obtain the classification output 850 corresponding to candidate mode 2. Based on this, the model parameters of the feature extraction module 8312 and the classification module 832 in the training model 830 can be adjusted according to the classification output 850 until the preset convergence condition is met, thus obtaining the object classification model.
[0217] The above sections provided illustrative examples of the first-stage training process and the second-stage training process of the object classification model. The following will utilize... Figure 9 The process of classifying the overall objects is explained.
[0218] Figure 9 An example schematic diagram of an object classification process according to an embodiment of the present disclosure is shown.
[0219] like Figure 9 As shown, in embodiment 900 of object classification, images of the object 910 to be classified in at least two candidate modalities can be obtained.
[0220] Target detection is performed on the image 921 of the reference modality in at least two candidate modalities using detection model M1, obtaining position indication information 930 for each object to be segmented. Target segmentation is then performed on the image 921 of the reference modality using segmentation model M2, resulting in an initial mask image. Multiple initial masks in the initial mask image are then morphologically adjusted to form boundary indication information between the adjusted masks that meet preset position conditions, resulting in a reference mask image 940. Based on this, threshold segmentation can be performed on the reference mask image 940 using the position indication information 930 as a reference, yielding an image segmentation result 950.
[0221] After obtaining the image segmentation result 950, the image segmentation result 950 and the image 922 of the target modality in at least two candidate modalities can be input into the feature extraction module M31 of the trained object classification model M3 to obtain the candidate sample features 960 of the target modality; the candidate sample features 960 of the target modality can be input into the classification module M32 of the trained object classification model M3 to obtain the category 970 of each object to be classified.
[0222] The above are merely exemplary embodiments, but are not limited thereto. Other object classification methods known in the art may also be included, as long as they can improve the efficiency and accuracy of object classification.
[0223] Based on the above image segmentation method, the present invention also provides an image segmentation apparatus. The following will be combined with... Figure 10 The device is described in detail.
[0224] Figure 10 A block diagram of an image segmentation apparatus according to an embodiment of the present disclosure is shown schematically.
[0225] like Figure 10 As shown, the image segmentation device 1000 may include a first acquisition module 1010, a shape adjustment module 1020, and a first image segmentation module 1030.
[0226] The first acquisition module 1010 is used to acquire the initial mask image of the image to be segmented, wherein the initial mask image includes the initial masks of multiple objects to be segmented.
[0227] The shape adjustment module 1020 is used to adjust the shape of multiple initial masks respectively, so that boundary indication information is formed between the adjusted masks of the initial masks that meet the preset position conditions, and a reference mask image is obtained.
[0228] The first image segmentation module 1030 is used to segment multiple objects to be segmented in the image to be segmented based on the reference mask image, so as to obtain the image segmentation result.
[0229] Based on the above object classification method, the present invention also provides an object classification device. The following will be combined with... Figure 11 The device is described in detail.
[0230] Figure 11 A block diagram of an object classification apparatus according to an embodiment of the present disclosure is shown schematically.
[0231] like Figure 11 As shown, the object classification device 1100 may include a second acquisition module 1110, a second image segmentation module 1120, and an object classification module 1130.
[0232] The second acquisition module 1110 is used to acquire images of the object to be classified in at least two candidate modalities, wherein the images of different candidate modalities reflect different physical properties of the object to be classified.
[0233] The second image segmentation module 1120 is used to perform image segmentation processing on the reference modality image in at least two images to obtain the image segmentation result.
[0234] The object classification module 1130 is used to input images and image segmentation results under at least two candidate modalities into a trained object classification model to obtain the category of each object to be classified.
[0235] Any one or more of the modules according to embodiments of this disclosure, or at least a portion thereof, may be implemented in a single module. Any one or more of the modules according to embodiments of this disclosure may be implemented by dividing them into multiple modules. Any one or more of the modules according to embodiments of this disclosure may be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules according to embodiments of this disclosure may be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0236] It should be noted that the image segmentation device part in the embodiments of this disclosure corresponds to the image segmentation method part in the embodiments of this disclosure. For a detailed description of the image segmentation device part, please refer to the image segmentation method part, and it will not be repeated here. Similarly, the object classification device part in the embodiments of this disclosure corresponds to the object classification method part in the embodiments of this disclosure. For a detailed description of the object classification device part, please refer to the object classification method part, and it will not be repeated here.
[0237] Figure 12 A block diagram schematically illustrates an electronic device suitable for implementing an image segmentation method and an object classification method according to embodiments of the present disclosure. Figure 12 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0238] like Figure 12 As shown, a computer electronic device 1200 according to an embodiment of the present disclosure includes a processor 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage portion 1209 into a random access memory (RAM) 1203. The processor 1201 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1201 may also include onboard memory for caching purposes. The processor 1201 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0239] RAM 1203 stores various programs and data required for the operation of electronic device 1200. Processor 1201, ROM 1202 and RAM 1203 are interconnected via bus 1204.
[0240] According to embodiments of this disclosure, the electronic device 1200 may further include an input / output (I / O) interface 1205, which is also connected to a bus 1204. The electronic device 1200 may also include one or more of the following components connected to the input / output (I / O) interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output (I / O) interface 1205 as needed. A removable medium 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1210 as needed so that computer programs read from it can be installed into the storage section 1208 as needed.
[0241] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the image segmentation method and object classification method according to the embodiments of this disclosure.
[0242] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0243] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the image segmentation method and object classification method provided in the embodiments of this disclosure.
[0244] When the computer program is executed by the processor 1201, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0245] According to embodiments of this disclosure, program code for executing computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages.
[0246] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. It should also be noted that in some alternative implementations, the functions indicated in the boxes may occur in a different order than those shown in the drawings.
[0247] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. An image segmentation method, comprising: Obtain an initial mask image of the image to be segmented, wherein the initial mask image includes initial masks for each of the multiple objects to be segmented; The initial masks are morphologically adjusted to form boundary indication information between the adjusted masks that meet preset position conditions, thereby obtaining a reference mask image; and Based on the reference mask image, multiple objects to be segmented in the image to be segmented are segmented to obtain the image segmentation result.
2. The method according to claim 1, wherein, The adjusted mask was obtained in the following way: Based on attribute information obtained from multiple initial masks, shape adjustment parameters are determined, wherein the shape adjustment parameters indicate the required shape adjustment range to form the boundary indication information between the adjusted masks; and The initial mask is shaped according to the morphological adjustment parameters to obtain the adjusted mask.
3. The method according to claim 2, wherein, The determination of the shape adjustment parameters based on attribute information obtained from multiple initial masks includes at least one of the following: The shape adjustment parameters are determined based on first statistical information obtained from the area ratio of each initial mask relative to the initial mask image; as well as The morphological adjustment parameters are determined based on second statistical information obtained from the relative distances between the initial masks.
4. The method according to claim 2, wherein, The shape adjustment includes dilation, and the shape adjustment parameters include the size of the convolutional kernel used for dilation. The step of adjusting the shape of the initial mask according to the shape adjustment parameters to obtain the adjusted mask includes: The initial mask is traversed using a convolution kernel of the specified size to obtain the adjusted mask.
5. The method according to any one of claims 1 to 4, wherein, The step of segmenting multiple objects in the image to be segmented based on the reference mask image to obtain an image segmentation result includes: Using the boundary indication information, the binary image obtained by binarization of the image to be segmented is adjusted to obtain an adjusted image; and Using the location indication information obtained by target detection in the image to be segmented as a reference, the adjusted image is corrected to obtain the image segmentation result, wherein the location indication information represents the position of the object to be segmented in the image to be segmented.
6. The method according to claim 5, wherein, The object to be segmented in the binary image is displayed with a first gray value, and other regions in the binary image other than the object to be segmented are displayed with a second gray value; The step of adjusting the binary image obtained by binarization of the image to be segmented using the boundary indication information to obtain an adjusted image includes: Based on the boundary indication information, the value of the pixel position corresponding to the boundary indication information in the binary image is set to the second gray value to form a dividing line in the overlapping area between the adjusted masks, thereby obtaining the adjusted image.
7. The method according to claim 5, wherein, The step of correcting the adjusted image based on the location indication information obtained by target detection in the image to be segmented, to obtain the image segmentation result, includes: Connectivity analysis is performed on the adjusted image to distinguish the boundaries of each of the objects to be segmented, resulting in an object region image, wherein the object region image includes candidate regions for each of the multiple objects to be segmented; and Based on the location indication information, each candidate region is matched to obtain the image segmentation result.
8. The method according to claim 7, wherein, The step of matching each candidate region based on the location indication information to obtain the image segmentation result includes: The candidate region is matched with the location indication information to obtain the matching result; In response to the matching result indicating that there is a candidate region that does not match any of the location indication information, redundant candidate regions are deleted; and The image segmentation result is obtained based on the candidate regions corresponding to each of the location indication information.
9. The method according to any one of claims 1 to 8, wherein, The object to be segmented includes ore, and the preset position conditions include that the ore corresponding to the initial mask is stuck together or overlapping in the image to be segmented, and the boundary indication information is used to separate the stuck or overlapping ore.
10. An object classification method, comprising: Acquire images of the object to be classified in at least two candidate modalities, wherein the images of different candidate modalities reflect different physical properties of the object to be classified; Using the image segmentation method according to any one of claims 1 to 9, image segmentation processing is performed on at least two reference modal images to obtain image segmentation results; and The images and image segmentation results of at least two candidate modalities are input into a trained object classification model to obtain the category of each object to be classified.
11. The method according to claim 10, wherein, The trained object classification model was obtained through the following method: Obtain sample images of the sample object in at least two of the candidate modalities; The sample image of each candidate modality is input into the feature extraction module of the model to be trained to obtain the candidate sample features of at least two candidate modalities. The candidate sample features of at least two candidate modalities are input into the classification module of the model to be trained to obtain the classification output corresponding to each of the at least two candidate modalities. as well as The model to be trained is trained based on the differences between at least two of the classification outputs to obtain an object classification model.
12. The method according to claim 11, wherein, The step of training the model to be trained based on the difference between at least two of the classification outputs to obtain an object classification model includes: Based on the difference between at least two of the classification outputs, at least one target mode is determined from the at least two candidate modes; and The model parameters of the feature extraction module and the classification module corresponding to the target modality in the model to be trained are adjusted using the sample images until the preset convergence condition is met, thereby obtaining the object classification model; The object classification model is configured to classify objects based on images of the target modality.
13. The method according to claim 12, wherein, The step of determining at least one target mode from at least two candidate modes based on the difference between at least two classification outputs includes: Based on the relative distance between any two classification outputs, determine the difference of each classification output relative to the other classification outputs; and The candidate mode corresponding to the classification output associated with the largest numerical difference is determined as the target mode.
14. The method according to any one of claims 10 to 13, wherein, The candidate modes include at least two of the following: visible light imaging mode, X-ray imaging mode at a first energy level, and X-ray imaging mode at a second energy level, wherein the energy range corresponding to the first energy level is different from the energy range corresponding to the second energy level.
15. The method according to any one of claims 10 to 13, wherein, The objects to be classified include ores, and the reference mode includes X-ray imaging mode.
16. An image segmentation apparatus, comprising: The first acquisition module is used to acquire the initial mask image of the image to be segmented, wherein the initial mask image includes the initial masks of multiple objects to be segmented. A shape adjustment module is used to adjust the shape of multiple initial masks respectively, so that boundary indication information is formed between the adjusted masks that meet preset position conditions, thereby obtaining a reference mask image; and The first image segmentation module is used to segment multiple objects to be segmented in the image to be segmented based on the reference mask image, so as to obtain an image segmentation result.
17. An object classification device, comprising: The second acquisition module is used to acquire images of the object to be classified in at least two candidate modalities, wherein the images of different candidate modalities reflect different physical properties of the object to be classified. The second image segmentation module is configured to use the image segmentation apparatus of claim 16 to perform image segmentation processing on at least two reference modal images to obtain an image segmentation result; and The object classification module is used to input the images under at least two candidate modalities and the image segmentation results into the trained object classification model to obtain the category of each object to be classified.
18. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9 or any one of claims 10 to 15.
19. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9 or any one of claims 10 to 15.
20. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9 or any one of claims 10 to 15.