Image semantic recognition method and electronic device
Patent Information
- Application Number
- CN202610805566.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-05
AI Technical Summary
[0004]有鉴于此,本公开实施例致力于提供一种图像语义识别方法、存储介质、设备及产品,以解决现有技术中难以同时兼顾图像处理速度和语义分割精度的问题
[0022]第五方面,本公开一实施例提供了一种计算机程序产品,该计算机程序产品包括指令,该指令在电子设备上执行时使电子设备实现第一方面所述的图像语义识别方法。
Smart Images

Figure CN122336756B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer vision technology, specifically to an image semantic recognition method and electronic device. Background Technology
[0002] With the continuous development of artificial intelligence technology, intelligent devices such as robots are widely used in applications such as home services and intelligent companionship to provide services to humans. In these application scenarios, intelligent devices often need to have a certain degree of environmental perception capabilities.
[0003] To achieve environmental perception, semantic recognition is usually required based on real-time acquired images. However, image semantic recognition methods in related technologies struggle to simultaneously balance image processing speed and semantic segmentation accuracy. Summary of the Invention
[0004] In view of this, the embodiments of this disclosure aim to provide an image semantic recognition method, storage medium, device and product to solve the problem that it is difficult to simultaneously achieve image processing speed and semantic segmentation accuracy in the prior art.
[0005] In a first aspect, one embodiment of this disclosure provides an image semantic recognition method, comprising: performing semantic recognition processing on a first image to obtain an object recognition result corresponding to at least one first object instance, the object recognition result including boundary information; constructing prompt information based on the boundary information corresponding to the first object instance; performing semantic segmentation processing on an image region in the first image corresponding to the first object instance according to the prompt information to obtain an object segmentation result corresponding to the first object instance, wherein the image segmentation accuracy of the semantic segmentation processing is higher than the image segmentation accuracy of the semantic recognition processing; and determining the semantic recognition result corresponding to the first image based on the object recognition result and the object segmentation result corresponding to at least one first object instance.
[0006] In conjunction with the first aspect, in some implementations of the first aspect, the object recognition result further includes a first mask; before performing semantic segmentation processing on the image region corresponding to the first object instance in the first image according to the prompt information to obtain the object segmentation result corresponding to the first object instance, the above image semantic recognition method further includes: performing contour extraction processing on the first object instance based on the first mask corresponding to the first object instance to obtain contour region information; obtaining the contour region image corresponding to the first object instance from the first image according to the contour region information; wherein, performing semantic segmentation processing on the image region corresponding to the first object instance in the first image according to the prompt information to obtain the object segmentation result corresponding to the first object instance includes: performing semantic segmentation processing on the contour region image according to the prompt information to obtain a second mask; fusing the first mask and the second mask to obtain the object segmentation result corresponding to the first object instance.
[0007] In conjunction with the first aspect, in some implementations of the first aspect, contour extraction processing is performed on the first object instance based on the first mask corresponding to the first object instance to obtain contour region information, including: performing morphological gradient operation processing on the first mask corresponding to the first object instance to extract the contour pixel set; and expanding the contour pixel set outward by a target number of pixels to obtain contour region information.
[0008] In conjunction with the first aspect, in some implementations of the first aspect, the first mask and the second mask are fused to obtain the object segmentation result corresponding to the first object instance, including: applying the second mask to the contour region corresponding to the first object instance, applying the first mask to other regions other than the contour region, and splicing the first mask and the second mask to obtain the object segmentation result corresponding to the first object instance.
[0009] In conjunction with the first aspect, in some implementations of the first aspect, after splicing the first mask and the second mask, the method further includes: feathering the splicing area between the first mask and the second mask.
[0010] In conjunction with the first aspect, in some implementations of the first aspect, the object recognition result also includes a confidence level; constructing prompt information based on the boundary information corresponding to the first object instance includes: if the first object instance meets the target conditions, then constructing prompt information based on the boundary information corresponding to the first object instance; wherein, the target conditions include at least one of the following: the number of first object instances is greater than a preset number threshold, the confidence level corresponding to at least one first object instance is less than a preset confidence threshold, and there is overlap between the boundary information corresponding to at least two first object instances; the method further includes: if the first object instance does not meet the target conditions, then determining the object recognition result corresponding to at least one first object instance as the semantic recognition result corresponding to the first image.
[0011] In conjunction with the first aspect, in some implementations of the first aspect, when there are multiple first object instances, semantic segmentation processing is performed on the image regions in the first image corresponding to the first object instances based on the prompt information to obtain object segmentation results corresponding to the first object instances. This includes: concatenating the prompt information constructed corresponding to multiple first object instances to obtain batch prompt information; and performing parallel semantic segmentation processing on the image regions in the first image corresponding to multiple first object instances based on the batch prompt information to obtain object segmentation results corresponding to each of the multiple first object instances.
[0012] In conjunction with the first aspect, in some implementations of the first aspect, after performing semantic segmentation processing on the image region in the first image corresponding to the first object instance based on the prompt information to obtain the object segmentation result corresponding to the first object instance, the method further includes: obtaining the historical object segmentation result corresponding to the first object instance from the semantic recognition result corresponding to the historical image; and performing smoothing processing on the object segmentation result corresponding to the first object instance based on the historical object segmentation result.
[0013] In conjunction with the first aspect, in some implementations of the first aspect, the object recognition result also includes appearance feature information; before obtaining the historical object segmentation result corresponding to the first object instance from the semantic recognition result corresponding to the historical image, the method further includes: obtaining the object segmentation result and appearance feature information corresponding to the second object instance from the semantic recognition result corresponding to the second image, wherein the second image is the previous frame image of the first image; determining the correlation degree between the first object instance and the second object instance based on the object segmentation result and appearance feature information corresponding to the first object instance and the second object instance respectively; if the correlation degree is greater than a preset correlation threshold, then determining that the first object instance and the second object instance are the same object instance, and uniformly identifying the first object instance and the second object instance.
[0014] In conjunction with the first aspect, in some implementations of the first aspect, after determining the correlation between the first object instance and the second object instance based on the object segmentation results and appearance feature information corresponding to the first object instance and the second object instance respectively, the method further includes: if there is a target second object instance that satisfies a first condition among at least one second object instance corresponding to the second image, then determine the number of consecutively lost target frames of the target second object instance, wherein the first condition includes that the correlation between the target second object instance and any first object instance is less than or equal to a preset correlation threshold; if the number of target frames reaches a first preset number of frames, then determine whether there is a target first object instance that satisfies a second condition among at least one first object instance corresponding to the first image, wherein the second condition includes that the distance between the target second object instance and the target second object instance is less than a preset distance threshold; if there is a target first object instance, then retain the target second object instance.
[0015] In conjunction with the first aspect, in some implementations of the first aspect, after determining the number of consecutive target frames that the target second object instance has been lost, the method further includes: if the target frame number reaches a second preset frame number, then delete the target second object instance.
[0016] In conjunction with the first aspect, in some implementations of the first aspect, semantic recognition processing is performed by an object detection model, which is a general pre-trained model; before performing semantic recognition processing on the first image to obtain object recognition results corresponding to at least one first object instance, the method further includes: performing structured removal processing on other category detection channels in the object detection model other than the object category to be detected according to the object category to be detected corresponding to the target application scenario; and deploying the processed object detection model.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, semantic segmentation processing is performed by an image segmentation model, which is a general pre-trained model; before performing semantic segmentation processing on the image region in the first image corresponding to the first object instance according to the prompt information to obtain the object segmentation result corresponding to the first object instance, the method further includes: performing structured removal processing on redundant prediction branches in the image segmentation model used to output multiple candidate segmentation results corresponding to the object segmentation result; and deploying the processed image segmentation model.
[0018] In conjunction with the first aspect, in some implementations of the first aspect, semantic recognition processing and semantic segmentation processing are executed in independent asynchronous computation streams and run in a parallel inter-frame pipeline manner, so that when semantic segmentation processing is performed on the current frame image, semantic recognition processing is performed on the next frame image in parallel.
[0019] Secondly, one embodiment of this disclosure provides an image semantic recognition device, comprising: a semantic recognition module, configured to perform semantic recognition processing on a first image to obtain an object recognition result corresponding to at least one first object instance, the object recognition result including boundary information; a construction module, configured to construct prompt information based on the boundary information corresponding to the first object instance; a semantic segmentation module, configured to perform semantic segmentation processing on an image region in the first image corresponding to the first object instance according to the prompt information to obtain an object segmentation result corresponding to the first object instance, the image segmentation accuracy of the semantic segmentation processing being higher than the image segmentation accuracy of the semantic recognition processing; and a determination module, configured to determine the semantic recognition result corresponding to the first image based on the object recognition result and the object segmentation result corresponding to at least one first object instance.
[0020] Thirdly, one embodiment of this disclosure provides a computer-readable storage medium storing a computer program for performing the image semantic recognition method described in the first aspect.
[0021] Fourthly, one embodiment of this disclosure provides an electronic device, the electronic device comprising: a processor; a memory for storing processor-executable instructions; the processor being configured to perform the image semantic recognition method described in the first aspect.
[0022] Fifthly, one embodiment of this disclosure provides a computer program product including instructions that, when executed on an electronic device, cause the electronic device to implement the image semantic recognition method described in the first aspect.
[0023] In this embodiment, semantic recognition processing is performed on the first image to obtain the boundary information of the first object instance, and prompt information is constructed based on the boundary information. Since the prompt information is constructed based on the boundary information, it does not rely on manual intervention, thus solving the problem of relying on manual prompts in autonomous perception scenarios. Furthermore, by performing semantic segmentation processing only on the image region corresponding to the first object instance according to the prompt information, the computational overhead caused by fine processing of the entire image is effectively avoided, thereby achieving fast real-time processing and high-precision semantic segmentation, effectively balancing image processing speed and semantic segmentation accuracy. Attached Figure Description
[0024] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0025] Figure 1 The diagram shown is a flowchart of an image semantic recognition method provided in an embodiment of this disclosure.
[0026] Figure 2 The diagram shown is a flowchart of obtaining the object segmentation result corresponding to the first object instance according to an embodiment of this disclosure.
[0027] Figure 3 The diagram shown is a schematic flowchart of a timing consistency post-processing provided in an embodiment of this disclosure.
[0028] Figure 4 The diagram shown is a structural schematic of an image semantic recognition device provided in an embodiment of this disclosure.
[0029] Figure 5 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0030] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0031] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods and means well-known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0032] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0033] Furthermore, the terms “first,” “second,” “third,” and “fourth” are used only for distinguishing descriptions and should not be interpreted as indicating or implying relative importance.
[0034] In the field of real-time visual perception technology for embedded platforms, to enable mobile robots to quickly obtain semantic category and contour information of surrounding objects, related technologies typically deploy lightweight single-stage instance segmentation networks, such as the You Only Look Once Version 8 Segmentation (YOLOv8-seg) and the You Only Look Once Version 11 Segmentation (YOLO11-seg). Specifically, by simultaneously outputting object detection boxes and pixel-level masks through a shared convolutional backbone, detection and segmentation are integrated into a single forward inference, achieving high computational efficiency. For example, after downsampling the input image, the backbone network extracts multi-scale features, the detection head predicts object bounding boxes and categories, and the parallel segmentation head directly generates instance masks based on the multi-scale features.
[0035] However, when the above-mentioned solution is applied to scenarios such as indoor companion robots that require detailed depiction of furniture outlines, the quality of the output mask is insufficient. The segmentation head of a lightweight single-stage segmentation network typically operates on a low-resolution feature map and restores the original image size through upsampling. This process smooths and blurs the details of object edges, resulting in jagged edges or significant deviations from the true contours of the segmentation mask at slender structures such as sofa armrests, chair legs, and bed edges. Meanwhile, general-purpose fine-grained segmentation models, represented by the Segment Anything Model (SAM) and its lightweight variants such as the Mobile Segment Anything Model (MobileSAM) and the Edge Segment Anything Model (EdgeSAM), while possessing excellent zero-shot segmentation capabilities and edge accuracy, rely on externally provided cues such as points or boxes, lacking reliable human cues for autonomous robot operation. If full-image traversal cues or full-image fine-grained segmentation is used, the single-frame inference time on embedded graphics processing units (GPUs) can reach tens of milliseconds, failing to meet the processing latency requirements of real-time closed-loop control. This creates a dilemma: lightweight networks prioritizing real-time performance sacrifice edge accuracy, while fine-grained segmentation models that guarantee edge accuracy are difficult to directly apply to autonomous perception on edge computing platforms due to the lack of automatic cues and excessive computational overhead.
[0036] To address the aforementioned technical problems, this disclosure provides an image semantic recognition method and electronic device. The image semantic recognition method includes: performing semantic recognition processing on a first image to obtain an object recognition result corresponding to at least one first object instance, the object recognition result including boundary information; constructing prompt information based on the boundary information corresponding to the first object instance; performing semantic segmentation processing on the image region in the first image corresponding to the first object instance according to the prompt information to obtain an object segmentation result corresponding to the first object instance, the image segmentation accuracy of the semantic segmentation processing being higher than the image segmentation accuracy of the semantic recognition processing; and determining the semantic recognition result corresponding to the first image based on the object recognition result and the object segmentation result corresponding to at least one first object instance. Thus, by performing semantic recognition processing on the first image to obtain the boundary information of the first object instance, and constructing prompt information based on the boundary information, the prompt information is constructed based on the boundary information and does not rely on manual intervention, solving the problem of relying on manual prompts in autonomous perception scenarios. Furthermore, by performing semantic segmentation processing only on the image region corresponding to the first object instance according to the prompt information, the computational overhead caused by fine processing of the entire image is effectively avoided, thereby achieving real-time semantic processing and high-precision semantics, effectively balancing image processing speed and semantic segmentation accuracy.
[0037] The following is combined Figures 1 to 3 The image semantic recognition method provided in this disclosure is described in detail.
[0038] Figure 1 The diagram shown is a schematic flowchart of an image semantic recognition method provided in an embodiment of this disclosure. This method can be applied to electronic devices; exemplarily, the electronic device can be a mobile phone, computer, or other device with information processing capabilities. Figure 1 As shown, the image semantic recognition method may include the following steps.
[0039] Step S110: Perform semantic recognition processing on the first image to obtain object recognition results corresponding to at least one first object instance.
[0040] In one implementation, the object recognition result can characterize a set of information output by semantic recognition processing and associated with the detected object instance. For example, the object recognition result may include boundary information, a first mask, confidence level, semantic category label, appearance feature information, or any combination thereof.
[0041] The first image can be the current frame image captured in real time by an image sensor. Semantic recognition processing can characterize the computational process of locating and initially segmenting first object instances that may exist in the first image and belong to a preset category set.
[0042] In some implementations, semantic recognition processing can be performed by a pre-trained lightweight image segmentation model. For example, this could be the YOLO11n-seg model. The YOLO11n-seg model takes a first image downsampled to 640x640 pixels as input and outputs an object recognition result corresponding to at least one first object instance. Specifically, the first object instance can represent each individual detected in the first image, such as a sofa, a chair, or a table in an indoor scene.
[0043] In some implementations, the object recognition result may include boundary information, which can characterize the data representation that defines the spatial extent of the first object instance in the first image. Specifically, the boundary information may include an axis-aligned bounding box, typically represented by a quadruple (x, y, w, h), representing the x-coordinate and y-coordinate of the bounding box center, the width of the bounding box, and the height of the bounding box, respectively. Optionally, the boundary information may further include a coarse-grained instance mask, which is a binary image corresponding to the spatial dimensions of the first image, where pixels belonging to the first object instance are marked as 1, and the rest as 0, used to more finely delineate the spatial contour of the first object instance, but with relatively low edge precision. In addition, the object recognition result may optionally include a confidence score and a semantic category label corresponding to each first object instance. The confidence score reflects the probability that the model considers the detection result to be reliable, while the semantic category label indicates which category of indoor closed words such as sofa, bed, and chair the first object instance belongs to.
[0044] In one implementation, the object recognition result corresponding to at least one first object instance output by the YOLO11n-seg model can satisfy the following formula (1).
[0045] in, The model identifies the first The object recognition result corresponding to the first object instance The coordinate values representing boundary information , Indicates the first The first mask corresponding to the first object instance , Indicates the confidence level. , Indicates the number of instances of the first object. Indicates the first The semantic category label corresponding to each first object instance .
[0046] Step S120: Construct prompt information based on the boundary information corresponding to the first object instance.
[0047] For example, boundary information corresponding to at least one first object instance can be extracted from the object recognition results. Based on the boundary information, prompt information is constructed to drive subsequent semantic segmentation processing. The process of constructing the prompt information requires no manual intervention.
[0048] In some implementations, constructing the cue information can involve directly encoding the bounding box coordinates into a high-dimensional vector. For example, the bounding box coordinates can be fed into a cue encoder, which is part of the subsequent semantic segmentation model. The cue encoder maps spatial coordinates into a fixed-dimensional cue embedding vector, such as a 256-dimensional feature vector.
[0049] In one implementation, the process of constructing the prompt message can satisfy the following formula (2).
[0050] in, Indicates the first The prompt message corresponding to the first object instance .
[0051] Furthermore, the methods for constructing prompts are not limited to the encoding methods described above. For example, boundary information can be directly used as input parameters and passed to the subsequent semantic segmentation processing part through the application programming interface; or the bounding box can be rendered as a guide image and fed into the model for semantic segmentation processing along with the first image.
[0052] Step S130: Based on the prompt information, perform semantic segmentation processing on the image region in the first image corresponding to the first object instance to obtain the object segmentation result corresponding to the first object instance.
[0053] For example, the object segmentation result can be a fine-grained mask corresponding to the first object instance, obtained through semantic segmentation processing. Semantic segmentation processing can characterize any computational process capable of pixel-level classification of specific regions in an image to generate segmentation results that define the precise contours of objects. Semantic segmentation processing can be performed by an image segmentation model used for semantic segmentation. The image segmentation model can be a SAM, or a lightweight variant of SAM, such as the EdgeSAM or MobileSAM model. The image segmentation accuracy of semantic segmentation processing is higher than that of semantic recognition processing.
[0054] In one implementation, during semantic segmentation, the cue information is fed into the decoder of the EdgeSAM model. The decoder of the EdgeSAM model combines the global or local image feature maps extracted from the first image by the image encoder with the cue information, and performs fine decoding operations on the image regions defined by the boundary information, satisfying the following formula (3).
[0055] in, Indicates the first The fine mask corresponding to the first object instance This represents a local image feature map of the first image. Indicates the first The prompt message corresponding to the first object instance.
[0056] In some implementations, local image regions typically occupy only 5% to 15% of the total area of the first image, thus the computational cost of a single semantic segmentation process is far less than that of processing the entire image. By using a prompt-driven local processing mechanism, an object segmentation result corresponding to the first object instance is output. Compared to extracting the feature map of the entire image, the computational cost of the EdgeSAM model can be reduced by 60% to 80%.
[0057] In addition, in some embodiments, the semantic recognition processing in step S110 and the semantic segmentation processing in step S130 are executed in independent asynchronous computation streams and run in a parallel inter-frame pipeline manner, so that when performing semantic segmentation processing on the current frame image, semantic recognition processing on the next frame image is performed in parallel.
[0058] For example, the semantic recognition processing in step S110 and the semantic segmentation processing in step S130 can be executed in two independent asynchronous computation streams. The first computation stream is dedicated to performing semantic recognition processing, while the second computation stream is dedicated to performing semantic segmentation processing. These two computation streams run in a pipelined parallel manner between frames. Specifically, while the second computation stream is performing semantic segmentation processing on the t-th frame image, the first computation stream has already started synchronously and begun performing semantic recognition processing on the next frame, i.e., the t+1-th frame image. The two computation streams are coordinated through a lightweight event synchronization mechanism to ensure that when the second computation stream needs the object recognition result of the t-th frame, the object recognition result is already ready; the inference process of the second computation stream and the process of the first computation stream processing the t+1-th frame are completely overlapping.
[0059] In this way, by distributing the semantic recognition processing and semantic segmentation processing tasks into independent asynchronous computation streams to achieve inter-frame pipeline parallelism, the time for semantic segmentation processing of the current frame can overlap with the time for semantic recognition processing of the next frame. This reduces the end-to-end latency from the sum of the processing times of the two stages to approximately the time of the longer stage, significantly improving throughput and real-time performance.
[0060] In addition, in some embodiments, when there are multiple first object instances, step S130 may specifically include: concatenating the prompt information constructed corresponding to multiple first object instances to obtain batch prompt information; and performing parallel semantic segmentation processing on the image regions in the first image corresponding to multiple first object instances based on the batch prompt information to obtain object segmentation results corresponding to each of the multiple first object instances.
[0061] For example, in scenarios with multiple first object instances, if a sequential processing approach is adopted, the total latency of semantic segmentation will increase linearly with the number of first object instances. Therefore, a parallel execution strategy can be adopted for scenarios with multiple first object instances. Specifically, multiple independent prompts (e.g., multiple prompt embedding vectors) constructed for each first object instance in the first image are concatenated to form batch prompts. Then, the batch prompts are fed into the decoder of the image segmentation model all at once. The image encoder of the image segmentation model only needs to run once, performing parallel semantic segmentation processing on the image regions in the first image corresponding to multiple first object instances to extract image features shared by all first object instances. The decoder can then perform parallel semantic segmentation processing on the image regions corresponding to each first object instance in the batch, thereby obtaining object segmentation results corresponding to each of the multiple first object instances in the first image all at once.
[0062] In this way, by concatenating the prompts of multiple first object instances into batch prompts and performing parallel semantic segmentation, the total latency no longer increases linearly with the number of first object instances when processing scenarios with multiple first object instances, but remains basically constant or increases slowly, thereby improving processing efficiency and frame rate stability in crowded and complex scenarios.
[0063] Step S140: Based on the object recognition result and object segmentation result corresponding to at least one first object instance, determine the semantic recognition result corresponding to the first image.
[0064] For example, the semantic recognition result may include a mask, boundary information, confidence level, semantic category label, appearance feature information, or any combination thereof corresponding to at least one first object instance. The mask corresponding to the first object instance is obtained by fusing a first mask and a second mask. Based on the object recognition result and object segmentation result corresponding to at least one first object instance, the semantic recognition result corresponding to the first image can be determined. In one implementation, the object recognition result and object segmentation result corresponding to at least one first object instance can be fused to obtain the semantic recognition result corresponding to the first image. Specifically, the semantic recognition result can be used to support subsequent downstream tasks of the robot, such as grasping, obstacle avoidance, and 3D scene reconstruction.
[0065] In one implementation, the object recognition result corresponding to the first object instance may include a fine mask of the first object instance. The fine mask of the first object instance and the first mask can be weighted and fused to obtain the object segmentation result corresponding to the first object instance. For example, weighted fusion can be performed using formula (4).
[0066] in, Indicates the first The object segmentation result corresponding to the first object instance Indicates the first The fine mask corresponding to the first object instance Indicates the first The first mask corresponding to the first object instance, where α represents the weight, and α can be 0.7.
[0067] In this embodiment, semantic recognition processing is performed on the first image to obtain the boundary information of the first object instance, and prompt information is constructed based on the boundary information. Since the prompt information is constructed based on the boundary information, it does not rely on manual intervention, thus solving the problem of relying on manual prompts in autonomous perception scenarios. Furthermore, by performing semantic segmentation processing only on the image region corresponding to the first object instance according to the prompt information, the computational overhead caused by fine processing of the entire image is effectively avoided, thereby achieving fast real-time processing and high-precision semantic segmentation, effectively balancing image processing speed and semantic segmentation accuracy.
[0068] To adapt to complex and ever-changing real-world operating scenarios and balance computational load with perception performance, the image semantic recognition method provided in this disclosure can not only directly output the object recognition result as the semantic recognition result to reduce processing latency when the confidence level of the object recognition result is high, but also construct prompt information to guide subsequent processing when the confidence level of the recognition result is low, thereby reducing computational overhead while ensuring recognition accuracy. Therefore, based on the above solution, the following preferred solutions are also provided.
[0069] In one implementation, the object recognition result may further include a confidence level; step S120 may further include: if the first object instance satisfies the target condition, then constructing prompt information based on the boundary information corresponding to the first object instance. If the first object instance does not satisfy the target condition, then determining the object recognition result corresponding to at least one first object instance as the semantic recognition result corresponding to the first image.
[0070] For example, in step S120 above, not every frame of image unconditionally constructs prompt information and initiates semantic segmentation processing; instead, a conditional judgment is performed first. Specifically, the operation of constructing prompt information and subsequent steps is only executed when the detected first object instance meets the preset target conditions. The target conditions include at least one of the following: the number of first object instances is greater than a preset number threshold; the confidence level corresponding to at least one first object instance is less than a preset confidence threshold; and there is overlap between the boundary information corresponding to at least two first object instances. In some implementations, the number of first object instances is greater than the preset number threshold, for example, 3; there is at least one first object instance whose corresponding confidence level is less than a preset confidence threshold, for example, 0.85; and there are at least two first object instances whose corresponding bounding box information overlaps, indicating that mutual occlusion has occurred.
[0071] For example, if the first object instance in the current frame image does not satisfy any of the above target conditions, the current scene is considered relatively simple and belongs to the light-load mode. In this mode, all subsequent semantic segmentation processing steps will be skipped. At this time, the object recognition result including the first mask obtained in step S210 will be directly determined as the semantic recognition result corresponding to the first image.
[0072] In this way, by setting target conditions for scene complexity and starting and stopping semantic segmentation processing as needed, high-overhead fine processing can be automatically skipped in simple, low-dynamic scenes, and the process can run with low latency (e.g., about 10 milliseconds end-to-end). Semantic segmentation processing based on prompts and subsequent processing steps are only started when dealing with dense, uncertain, or occluded complex scenes, thereby improving real-time response capabilities while ensuring perception accuracy.
[0073] Furthermore, in some embodiments, the above-described image semantic recognition method can be applied to the edge computing platform of a robot. To further optimize the computational efficiency of semantic segmentation processing in the above embodiments and address the problem of how to further reduce the overhead of semantic segmentation processing when the computing power of the edge computing platform is very limited, this disclosure also provides a preferred solution.
[0074] Based on this, the following is combined with Figure 2Describe in detail the process of obtaining the object segmentation result corresponding to the first object instance.
[0075] Figure 2 The diagram shown is a flowchart illustrating the process of obtaining an object segmentation result corresponding to a first object instance, according to an embodiment of this disclosure. Figure 1 Extending from the illustrated embodiment Figure 2 The illustrated embodiment will be described in detail below. Figure 2 The illustrated embodiments and Figure 1 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.
[0076] like Figure 2 As shown, the object recognition result includes a first mask; before performing semantic segmentation processing on the image region corresponding to the first object instance in the first image according to the prompt information to obtain the object segmentation result corresponding to the first object instance (step S130 above), the image semantic recognition method further includes the following steps.
[0077] Step S210: Based on the first mask corresponding to the first object instance, perform contour extraction processing on the first object instance to obtain contour region information.
[0078] For example, the object recognition result may include a first mask. Contour extraction processing can characterize computational operations used to extract the edge regions of a first object instance from the first mask or a first image. For example, it may include, but is not limited to: algorithmic steps such as performing morphological gradient operations on the first mask to obtain a set of reference contour pixels; or processing steps such as locating boundaries through edge detection filtering, active contour models, etc. Based on the first mask corresponding to the first object instance, contour extraction processing can be performed on the first object instance to obtain contour region information.
[0079] Step S220: Based on the contour region information, obtain the contour region image corresponding to the first object instance from the first image.
[0080] For example, the contour region information can characterize the spatial description of the contour used to calibrate the first object instance and its surrounding protected area. For example, it may include, but is not limited to: a set of coordinates of an edge-sensitive zone generated by extending the reference contour pixel set outward by a target number of pixels, or a closed polygon representation of the region generated based on the contour line. Based on the contour region information, a contour region image corresponding to the first object instance is cropped from the first image, i.e., a small local patch containing only the boundary region of the first object instance, such as an edge stripe of the first object instance.
[0081] In addition, in some embodiments, step S210 may specifically include: performing morphological gradient operation on the first mask corresponding to the first object instance to extract the contour pixel set; and expanding the target number of pixels outward based on the contour pixel set to obtain contour region information.
[0082] For example, the step of extracting the contour of the first mask to obtain contour region information can be implemented through morphological operations. Specifically, morphological gradient operations are performed on the first mask corresponding to the first object instance to extract a set of reference contour pixels, i.e., the contour pixel set. In one implementation, the morphological gradient operation can be specifically expressed as performing morphological dilation and morphological erosion on the first mask using the following formula (5), and then subtracting the dilation result from the erosion result. The dilation and erosion operations can be performed using... The rectangular structuring element of the pixel is used to extract the skeleton line of one pixel width at the edge of the first mask, i.e., the set of contour pixels.
[0083] in, Represents the set of outline pixels. Indicates morphological expansion. Indicates morphological corrosion, express A rectangular structural element of a pixel.
[0084] However, since the edges of the first mask are not perfectly accurate, the edges of the actual first object instance may exist within a certain range near this skeleton line. Therefore, the extracted contour pixel set is further expanded outward by the target number of pixels to form a contour region information with a protected area. Specifically, it can be expanded outward by 16 pixels, for example, by performing morphological dilation on the contour pixel set using a disk structuring element with a radius of 16 pixels. The expanded region is called the edge-sensitive zone, i.e., the contour region information.
[0085] In one implementation, the minimum bounding rectangle of the contour region information can be obtained by using the following formula (6), the corresponding local image block can be cropped from the first image that has not been downsampled, and it can be upsampled to a higher resolution (such as 2 times or 4 times).
[0086] in, This represents a high-resolution local image patch after upsampling. Represents a local image patch. This indicates the upsampling factor.
[0087] In this way, by expanding the extracted set of contour pixels outward to form contour region information, it is ensured that even if the edge of the first mask is not accurate enough, the edge of the first object instance can still fall into the contour region information with a high probability. This effectively avoids the failure of real edge repair or the appearance of gaps at the splicing point due to the small selection area or positioning deviation in the high-precision processing during subsequent high-precision repair.
[0088] Step S230: Based on the prompt information, perform semantic segmentation processing on the contour region image to obtain the second mask.
[0089] For example, according to the prompt information, semantic segmentation processing is performed only on the contour region image obtained in step S220 above to obtain a second mask. The edge accuracy of the second mask is higher than that of the first mask. The semantic segmentation process is similar to that in the previous embodiment, except that this embodiment only performs fine semantic segmentation processing on the contour region image, rather than processing the entire image region corresponding to the first object instance.
[0090] Step S240: The first mask and the second mask are fused to obtain the object segmentation result corresponding to the first object instance.
[0091] For example, the object segmentation result is obtained by fusing the first mask and the second mask. The fusing method can be, for example, weighted fusing, stitching fusing, etc.
[0092] Taking the stitching and fusion process as an example, in some embodiments, the first mask and the second mask are fused to obtain the object segmentation result corresponding to the first object instance. Specifically, this may include: using the second mask for the contour region corresponding to the first object instance, using the first mask for other regions besides the contour region, and stitching the first mask and the second mask together to obtain the object segmentation result corresponding to the first object instance.
[0093] For example, the contour region can be the region corresponding to the contour pixel set or the region corresponding to the edge sensitive zone. In one implementation, the contour region corresponding to the first object instance adopts the second mask, and other regions other than the contour region adopt the first mask. The first mask and the second mask can be spliced based on the following formula (7) to obtain the object segmentation result corresponding to the first object instance.
[0094] in, This represents the object segmentation result obtained through concatenation processing, corresponding to the first object instance. Indicates the second mask. Indicates the first mask. Indicates the edge-sensitive zone.
[0095] In this way, by using region selection to synthesize the object segmentation results, the accuracy of the edge contour of the first object instance is guaranteed, while avoiding false positive noise or additional computational overhead that may be introduced due to overprocessing outside the contour region. This results in a balance between edge sharpness and contour region integrity in the final object segmentation result.
[0096] To prevent visually abrupt seams that may result from direct splicing, in some embodiments, after splicing the first mask and the second mask, the image semantic recognition method may further include feathering the splicing area between the first mask and the second mask.
[0097] In some examples, the stitching region can represent the boundary transition zone formed by extending the aforementioned contour region outward by a predetermined number of pixels (e.g., 4 pixels). Feathering can be a smoothing process such as linear interpolation or filtering. For example, within the boundary transition zone, the values of the first mask and the second mask can be feathered using linear interpolation, allowing the mask values from two different sources to transition smoothly.
[0098] In this way, by feathering the transition at the junction of the first and second masks, a smooth transition between regions of different precision is achieved, thereby eliminating possible hard edges or visual artifacts and improving the visual coherence and naturalness of the object segmentation result corresponding to the first object instance.
[0099] In this embodiment, since the large flat area of the first object instance, such as the seat surface of a sofa or the panel of a table, has consistent semantics and uniform color and texture, the first mask can already provide relatively accurate object segmentation results. Therefore, by running semantic segmentation processing only in the contour area, the amount of computation can be reduced without losing the overall segmentation accuracy, effectively alleviating the computing power bottleneck of the edge computing platform.
[0100] Furthermore, in robot video stream perception, processing each frame independently can cause inconsistent jitter or flickering in the object segmentation results over time, resulting in mask holes. Therefore, a temporal consistency post-processing step can be included after obtaining the object segmentation results.
[0101] Figure 3 The diagram shown is a schematic flowchart of a timing consistency post-processing method provided in an embodiment of this disclosure. Figure 1 Extending from the illustrated embodiment Figure 3 The illustrated embodiment will be described in detail below. Figure 3 The illustrated embodiments and Figure 1 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.
[0102] like Figure 3 As shown, after step S130 above, the image semantic recognition method may further include the following steps.
[0103] Step S310: Obtain the historical object segmentation result corresponding to the first object instance from the semantic recognition result corresponding to the historical image.
[0104] In some examples, the historical object segmentation result may include the final mask of the historical object instance in the historical image corresponding to the first object instance.
[0105] For example, historical object segmentation results corresponding to the first object instance are retrieved and obtained from a specific storage structure. The storage structure can be a sliding window buffer, which can store historical object segmentation results of the first object instance tracked in the most recent few frames (e.g., 5 frames). Based on the obtained historical object segmentation results, smoothing processing is performed on the object segmentation result corresponding to the first object instance in the current frame.
[0106] In one implementation, a fixed-length sliding window can be maintained to store the historical object segmentation results of the most recent 5 frames. The sliding window can satisfy the following formula (8).
[0107] in, Indicates a sliding window. This represents the final mask of the historical object instance corresponding to the first object instance in the most recent 5 historical images.
[0108] Step S320: Based on the historical object segmentation results, smooth the object segmentation results corresponding to the first object instance.
[0109] In one implementation, the object segmentation result may include a mask corresponding to the first object instance, obtained after stitching or merging. Smoothing methods may include, for example, Kalman smoothing or Bayesian smoothing.
[0110] For example, for each first object instance in the first image, the object segmentation result corresponding to the first object instance can be smoothed using historical statistical stability based on the following formula (9).
[0111] in, This represents the final, smoothed mask corresponding to the first object instance in the current frame image (i.e., the first image). This represents the mask obtained after stitching or merging the first object instance in the current frame image. This represents the average value of the historical mask (e.g., the average value of the mask of the first object instance in the most recent T-frame historical image). This represents the variance of the history mask. It represents the reciprocal of the confidence level of the current frame image.
[0112] In this embodiment, by using the historical object segmentation results corresponding to the first instance object to smooth the object segmentation results of the current frame corresponding to the first object instance, the noise introduced by the instability of a single frame is effectively suppressed, thereby reducing mask holes and edge jitter in the video stream and improving the temporal consistency of robot visual perception.
[0113] Furthermore, to achieve the aforementioned smoothing process, it is necessary to first establish associations for the same object instance across different frame images. Based on this, in some embodiments, the object recognition result also includes appearance feature information; before obtaining the historical object segmentation result corresponding to the first object instance from the semantic recognition result corresponding to the historical image, the image semantic recognition method may further include: obtaining the object segmentation result and appearance feature information corresponding to the second object instance from the semantic recognition result corresponding to the second image, where the second image is the previous frame image of the first image; determining the association degree between the first object instance and the second object instance based on the object segmentation result and appearance feature information corresponding to the first object instance and the second object instance respectively; if the association degree is greater than a preset association threshold, then determining that the first object instance and the second object instance are the same object instance, and uniformly identifying the first object instance and the second object instance.
[0114] For example, before obtaining historical object segmentation results from historical images, it is necessary to first determine which first object instance in the current frame image and which second object instance in the historical frame image are the same object instance. Specifically, the semantic recognition result corresponding to the second image is obtained. The second image is the previous frame image of the first image. The object segmentation result (such as a mask) and appearance feature information corresponding to the second object instance are obtained from the semantic recognition result. The appearance feature information can be a feature vector representing the appearance of the object instance output by the intermediate layer of the model used for semantic recognition processing, such as the 256-dimensional embedding vector extracted for each object instance by the backbone network of the YOLO11n-seg model.
[0115] In some examples, the degree of association can be an indicator that characterizes the degree of association between different object instances, and this indicator can be a joint indicator. For example, the degree of association between two object instances can be determined based on the intersection-union ratio (IU) of the masks corresponding to the two object instances and the cosine similarity of the appearance feature information of the two object instances. Among them, the IU measures the degree of overlap between the two object instances in spatial location and geometry; the cosine similarity measures the consistency between the two object instances in appearance semantics. In a specific implementation, the product of the IU and the cosine similarity can be used as the degree of association between the first object instance and the second object instance, as shown in the following formula (10).
[0116] in, Indicates the degree of relevance. Indicates intersection, union, and ratio. Represents cosine similarity. This represents the history mask for frame t. This represents the history mask for frame t-1. This represents the appearance feature information of frame t. This represents the appearance feature information of the (t-1)th frame.
[0117] For example, when the correlation between the first object instance and the second object instance is greater than a preset correlation threshold (e.g., 0.75), it can be determined that the first object instance and the second object instance are the same object instance. A unified identifier is then assigned to these two object instances, such as the first object instance inheriting the globally unique ID corresponding to the second object instance, thereby achieving cross-frame object instance tracking. When the correlation between the first object instance and all second object instances is less than or equal to the preset correlation threshold (e.g., 0.75), the first object instance is determined to be a newly entered instance, and a new identifier, such as a new globally unique ID, is assigned.
[0118] In this way, by combining the similarity of spatial geometric overlap and semantic appearance features of object instances to calculate the correlation, it surpasses the matching method that relies on only a single feature. Thus, even in complex situations such as rapid movement, brief deformation, or partial occlusion of object instances, it can accurately determine the identity of object instances across frames and effectively prevent the jumping of object instance identifiers.
[0119] In addition, in some embodiments, after determining the correlation between the first object instance and the second object instance based on the object segmentation results and appearance feature information corresponding to the first object instance and the second object instance respectively, the image semantic recognition method may further include: if there is a target second object instance that satisfies a first condition among at least one second object instance corresponding to the second image, then determine the number of consecutively lost target frames of the target second object instance, wherein the first condition includes that the correlation between the target second object instance and any first object instance is less than or equal to a preset correlation threshold; if the number of target frames reaches a first preset number of frames, then determine whether there is a target first object instance that satisfies a second condition among at least one first object instance corresponding to the first image, wherein the second condition includes that the distance between the target second object instance and the target second object instance is less than a preset distance threshold; if there is a target first object instance, then retain the target second object instance.
[0120] For example, after obtaining the correlation between the first object instance and each historical second object instance, it may be found that some second object instances in the previous frame cannot be successfully matched with any first object instance in the current frame (i.e., the correlation is all lower than or equal to the preset correlation threshold). These second object instances can be marked as potentially lost object instances, and the number of consecutive target frames of potentially lost object instances can be accumulated.
[0121] In some implementations, when the number of consecutively lost target frames for the second target object instance reaches a first preset number of frames (e.g., 3 frames), the system continues to check whether there is a target first object instance that satisfies a second condition among at least one first object instance corresponding to the first object. For example, whether there is an instance whose spatial location is very close to the historical location of the second object instance. Specifically, it is determined whether there is a target first object instance whose distance to the target second object instance is less than a preset distance threshold (e.g., after back-projecting the centers of the bounding boxes of the target first object instance and the target second object instance to the world coordinate system using camera intrinsics, their distance is less than 0.5 meters). If there is a target first object instance, it can be determined that the target second object instance has reappeared after a brief occlusion. At this time, a retention action is performed: the original identifier of the target second object instance is retained and not deleted, the historical buffer it occupies is not released, and the historical object segmentation result of the target second object instance can be used as a placeholder output of the first image to maintain its perceptual continuity.
[0122] In this way, by setting a first condition and then reconfirming it in conjunction with spatial location information when the first condition is met, it is possible to distinguish between two scenarios: temporary occlusion and actual departure. This avoids the incorrect deletion and reinitialization of the object instance identifier due to occasional detection loss or temporary occlusion, thereby maintaining the long-term stability and continuity of the object instance identifier.
[0123] In addition, in some embodiments, after determining the number of target frames that the target second object instance has been continuously lost, the image semantic recognition method may further include: if the number of target frames reaches a second preset number of frames, then delete the target second object instance.
[0124] For example, the second preset frame number can be a value greater than or equal to the first preset frame number. For instance, when the first preset frame number is 3 frames, the second preset frame number can be 4 frames. After determining the number of consecutive target frames that the target second object instance has lost, if the target frame number further accumulates and eventually reaches the second preset frame number, it can be determined that the target second object instance has left the scene. At this time, a deletion operation can be performed on the target second object instance, removing its relevant information, including its historical object recognition results, appearance feature information, identifiers, etc., from the memory buffer and the active instance list.
[0125] In this way, by performing the deletion operation after the target frame number reaches the second preset frame number, the storage space and computing resources occupied by object instances that have been confirmed to have left the scene can be released in a timely manner. This reduces the problem of memory leaks or decreased association matching efficiency caused by the infinite accumulation of information related to invalid object instances over time, and improves the stability and timeliness of long-term operation.
[0126] Furthermore, when deploying models used for semantic recognition processing to resource-constrained edge computing platforms, the general-purpose pre-trained models they rely on often suffer from severe storage and computational redundancy in order to cover a massive number of open vocabulary categories. To eliminate this redundancy, this disclosure provides a scheme for customized pruning before model deployment.
[0127] In addition, in some embodiments, semantic recognition processing is performed by an object detection model, which is a general pre-trained model; before performing semantic recognition processing on the first image to obtain an object recognition result corresponding to at least one first object instance, the image semantic recognition method may further include: performing structured removal processing on other category detection channels in the object detection model other than the object category to be detected according to the object category to be detected corresponding to the target application scenario; and deploying the processed object detection model.
[0128] For example, the semantic recognition processing in step S110 can be performed by an object detection model for semantic recognition processing. The object detection model can be a YOLO11n-seg model, which can have 80 open vocabulary classes. Before the object detection model is deployed and run online, it needs to undergo a pruning phase. Specifically, the category of the object instance to be detected corresponding to the target application scenario can be determined first, where the target application scenario can be the scenario that the method is actually applied to. For example, for an indoor companion robot for home use, the target application scenario can be an indoor scene, and the category of the object instance to be detected can be limited to a closed set consisting of 30 categories of indoor home furnishings. For example, sofa, single sofa, bed, bedside table, dining table, dining chair, desk, office chair, wardrobe, refrigerator, washing machine, television, TV cabinet, coffee table, side table, bookshelf, cabinet, stove, sink, toilet, bathtub, shower room, curtains, carpet, door, window, stairs, elevator, corridor, and others. Then, a structured removal process is performed on the detection channels of all categories other than the 30 categories of target object instances in the original object detection model, which was pre-trained on 80 open vocabulary classes. Specifically, during the model engine construction phase, the output channel weights of the classification convolutional layer responsible for predicting the scores of the other 50 non-target object instance categories are reset to zero, and the constant folding technique of the inference engine is used to eliminate these channels and all their downstream dependent computational subgraphs. The non-maximum suppression kernel function traverses the candidate boxes by category. After pruning, only the anchor indexes corresponding to the 30 indoor vocabulary classes need to be processed, reducing overhead and post-processing time by about 15%.
[0129] In this way, deploying the processed target detection model can reduce the weight storage size of its classification head by more than 60%, and correspondingly reduce runtime memory usage and memory bandwidth pressure. By structurally removing detection channels unrelated to the target scene and having the inference engine perform constant folding and computation subgraph elimination, the model's weight storage volume and runtime memory usage are reduced, and redundant computations of irrelevant categories are eliminated, making it more suitable for the limited memory and computing resources of edge computing platforms.
[0130] Similarly, the image segmentation model used to perform semantic segmentation processing also suffers from computational redundancy designed for general scenarios. This disclosure also provides a pre-deployment pruning scheme for this model. Furthermore, in some embodiments, the semantic segmentation processing is performed by an image segmentation model, which is a general pre-trained model; before performing semantic segmentation processing on the image region corresponding to the first object instance in the first image according to the prompt information to obtain the object segmentation result corresponding to the first object instance, the image semantic recognition method further includes: performing structured removal processing on the redundant prediction branches in the image segmentation model used to output multiple candidate segmentation results corresponding to the object segmentation result; and deploying the processed image segmentation model.
[0131] For example, the semantic segmentation processing in step S130 is performed by an image segmentation model for semantic segmentation processing. This image segmentation model can be a SAM model or a variant thereof. Specifically, the decoder of a typical general-purpose pre-trained segmentation model, such as a SAM model or a variant thereof, usually outputs multiple candidate segmentation results, from which the optimal one is selected by an internal prediction branch. However, in the application scenario of this embodiment, the prompt information input to the image segmentation model is constructed from a detection box that has already completed specific category recognition; its semantics are clear and almost unambiguous. Therefore, the prediction branches corresponding to multiple candidate segmentation results constitute structural redundancy.
[0132] Based on this, before deploying an image segmentation model online, the structure designed to handle ambiguity in its decoder can be pruned. Specifically, redundant prediction branches in the image segmentation model that originally output multiple candidate segmentation results corresponding to the object segmentation result can be structurally removed. For example, the EdgeSAM model's decoder outputs 4 mask tokens by default, corresponding to 4 candidate masks. The number of output mask tokens can be reduced from 4 to 1, retaining only a single mask prediction branch, and removing the 3 sets of multilayer perceptron layers and their corresponding mask prediction heads associated with the other 3 redundant tokens. After pruning, a more streamlined image segmentation model is deployed. In one implementation, standard layer-by-layer 8-bit integer quantization is performed on the EdgeSAM model, retaining only 16-bit half-precision floating-point precision for the first two layers of the image encoder and the last layer of the mask decoder, balancing accuracy and speed within the native support range of the EdgeSAM model framework.
[0133] In this way, by analyzing the determinism of the prompt information in specific application scenarios and removing the redundant decoding branches designed for outputting and selecting multiple candidate results within the image segmentation model, the number of parameters and computational load of the image segmentation model are directly reduced. While maintaining output accuracy, the pressure on video memory bandwidth and computational overhead are reduced.
[0134] In one implementation, fixed-size GPU workspaces are allocated to both the YOLO11n-seg and EdgeSAM models during the model building phase. Specifically, the YOLO11n-seg model has a workspace size of 800 MB, and the EdgeSAM model has a workspace size of 1.1 Gigabit. The CPU-side sliding window buffer, robot operating system deployment queue, and input layer unified memory zero-copy image buffer together occupy approximately 300 MB. The zero-copy buffer can be managed using a circular queue, with a fixed allocation of 4 frames (1280×720×4 bits = 3.6 MB) of red, green, and blue transparency frame buffers. The semantic recognition process can be bound to a fixed CPU core, and approximately 4.2 Gigabits of memory are reserved for navigation, voice, and system daemons.
[0135] In one implementation, the YOLO11n-seg model and the EdgeSAM model are compiled into independent model engine files, thus avoiding subgraph splitting and redundant memory allocation caused by hybrid networks. CUDA Graph Capture is performed on each independent model engine, solidifying the operations of the central processing unit, such as kernel startup, memory barriers, and event synchronization, into a single-submit GPU command graph, eliminating CPU launch overhead per frame (0.5–2 ms). The unsampled 1280×720 original image is handled through unified memory allocation, with the CPU and GPU sharing the same physical memory page. The processed image segmentation model and the processed object detection model operate directly on this shared memory, avoiding approximately 3.5 MB of bidirectional CPU-GPU transfer per frame. The processed image segmentation model and the processed object detection model can be used on a Jetson Orin Nano (8 GB version), with total GPU memory usage controlled within 2.5 GB (30% of total memory), leaving remaining resources for parallel use by navigation, voice interaction, or robotic arm control modules.
[0136] In some implementations, the YOLO11n-seg model runs continuously at a fixed frequency. After each frame is completed, the object recognition results are written to a double-buffered ping-pong memory area, and a stream event is marked. The CPU-side dynamic scheduler listens for these stream events. If the... If the frame meets the target conditions, the scheduler waits for the stream event to complete before starting the EdgeSAM model to perform semantic segmentation processing; if the target conditions are not met, the object recognition result corresponding to at least one first object instance is determined as the semantic recognition result corresponding to the first image.
[0137] It should be noted that any parts not fully described in the above examples can be referred to the aforementioned embodiments, and will not be repeated here.
[0138] The above text combined Figures 1 to 3The present disclosure describes in detail embodiments of the image semantic recognition method. The following is a combination of... Figure 4 This document describes in detail embodiments of the image semantic recognition apparatus of this disclosure. It should be understood that the descriptions of the image semantic recognition method embodiments correspond to the descriptions of the image semantic recognition apparatus embodiments; therefore, any parts not described in detail can be referred to the foregoing method embodiments.
[0139] Figure 4 The diagram shown is a structural schematic of an image semantic recognition device provided in an embodiment of this disclosure. Figure 4 As shown, the image semantic recognition device 400 provided in this embodiment includes: The semantic recognition module 410 is used to perform semantic recognition processing on the first image to obtain an object recognition result corresponding to at least one first object instance, and the object recognition result includes boundary information. Construction module 420 is used to construct prompt information based on the boundary information corresponding to the first object instance; The semantic segmentation module 430 is used to perform semantic segmentation processing on the image region in the first image corresponding to the first object instance according to the prompt information, so as to obtain the object segmentation result corresponding to the first object instance. The image segmentation accuracy of the semantic segmentation processing is higher than that of the semantic recognition processing. The determining module 440 is used to determine the semantic recognition result corresponding to the first image based on the object recognition result and object segmentation result corresponding to at least one first object instance.
[0140] In one embodiment of this disclosure, the object recognition result further includes a first mask; the semantic segmentation module 430 is further configured to: perform contour extraction processing on the first object instance based on the first mask corresponding to the first object instance to obtain contour region information; obtain a contour region image corresponding to the first object instance from the first image according to the contour region information; perform semantic segmentation processing on the contour region image according to the prompt information to obtain a second mask; and perform fusion processing on the first mask and the second mask to obtain an object segmentation result corresponding to the first object instance.
[0141] In one embodiment of this disclosure, the semantic segmentation module 430 is further configured to: perform morphological gradient operation on the first mask corresponding to the first object instance to extract the contour pixel set; and expand the target number of pixels outward based on the contour pixel set to obtain contour region information.
[0142] In one embodiment of this disclosure, the semantic segmentation module 430 is further configured to: apply a second mask to the contour region corresponding to the first object instance, apply a first mask to other regions besides the contour region, and perform a splicing process on the first mask and the second mask to obtain an object segmentation result corresponding to the first object instance.
[0143] In one embodiment of this disclosure, the semantic segmentation module 430 is further configured to: feather the splicing region between the first mask and the second mask.
[0144] In one embodiment of this disclosure, the object recognition result further includes a confidence level; the construction module 420 is further configured to: if the first object instance satisfies the target condition, construct prompt information based on the boundary information corresponding to the first object instance; wherein, the target condition includes at least one of the following: the number of first object instances is greater than a preset number threshold, the confidence level corresponding to at least one first object instance is less than a preset confidence threshold, and there is overlap between the boundary information corresponding to at least two first object instances; if the first object instance does not satisfy the target condition, the object recognition result corresponding to at least one first object instance is determined as the semantic recognition result corresponding to the first image.
[0145] In one embodiment of this disclosure, when there are multiple first object instances, the semantic segmentation module 430 is further configured to: concatenate the prompt information constructed corresponding to the multiple first object instances to obtain batch prompt information; and perform parallel semantic segmentation processing on the image regions in the first image corresponding to the multiple first object instances according to the batch prompt information to obtain object segmentation results corresponding to the multiple first object instances respectively.
[0146] In one embodiment of this disclosure, the semantic segmentation module 430 is further configured to: obtain a historical object segmentation result corresponding to the first object instance from the semantic recognition result corresponding to the historical image; and perform smoothing processing on the object segmentation result corresponding to the first object instance based on the historical object segmentation result.
[0147] In one embodiment of this disclosure, the object recognition result further includes appearance feature information; the semantic segmentation module 430 is further configured to: obtain the object segmentation result and appearance feature information corresponding to the second object instance from the semantic recognition result corresponding to the second image, wherein the second image is the previous frame image of the first image; determine the correlation degree between the first object instance and the second object instance based on the object segmentation result and appearance feature information corresponding to the first object instance and the second object instance respectively; if the correlation degree is greater than a preset correlation threshold, determine that the first object instance and the second object instance are the same object instance, and uniformly identify the first object instance and the second object instance.
[0148] In one embodiment of this disclosure, the semantic segmentation module 430 is further configured to: if there is a target second object instance that satisfies a first condition among at least one second object instance corresponding to the second image, then determine the number of consecutive target frames that the target second object instance has lost, wherein the first condition includes the correlation degree between the target second object instance and any first object instance being less than or equal to a preset correlation threshold; if the number of target frames reaches a first preset number of frames, then determine whether there is a target first object instance that satisfies a second condition among at least one first object instance corresponding to the first image, wherein the second condition includes the distance between the target second object instance and the target second object instance being less than a preset distance threshold; if there is a target first object instance, then retain the target second object instance.
[0149] In one embodiment of this disclosure, the semantic segmentation module 430 is further configured to: delete the target second object instance if the target frame number reaches a second preset frame number.
[0150] In one embodiment of this disclosure, semantic recognition processing is performed by an object detection model, which is a general pre-trained model; the semantic recognition module 410 is further configured to: perform structured removal processing on other category detection channels in the object detection model other than the object category to be detected according to the object category to be detected corresponding to the target application scenario; and deploy the processed object detection model.
[0151] In one embodiment of this disclosure, semantic segmentation processing is performed by an image segmentation model, which is a general pre-trained model; the semantic segmentation module 430 is further configured to: perform structured removal processing on redundant prediction branches in the image segmentation model that are used to output multiple candidate segmentation results corresponding to the object segmentation results; and deploy the processed image segmentation model.
[0152] In one embodiment of this disclosure, the determining module 440 is further configured to: execute semantic recognition processing and semantic segmentation processing in independent asynchronous computation streams and run in a parallel inter-frame pipeline manner, so that when performing semantic segmentation processing on the current frame image, semantic recognition processing on the next frame image is performed in parallel.
[0153] Below, for reference Figure 5 To describe an electronic device according to embodiments of the present disclosure. Figure 5 The diagram shown is a structural schematic of an electronic device provided in an exemplary embodiment of this disclosure.
[0154] like Figure 5 As shown, the electronic device 500 includes one or more processors 501 and memory 502.
[0155] The processor 501 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.
[0156] The memory 502 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 501 may execute the program instructions to implement the image semantic recognition methods of the various embodiments of this disclosure described above and / or other desired functions. Various content, such as object recognition results, prompts, and object segmentation results, may also be stored in the computer-readable storage medium.
[0157] In one example, the electronic device 500 may also include an input device 503 and an output device 504, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0158] The input device 503 may include, for example, a keyboard, a mouse, etc.
[0159] The output device 504 can output various information to the outside, including object recognition results, prompts, object segmentation results, etc. The output device 504 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0160] Of course, for the sake of simplicity, Figure 5 Only some of the components of the electronic device 500 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 500 may include any other suitable components depending on the specific application.
[0161] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products, including computer program instructions that, when executed by a processor, cause the processor to perform the steps in the image semantic recognition methods according to various embodiments of this disclosure described above.
[0162] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0163] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the image semantic recognition methods according to various embodiments of this disclosure described above.
[0164] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0165] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0166] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0167] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0168] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0169] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. An image semantic recognition method, characterized in that, For use in edge computing platforms, the method includes: The first image is subjected to semantic recognition processing to obtain an object recognition result corresponding to at least one first object instance, wherein the object recognition result includes boundary information and a first mask; Based on the boundary information corresponding to the first object instance, construct the prompt information; Perform morphological gradient operations on the first mask corresponding to the first object instance to extract the contour pixel set; Based on the contour pixel set, expand the target number of pixels outward to obtain contour region information; Based on the contour region information, obtain the contour region image corresponding to the first object instance from the first image; Based on the prompt information, semantic segmentation processing is performed on the contour region image to obtain a second mask. The image segmentation accuracy of the semantic segmentation processing is higher than that of the semantic recognition processing. The contour region corresponding to the first object instance is covered by the second mask, and other regions other than the contour region are covered by the first mask. The first mask and the second mask are spliced together, and the spliced region between the first mask and the second mask is feathered to obtain the object segmentation result corresponding to the first object instance. Based on the object recognition result and the object segmentation result corresponding to at least one of the first object instances, a semantic recognition result corresponding to the first image is determined.
2. The method according to claim 1, characterized in that, The object recognition result also includes a confidence level; the step of constructing prompt information based on the boundary information corresponding to the first object instance includes: If the first object instance satisfies the target condition, then a prompt message is constructed based on the boundary information corresponding to the first object instance; The target conditions include at least one of the following: the number of the first object instances is greater than a preset number threshold, the confidence level of at least one of the first object instances is less than a preset confidence threshold, and there is overlap between the boundary information corresponding to at least two of the first object instances. The method further includes: If the first object instance does not meet the target condition, the object recognition result corresponding to at least one of the first object instances will be determined as the semantic recognition result corresponding to the first image.
3. The method according to claim 1, characterized in that, When there are multiple instances of the first object, the step of performing semantic segmentation processing on the contour region image based on the prompt information to obtain a second mask includes: The prompt information constructed corresponding to multiple instances of the first object is concatenated to obtain batch prompt information; Based on the batch prompt information, parallel semantic segmentation processing is performed on the contour region images corresponding to multiple first object instances to obtain second masks corresponding to each of the multiple first object instances.
4. The method according to claim 1, characterized in that, After applying the second mask to the contour region corresponding to the first object instance, applying the first mask to other regions besides the contour region, performing a stitching process on the first mask and the second mask, and feathering the stitching region between the first mask and the second mask to obtain the object segmentation result corresponding to the first object instance, the method further includes: Obtain the historical object segmentation result corresponding to the first object instance from the semantic recognition results corresponding to the historical images; Based on the historical object segmentation results, the object segmentation results corresponding to the first object instance are smoothed.
5. The method according to claim 4, characterized in that, The object recognition result also includes appearance feature information; Before obtaining the historical object segmentation result corresponding to the first object instance from the semantic recognition result corresponding to the historical image, the method further includes: From the semantic recognition result corresponding to the second image, obtain the object segmentation result and appearance feature information corresponding to the second object instance, wherein the second image is the previous frame image of the first image; Based on the object segmentation results and appearance feature information corresponding to the first object instance and the second object instance respectively, the degree of association between the first object instance and the second object instance is determined. If the correlation degree is greater than the preset correlation threshold, then the first object instance and the second object instance are determined to be the same object instance, and the first object instance and the second object instance are uniformly identified.
6. The method according to claim 5, characterized in that, After determining the correlation between the first object instance and the second object instance based on the object segmentation results and appearance feature information corresponding to the first object instance and the second object instance respectively, the method further includes: If at least one of the second object instances corresponding to the second image contains a target second object instance that satisfies the first condition, then the number of target frames that the target second object instance has been continuously lost is determined, wherein the first condition includes that the correlation degree between the target second object instance and any first object instance is less than or equal to the preset correlation threshold. If the target frame count reaches a first preset frame count, then it is determined whether there is a target first object instance that satisfies a second condition among at least one first object instance corresponding to the first image, wherein the second condition includes the distance between the target second object instance and the target second object instance being less than a preset distance threshold; If the first target object instance exists, then the second target object instance is retained.
7. The method according to claim 6, characterized in that, After determining the number of consecutive target frames that the target second object instance has lost, the method further includes: If the target frame count reaches the second preset frame count, then the target second object instance is deleted.
8. The method according to any one of claims 1 to 7, characterized in that, The semantic recognition process is performed by an object detection model, which is a general pre-trained model. Before performing semantic recognition processing on the first image to obtain object recognition results corresponding to at least one first object instance, the method further includes: Based on the category of the object to be detected corresponding to the target application scenario, the detection channels of other categories in the target detection model, except for the category of the object to be detected, are subjected to structured removal processing. The processed target detection model is then deployed.
9. The method according to any one of claims 1 to 7, characterized in that, The semantic segmentation process is performed by an image segmentation model, which is a general pre-trained model. Before performing semantic segmentation on the contour region image based on the prompt information to obtain the second mask, the method further includes: The redundant prediction branches in the image segmentation model used to output multiple candidate segmentation results corresponding to the object segmentation result are structurally removed. The processed image segmentation model is then deployed.
10. The method according to any one of claims 1 to 7, characterized in that, The semantic recognition process and the semantic segmentation process are executed in independent asynchronous computation streams and run in a parallel inter-frame pipeline manner, so that when the semantic segmentation process of the current frame image is performed, the semantic recognition process of the next frame image is performed in parallel.
11. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the image semantic recognition method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Remote sensing video segmentation method and segmentation system based on text guidance
CN121074762A
Real-time medical image segmentation method and system
CN121725003A