Image processing method and device, electronic equipment, chip and storage medium
By analyzing the differences between the previous and subsequent frames and semantic segmentation results in the single-frame semantic segmentation model, the segmentation results of the current frame are optimized and adjusted, and the stability of image semantic segmentation is solved, and the time domain smoothing effect is realized during single-frame recognition. It is suitable for the fields of image semantic segmentation, video semantic segmentation, instance segmentation and target tracking.
Patent Information
- Application Number
- CN202411337244.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-07-25
AI Technical Summary
There is a lack of effective means in the prior art to improve the stability of image semantic segmentation results, especially when ignoring the timing relationship and unable to accurately identify changing scenarios.
By obtaining the semantic segmentation results of the first image frame and its historical frame, analyzing the inter-frame differences, and optimizing and adjusting the segmentation results of the current frame in combination with the semantic segmentation results, directly post-processing the output results of the single-frame semantic segmentation model to avoid extra information calculation.
It improves the time domain stability of the single-frame semantic segmentation model, reduces the amount of calculation, improves the stability and accuracy of semantic segmentation results, and is suitable for image processing tasks in multiple fields.
Smart Images

Figure CN120374969A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to an image processing method, apparatus, electronic device, chip, and storage medium. Specifically, it relates to the fields of image semantic segmentation and video semantic segmentation, and can be extended to multiple fields such as instance segmentation and object tracking. Background Art
[0002] Image semantic segmentation is to assign a semantic label to each pixel in an image, which is a basic task in image processing. The video semantic segmentation task is to assign a semantic label to each pixel in each image frame of a video based on the single-frame image semantic segmentation. It should be noted that the related technologies for improving the stability of semantic segmentation results mainly focus on using the temporal information between video frames. A video is composed of a sequence of consecutive single frames, and the sequence contains temporal information. By using the temporal information, the correlation of objects with the same spatial features can be obtained, thereby improving the consistency of multi-frame image segmentation results.
[0003] However, there is currently a lack of effective means for improving the stability of semantic segmentation results. Summary of the Invention
[0004] The present disclosure provides an image processing method, apparatus, electronic device, chip, and storage medium, which can solve the disadvantages of a single-frame semantic segmentation model in multi-frame recognition (such as ignoring temporal relationships, being unable to accurately recognize changing scenes, etc.), and can improve the temporal stability of the single-frame semantic segmentation model, which is extremely important for video segmentation.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:
[0006] Obtain a first image frame and a first semantic segmentation result of the first image frame;
[0007] Obtain a second image frame and a second semantic segmentation result of the second image frame, where the second image frame is one of the historical frames of the first image frame;
[0008] Obtain a final semantic segmentation result of the first image frame according to the difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result.
[0009] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:
[0010] A first obtaining module, configured to obtain a first image frame and a first semantic segmentation result of the first image frame;
[0011] A second acquisition module, configured to acquire a second image frame and a second semantic segmentation result of the second image frame, where the second image frame is one of the historical frames of the first image frame;
[0012] An adjustment module, configured to obtain a final semantic segmentation result of the first image frame according to a difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result.
[0013] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0014] One or more processors;
[0015] Wherein, the processor is configured to call an instruction to cause the electronic device to execute the image processing method described in the foregoing first aspect.
[0016] According to a fourth aspect of the embodiments of the present disclosure, there is provided a chip, including:
[0017] One or more processors;
[0018] Wherein, the processor is configured to call an instruction to cause the chip to execute the image processing method described in the foregoing first aspect.
[0019] According to a fifth aspect of the embodiments of the present disclosure, there is provided a storage medium storing instructions, and when the instructions run on a processor, the processor is caused to execute the image processing method described in the foregoing first aspect.
[0020] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method described in the foregoing first aspect are implemented.
[0021] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: The final semantic segmentation result of the first image frame can be obtained by optimizing and adjusting the first semantic segmentation result according to the difference between the first image frame and the second image frame, and at least one of the semantic segmentation results of the first image frame and the second image frame. Thus, it can be seen that the present disclosure can directly process the result output by single-frame semantic segmentation without using a multi-frame image input model or pre-performing edge extraction or foreground extraction on the image, reducing a large amount of computation and lowering the algorithm complexity. The previous image frame and the semantic segmentation result are both the input and output of the model itself and can be directly called. Therefore, the present disclosure does not need to introduce or calculate additional information for processing. Since the present disclosure does not rely on additional information such as contours or foregrounds and directly performs post-processing on the output result of the model, it can be directly deployed after any scene and any single-frame semantic segmentation model to improve the temporal smoothing effect. In addition, the present disclosure improves on the traditional temporal smoothing processing method for front and rear frames. When the single-frame semantic segmentation model outputs the semantic segmentation result of the current image frame, the previous image frame and its semantic segmentation result (i.e., the semantic segmentation mask) are used to jointly adjust the semantic segmentation output result of the current image frame, thereby improving the temporal smoothing effect and avoiding increasing the use of complex calculation methods to increase additional computation.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0024] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment.
[0025] Figure 2 is a flowchart of an image processing method shown according to an exemplary embodiment.
[0026] Figure 3 is a schematic diagram of the single-frame semantic segmentation process provided by the embodiments of the present disclosure.
[0027] Figure 4 is a schematic diagram of the temporal smoothing post-processing process provided by the embodiments of the present disclosure.
[0028] Figure 5 is a comparison example diagram of the segmentation results of image processing using the benchmark method (a) and the method (b) provided by the present disclosure.
[0029] Figure 6 It is a block diagram of an image processing apparatus shown according to an exemplary embodiment.
[0030] Figure 7 It is a block diagram of an electronic device 700 shown according to an exemplary embodiment.
[0031] Figure 8 It is a schematic structural diagram of a chip 800 proposed in an embodiment of the present disclosure. Detailed implementation manners
[0032] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0033] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present disclosure. The singular forms "a" and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0034] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present disclosure to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the embodiments of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "when" or "in response to a determination".
[0035] It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solution of the present disclosure all comply with the relevant regulations of national laws and regulations, and do not violate public order and good customs.
[0036] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by users or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0037] It should be noted that in the embodiments of the present disclosure, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned. They should be considered exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present disclosure, but it does not mean that the applicant has already or necessarily used this solution.
[0038] The image processing method, apparatus, electronic device, chip, and storage medium according to the embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be noted that the image processing method provided by the embodiments of the present disclosure can be used in the post-processing stage of the image semantic segmentation task for post-processing oriented to the temporal stability of semantic segmentation.
[0039] Figure 1 It is a flowchart of an image processing method shown according to an exemplary embodiment. As Figure 1 shown, the image processing method may include but is not limited to the following steps.
[0040] In step 101, a first image frame and a first semantic segmentation result of the first image frame are obtained.
[0041] In some embodiments, the first image frame may be an image frame in a video, and the video may be a video for which a semantic segmentation task needs to be performed. Exemplarily, the first image frame may be the current image frame to be processed, that is, the currently to-be-processed image frame.
[0042] In some embodiments, the semantic segmentation result of the first image frame may be obtained by performing semantic segmentation on the first image frame. In some embodiments, the first image frame may be input into a semantic segmentation model to obtain the first semantic segmentation result of the first image frame. Among them, the semantic segmentation model may be a pre-trained model. Exemplarily, during training, a dataset including multiple images (such as including 30,000 images) that have been labeled may be used. The data categories may be multiple categories, such as 11 categories, and may include but are not limited to: people, mouths, faces, green plants, eyes, ground, sky, eyebrows, human skin, hair, and other categories, etc. At least some of the labeled images (such as 5,000 images) may be randomly selected as the test set. The data of the test set and the training set may not overlap.
[0043] In some embodiments, the semantic segmentation model may be a semantic segmentation model trained on a single-frame image. That is to say, the semantic segmentation model may be a single-frame semantic segmentation model. In some embodiments, the single-frame semantic segmentation model may use any single-frame model without the need to design an additional network structure for different scenarios to enhance the temporal stability of the output.
[0044] In step 102, a second image frame and a second semantic segmentation result of the second image frame are obtained.
[0045] In some embodiments, the second image frame may be one of the historical frames of the first image frame, that is, the second image frame may be an image frame before the first image frame. Exemplarily, the second image frame may be the previous image frame of the first image frame. Exemplarily, the first image frame and the second image frame may be two adjacent image frames in a video (or may also be referred to as two adjacent image frames). For example, the first image frame may be the current image frame to be processed, and the second image frame may be the previous image frame of the first image frame.
[0046] In some embodiments, the semantic segmentation result of the second image frame may be obtained by performing semantic segmentation on the second image frame. In some embodiments, the second image frame may be input into a semantic segmentation model (such as a single-frame semantic segmentation model) to obtain the second semantic segmentation result of the second image frame. In some embodiments, the second image frame may be the input image input into the semantic segmentation model. The second semantic segmentation result may be jointly adjusted using the semantic segmentation result of the previous image frame of the second image frame to optimize the semantic segmentation result of the second image frame, so as to obtain the final semantic segmentation result of the second image frame, that is: the optimized final semantic segmentation result obtained by jointly adjusting the segmentation output result of the second image frame using the third image frame (the previous image frame of the second image frame) and the semantic segmentation result of the third image frame.
[0047] In step 103, according to the difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result, the final semantic segmentation result of the first image frame is obtained.
[0048] In some embodiments, the difference between the first image frame and the second image frame described above may be based on the similarity between each pixel point in the first image frame and the pixel point at the same position in the second image frame. Exemplarily, the difference between the first image frame and the second image frame may be determined according to the similarity between each pixel point in the first image frame and the pixel point at the same position in the second image frame. That is to say, the difference analysis can be performed on each pixel point at the same position in the first image frame and the second image frame to obtain the difference analysis result of each pixel point in the first image frame on the two image frames.
[0049] In some embodiments, the difference analysis may be performed on the pixel points at the same position in the first image frame and the second image frame. Among them, the pixel points at the same position can be understood as: the pixel points with the same coordinate position in the first image frame and the second image frame, that is, the pixel points at the same coordinate position in the first image frame and the second image frame.
[0050] In a possible implementation, an image similarity comparison method can be adopted to perform a difference analysis on the pixel points at the same positions in the first image frame and the second image frame, so as to obtain the difference analysis results of each pixel point in the first image frame on the two image frames. Exemplarily, the greater the similarity of the pixel points at the same positions in the first image frame and the second image frame, the smaller the change amplitude of the pixel point on the two image frames can be indicated; the smaller the similarity of the pixel points at the same positions in the first image frame and the second image frame, the greater the change amplitude of the pixel point on the two image frames can be indicated. Among them, the image similarity comparison method can include but is not limited to any one of the following: histogram algorithm; grayscale image algorithm; hashing algorithm; cosine similarity; Euclidean distance, etc.
[0051] It should be noted that in some embodiments, other means can also be adopted to implement the change analysis of each pixel point at the same position in the first image frame and the second image frame. For example, an image offset detection method can be used to analyze whether there are changes in the pixel points at the same position in the first image frame and the second image frame, that is: by obtaining an image sequence, detecting the image texture, performing binarization processing on the image texture, and compressing every predetermined number of pixel points in the binarized image to form a compressed image block. Then, a shift matching operation is performed on the compressed image blocks in two adjacent image frames, and the rigid displacement amount of two adjacent image frames is obtained by using the offset amount of the matching image blocks in two adjacent image frames. Another example is that an optical flow method can also be used to analyze whether there are changes in the pixel points at the same position in two image frames, that is: by calculating the motion vectors (i.e., optical flow) of pixels or regions between adjacent frames, the dynamic changes in the image can be detected. Here, the present disclosure does not make any limitations in this regard and will not elaborate further.
[0052] After determining the difference between the first image frame and the second image frame, the final semantic segmentation result of the first image frame can be obtained according to the difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result. Exemplarily, the difference analysis results of each pixel point in the first image frame on the two image frames, and according to at least one of the second semantic segmentation result and the first semantic segmentation result, the class segmentation mask of the pixel points in the first semantic segmentation result can be correspondingly processed to obtain the final semantic segmentation result of the first image frame.
[0053] In some embodiments, the difference analysis results of the above-mentioned pixel points on the two image frames can be understood as the degree of change or the amplitude of the pixel points on the two image frames. According to the difference analysis results of each pixel point in the first image frame on the two image frames, and according to at least one of the second semantic segmentation result and the first semantic segmentation result, the first semantic segmentation result of the first image frame can be post-processed in the time domain, so as to obtain the final semantic segmentation result of the first image frame. Exemplarily, based on the difference analysis results of each pixel point in the first image frame on the two image frames, the category segmentation mask of the pixel point with a larger change amplitude in the first semantic segmentation result can be retained, and the category segmentation mask of the pixel point with a smaller change amplitude but a higher reliability of the semantic segmentation result can be retained in the first semantic segmentation result; for the pixel point with a smaller change amplitude but a lower reliability of the semantic segmentation result in the first semantic segmentation result, the category segmentation mask of the pixel point in the first semantic segmentation result is replaced with the category segmentation mask in the second semantic segmentation result, that is, the semantic segmentation mask result of the previous frame of the pixel point is used as the semantic segmentation mask result of the current frame.
[0054] Exemplarily, for each pixel in the first image frame, if the difference analysis result of the pixel on the two image frames (i.e., the second image frame and the first image frame) is that the change amplitude is relatively large (i.e., the pixel has undergone a large change in the two image frames, such as a large change amplitude, such as a change amplitude greater than or equal to a threshold), the first semantic segmentation result of the first image frame shall prevail, that is, the category segmentation mask of the pixel in the first semantic segmentation result can be retained. If the change analysis result of the pixel on the two image frames (i.e., the second image frame and the first image frame) is that the change amplitude is not large (i.e., the pixel has not undergone a large change in the two image frames, such as a small change amplitude, such as a change amplitude less than a threshold), the category segmentation mask of the pixel in the first semantic segmentation result can be adjusted based on the reliability of the category segmentation mask of the pixel in the first semantic segmentation result. For example, the higher the reliability of the pixel point in the first semantic segmentation result, the more accurate the semantic segmentation result of the pixel point can be considered, and the category segmentation mask of the pixel point in the first semantic segmentation result can be retained; for another example, the lower the reliability of the pixel point in the first semantic segmentation result, the lower the accuracy of the semantic segmentation result of the pixel point can be considered, and the category segmentation mask of the pixel point in the second image frame can be used as the category segmentation mask of the pixel point in the first image frame, and so on. Each pixel point is processed through the above steps, so as to obtain the final semantic segmentation result of the first image frame.
[0055] That is to say, for each pixel in the image, the present disclosure analyzes the changes of the pixel in the previous and subsequent frames, and performs different post-processing methods for temporal smoothing based on the differences in the changes, that is, different post-processing methods are used to optimize and adjust the semantic segmentation result of the current image frame based on the differences in the changes, so as to obtain the final semantic segmentation result of the current image frame.
[0056] In the above embodiment, for the semantic segmentation task, the present disclosure re-designs the temporal post-processing algorithm. After the output stage of the single-frame semantic segmentation model, the input image and the output mask of the single-frame semantic model are comprehensively utilized, and a good temporal performance can be achieved with a single-frame semantic segmentation model. The single-frame semantic segmentation model here can use any single-frame model, and there is no need to design an additional network structure for different scenarios to enhance the temporal stability of the output. In addition, by introducing the comparison of the images of the previous and subsequent frames, the present disclosure can more directly and accurately determine the changes in the images of the previous and subsequent frames, and perform different processing on the regions where the images of the previous and subsequent frames change. Compared with the traditional temporal post-processing method, the present disclosure can perform targeted smoothing on the regions where the images change without affecting the regions with no obvious changes such as the background, which can improve the accuracy, thereby improving the effect of temporal smoothing and enhancing the stability of the semantic segmentation result.
[0057] Figure 2 It is a flowchart of an image processing method shown according to an exemplary embodiment. As Figure 2 shown, the image processing method may include but is not limited to the following steps.
[0058] In step 201, a first image frame and a first semantic segmentation result of the first image frame are obtained.
[0059] In the embodiments of the present disclosure, step 201 can be implemented in any one of the embodiments of the present disclosure. The embodiments of the present disclosure do not limit this and will not be described in detail.
[0060] In step 202, a second image frame and a second semantic segmentation result of the second image frame are obtained.
[0061] In the embodiments of the present disclosure, step 202 can be implemented in any one of the embodiments of the present disclosure. The embodiments of the present disclosure do not limit this and will not be described in detail.
[0062] In step 203, the similarity between each pixel in the first image frame and the pixel at the same position in the second image frame is determined.
[0063] In some embodiments, a similarity algorithm can be used to determine the similarity between each pixel point in the first image frame and the pixel point at the same position in the second image frame. Exemplarily, for each pixel point in the first image frame, a similarity algorithm can be used to calculate the similarity between this pixel point and the pixel point at the same position in the second image frame. In some embodiments, the similarity algorithm can include, but is not limited to, any one of the following: distance (such as Euclidean distance) similarity calculation method; cosine similarity calculation method, etc. Exemplarily, the Euclidean distance similarity calculation method can be used to calculate the Euclidean distance between each pixel point in the first image frame and the pixel point at the same position in the second image frame, and this Euclidean distance can be used as the similarity.
[0064] In step 204, based on the similarity between each pixel point in the first image frame and the pixel point at the same position in the second image frame, the difference between the first image frame and the second image frame is determined.
[0065] It should be noted that the closer the pixel point in the first image frame is to the pixel point at the same position in the second image frame, the greater their similarity; the more distant the pixel point in the first image frame is from the pixel point at the same position in the second image frame, the smaller their similarity.
[0066] In some embodiments, the similarity between each pixel point in the first image frame and the pixel point at the same position in the second image frame can be compared with a similarity threshold, and based on the obtained comparison result, the difference analysis result of each pixel point in the first image frame in the two image frames can be determined. Exemplarily, if the similarity between the pixel point in the first image frame and the pixel point at the same position in the second image frame is less than the similarity threshold, it can be explained that the similarity of these two pixel points at the same position is small, that is, the change amplitude of this pixel point in these two image frames is large. If the similarity between the pixel point in the first image frame and the pixel point at the same position in the second image frame is greater than or equal to the similarity threshold, it can be explained that the similarity of these two pixel points at the same position is large, that is, the change amplitude of this pixel point in these two image frames is small.
[0067] Exemplarily, taking the Euclidean distance as an example, the Euclidean distance between each pixel point in the first image frame and the pixel point at the same position in the second image frame is compared with a distance threshold. If the Euclidean distance between the pixel point in the first image frame and the pixel point at the same position in the second image frame is greater than or equal to this distance threshold, it can be explained that the similarity of these two pixel points at the same position is small, that is, the change amplitude of this pixel point in these two image frames is large. If the Euclidean distance between the pixel point in the first image frame and the pixel point at the same position in the second image frame is less than this distance threshold, it can be explained that the similarity of these two pixel points at the same position is large, that is, the change amplitude of this pixel point in these two image frames is small.
[0068] In step 205, based on the difference between the first image frame and the second image frame, and based on at least one of the second semantic segmentation result and the first semantic segmentation result, the final semantic segmentation result of the first image frame is obtained.
[0069] In some embodiments, for each pixel point in the first image frame, according to the similarity between the pixel point and the pixel point at the same position in the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result, the class segmentation mask of the pixel point in the first semantic segmentation result is processed accordingly to obtain the final semantic segmentation result of the first image frame.
[0070] In some embodiments, if the similarity between the pixel point and the pixel point at the same position in the second image frame is less than the similarity threshold, the class segmentation mask of the pixel point in the first semantic segmentation result is retained; or, if the similarity between the pixel point and the pixel point at the same position in the second image frame is greater than or equal to the similarity threshold, the class segmentation mask of the pixel point in the first semantic segmentation result is adjusted based on the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result.
[0071] Exemplarily, taking any pixel point A in the first image frame as an example, if the difference analysis result of pixel point A in the first image frame and the second image frame is that the similarity between pixel point A and the pixel point at the same position in the second image frame is less than the similarity threshold, it can be explained that the change amplitude of pixel point A in the first image frame and the second image frame is relatively large. At this time, the class segmentation mask mask of pixel point A in the first semantic segmentation result can be retained. If the difference analysis result of pixel point A in the first image frame and the second image frame is that the similarity between pixel point A and the pixel point at the same position in the second image frame is greater than or equal to the similarity threshold, it can be explained that the change amplitude of pixel point A in the first image frame and the second image frame is relatively small. At this time, the class segmentation mask mask of pixel point A in the first semantic segmentation result can be adjusted based on the confidence of the class segmentation mask of pixel point A in the first semantic segmentation result. Among them, the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result can be the probability corresponding to the class segmentation mask of the pixel point in the first semantic segmentation result.
[0072] In some embodiments, if the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result is greater than or equal to the confidence threshold, the class segmentation mask of the pixel point in the first semantic segmentation result is retained; or, if the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result is less than the confidence threshold, the class segmentation mask of the pixel point in the first semantic segmentation result is replaced with the class segmentation mask of the corresponding pixel point in the second semantic segmentation result.
[0073] Exemplarily, taking any pixel point A in the first image frame as an example, if the confidence of the class segmentation mask of pixel point A in the first semantic segmentation result is greater than or equal to the confidence threshold, the class segmentation mask mask of pixel point A in the first semantic segmentation result can be retained. If the confidence of the class segmentation mask of pixel point A in the first semantic segmentation result is less than the confidence threshold, the class segmentation mask of pixel point A in the first semantic segmentation result can be replaced with the class segmentation mask of the corresponding pixel point A in the second semantic segmentation result. For example, assume that the class segmentation mask of pixel point A in the first semantic segmentation result is an eye class segmentation mask with a corresponding confidence of 95%. The class segmentation mask of pixel point A in the second semantic segmentation result is an eyebrow class segmentation mask. Assuming the confidence threshold is 90%, and the confidence of the class segmentation mask of pixel point A in the first semantic segmentation result is greater than the confidence threshold, the class segmentation mask mask of pixel point A in the first semantic segmentation result can be retained, that is, the class segmentation mask of pixel point A in the first semantic segmentation result is an eye class segmentation mask. Another example, assume that the class segmentation mask of pixel point A in the first semantic segmentation result is an eye class segmentation mask with a corresponding confidence of 15%. Since the confidence of the class segmentation mask of pixel point A in the first semantic segmentation result is less than the confidence threshold, the class segmentation mask of pixel point A in the first semantic segmentation result can be replaced with the class segmentation mask of the corresponding pixel point in the second semantic segmentation result, that is, the class segmentation mask of pixel point A in the first semantic segmentation result is updated from "eye class segmentation mask" to "eyebrow class segmentation mask". By performing the above-mentioned temporal smoothing post-processing on each pixel point in the first image frame, the final semantic segmentation result of the first image frame can be obtained.
[0074] To facilitate those skilled in the art to understand the present disclosure more clearly, the following will be described in conjunction with embodiments.
[0075] For example, as Figure 3 shown, the technical solution provided by the present disclosure can be at the position of temporal smoothing post-processing in the single-frame semantic segmentation process. Temporal smoothing post-processing can be performed after the model output of the current image frame (also referred to as the current frame), in cooperation with the image and mask of the previous frame (i.e., the semantic segmentation result of the previous frame), to optimize and generate the final mask of the current image frame (i.e., the final semantic segmentation result). Exemplarily, the current image frame can be input into a semantic segmentation model to obtain the current frame mask output by the semantic segmentation model (i.e., the semantic segmentation result of the current frame). According to the current frame mask, in combination with the image of the previous frame and the previous frame mask (i.e., the semantic segmentation result of the previous frame), temporal smoothing post-processing is performed on the current frame mask, so as to obtain the optimized current frame mask, that is, the final mask of the current frame.
[0076] AsFigure 4 As shown, it shows the optimization process of post-processing the mask of a certain pixel point of an image through temporal smoothing. For each pixel point of the image frame, the process shown in Figure 4 will be traversed and judged. Optionally, the optional implementation of the above temporal smoothing post-processing process is as follows: judge the Euclidean distance between the pixel point of the current frame and the corresponding pixel point of the previous frame. If the Euclidean distance is greater than or equal to a certain threshold, it means that the pixel point at this position has changed significantly between the front and rear frames. At this time, the result of the current frame can be adopted, that is, the mask of this pixel point of the current frame is adopted (that is, the mask of this pixel point of the current frame is retained); if the Euclidean distance is less than the threshold, it means that the change range of this pixel point between the front and rear frames is small. At this time, it can be judged whether the confidence of the classification result of this pixel point in the current frame is greater than or equal to a certain threshold. If the confidence is greater than or equal to the threshold, the category of the current pixel point can use the category judgment result of the current frame, otherwise the mask result of this pixel point in the previous frame can be adopted as the mask result in the current frame.
[0077] In summary, in order to avoid the complex calculations of the temporal post-processing method, the present disclosure directly processes based on the results output by single-frame semantic segmentation, without using a multi-frame image input model, or pre-performing edge extraction or foreground extraction on the image, reducing a large amount of computational effort. The previous image frame and the semantic segmentation result are both the input and output of the model itself and can be directly called. Therefore, the present disclosure does not need to introduce or calculate additional information for processing. Since the present disclosure does not rely on additional information such as contours or foregrounds and directly performs post-processing on the output result of the model, the present disclosure can be directly deployed after any scenario and any single-frame semantic segmentation model, and can improve the temporal smoothing effect. In addition, the present disclosure improves on the traditional temporal smoothing processing method between the front and rear frames. When the single-frame semantic segmentation model outputs the semantic segmentation result of the current image frame, the previous image frame and its semantic segmentation result (that is, the semantic segmentation mask) are used to jointly adjust the semantic segmentation output result of the current image frame, so as to improve the temporal smoothing effect and avoid adding complex calculation methods to increase additional computational effort.
[0078] It should be noted that the present disclosure can significantly improve the temporal stability of semantic segmentation, including the temporal performance of all categories as a whole. Figure 5 FIG. is a comparison diagram of the effects of semantic segmentation results obtained after image processing using the benchmark method and the method provided by the embodiment of the present disclosure respectively. Among them, the definition of the benchmark method is a scheme that does not use temporal post-processing, and other settings are exactly the same. As Figure 5As shown, the method provided by the present disclosure can significantly reduce the flickering of background objects and does not affect the segmentation results of fast-moving objects. For example, it maintains the segmentation of a human body waving an arm by the model and does not perform incorrect smoothing operations on parts that should be smoothed in the time domain due to post-processing.
[0079] It should also be noted that in the post-processing stage of the image semantic segmentation task for the technical solution provided by the embodiments of the present disclosure, after the semantic segmentation model (which can also be referred to as the semantic segmentation neural network model) outputs the result (i.e., the image mask), a post-processing algorithm is used to optimize the output of the image mask. The time-domain post-processing algorithm, in cooperation with the trained semantic segmentation model, can be used in the digital graphics processing process of image acquisition devices (such as mobile phones, cameras, video cameras, tablets, etc.) during video recording and in the field of autonomous driving.
[0080] Figure 6 It is a block diagram of an image processing device shown according to an exemplary embodiment. Referring to Figure 6 , the image processing device includes: a first acquisition module 601, a second acquisition module 602, and an adjustment module 603.
[0081] Among them, the first acquisition module 601 is used to acquire a first image frame and a first semantic segmentation result of the first image frame. In some embodiments, the first acquisition module 601 is used to input the first image frame into the semantic segmentation model to obtain the first semantic segmentation result of the first image frame; wherein, the semantic segmentation model is a semantic segmentation model trained on a single-frame image.
[0082] The second acquisition module 602 is used to acquire a second image frame and a second semantic segmentation result of the second image frame. In some embodiments, the second image frame is one of the historical frames of the first image frame.
[0083] The adjustment module 603 is used to obtain a final semantic segmentation result of the first image frame according to the difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result. In some embodiments, the difference is based on the similarity between each pixel point in the first image frame and the pixel point at the same position in the second image frame.
[0084] In some embodiments, the adjustment module 603 is used to: for each pixel point in the first image frame, according to the similarity between the pixel point and the pixel point at the same position in the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result, perform corresponding processing on the class segmentation mask of the pixel point in the first semantic segmentation result to obtain the final semantic segmentation result of the first image frame.
[0085] In some embodiments, the adjustment module 603 is configured to: if the similarity between a pixel and the pixel at the same position in the second image frame is less than the similarity threshold, retain the class segmentation mask of the pixel in the first semantic segmentation result; or, if the similarity between a pixel and the pixel at the same position in the second image frame is greater than or equal to the similarity threshold, adjust the class segmentation mask of the pixel in the first semantic segmentation result based on the confidence of the class segmentation mask of the pixel in the first semantic segmentation result.
[0086] In some embodiments, the adjustment module 603 is configured to: if the confidence of the class segmentation mask of a pixel in the first semantic segmentation result is greater than or equal to the confidence threshold, retain the class segmentation mask of the pixel in the first semantic segmentation result; or, if the confidence of the class segmentation mask of a pixel in the first semantic segmentation result is less than the confidence threshold, replace the class segmentation mask of the pixel in the first semantic segmentation result with the class segmentation mask of the corresponding pixel in the second semantic segmentation result.
[0087] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0088] Figure 7 is a block diagram of an electronic device 700 shown according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, a vehicle terminal, etc.
[0089] Referring to Figure 7 , the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0090] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0091] The memory 704 is configured to store various types of data to support the operation of the device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, and the like. The memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0092] The power supply component 706 provides power to various components of the electronic device 700. The power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.
[0093] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0094] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.
[0095] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0096] The sensor assembly 714 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 700. For example, the sensor assembly 714 can detect the on / off state of the device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor assembly 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the change in temperature of the electronic device 700. The sensor assembly 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 714 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0097] The communication component 716 is configured to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0098] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0099] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 704 including instructions, is also provided. The above instructions can be executed by the processor 720 of the electronic device 700 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0100] In an exemplary embodiment, a computer program product is also provided, including a computer program that implements the steps of the above method when executed by a processor 720 of an electronic device 700.
[0101] Figure 8 FIG. 4 is a schematic structural diagram of a chip 800 proposed by an embodiment of the present disclosure. For the case where the electronic device may be a chip or a chip system, reference may be made to Figure 8 the schematic structural diagram of the chip 800 shown in FIG. 5, but not limited thereto. As Figure 8 shown in FIG. 6, the chip 800 may include one or more processors 801. The chip 800 is used to execute any of the above methods.
[0102] In some embodiments, the chip 800 further includes one or more interface circuits 802. Optionally, terms such as interface circuit, interface, and transceiver pin can be replaced with each other. In some embodiments, the chip 800 further includes one or more memories 803 for storing data. Optionally, all or part of the memories 803 may be outside the chip 800. Optionally, the interface circuit 802 is connected to the memory 803, and the interface circuit 802 can be used to receive data from the memory 803 or other devices, and the interface circuit 802 can be used to send data to the memory 803 or other devices. For example, the interface circuit 802 can read the data stored in the memory 803 and send the data to the processor 801.
[0103] In some embodiments, the interface circuit 802 executes at least one of the communication steps such as sending and / or receiving in the above method. The interface circuit 802 executing the communication steps such as sending and / or receiving in the above method means, for example, that the interface circuit 802 executes data interaction between the processor 801, the chip 800, the memory 803, or the transceiver device. In some embodiments, the processor 801 executes at least one of the other steps (such as step 101, step 102, step 103, steps 201 to 205, but not limited thereto) in the above method.
[0104] In each of the embodiments such as virtual devices, physical devices, and chips, the described modules and / or devices can be combined or separated arbitrarily according to the situation. Optionally, some or all of the steps can also be executed by multiple modules and / or devices in cooperation, which is not limited here.
[0105] Other embodiments of the present invention will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include known common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the invention are pointed out by the following claims.
[0106] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, Comprising: Obtaining a first image frame and a first semantic segmentation result of the first image frame; Obtaining a second image frame and a second semantic segmentation result of the second image frame, where the second image frame is one of the historical frames of the first image frame; Obtaining a final semantic segmentation result of the first image frame according to the difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result.
2. The method according to claim 1, characterized in that, The difference is based on the similarity between each pixel point in the first image frame and the pixel point at the same position in the second image frame.
3. The method according to claim 1 or 2, characterized in that, The obtaining the final semantic segmentation result of the first image frame according to the difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result includes: For each pixel point in the first image frame, according to the similarity between the pixel point and the pixel point at the same position in the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result, performing corresponding processing on the class segmentation mask of the pixel point in the first semantic segmentation result to obtain the final semantic segmentation result of the first image frame.
4. The method according to claim 3, characterized in that, The performing corresponding processing on the class segmentation mask of the pixel point in the first semantic segmentation result according to the similarity between the pixel point and the pixel point at the same position in the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result includes: If the similarity between the pixel point and the pixel point at the same position in the second image frame is less than a similarity threshold, retaining the class segmentation mask of the pixel point in the first semantic segmentation result; or, If the similarity between the pixel point and the pixel point at the same position in the second image frame is greater than or equal to the similarity threshold, adjusting the class segmentation mask of the pixel point in the first semantic segmentation result based on the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result.
5. The method according to claim 4, characterized in that, The adjusting the class segmentation mask of the pixel point in the first semantic segmentation result based on the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result includes: If the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result is greater than or equal to a confidence threshold, retaining the class segmentation mask of the pixel point in the first semantic segmentation result; or, If the confidence of the class segmentation mask of the pixel point in the first semantic segmentation result is less than the confidence threshold, replacing the class segmentation mask of the pixel point in the first semantic segmentation result with the class segmentation mask of the corresponding pixel point in the second semantic segmentation result.
6. The method according to claim 1, wherein Obtaining the first semantic segmentation result of the first image frame includes: Inputting the first image frame into a semantic segmentation model to obtain the first semantic segmentation result of the first image frame; Wherein, the semantic segmentation model is a semantic segmentation model trained on a single-frame image.
7. An image processing apparatus, characterized in that, Comprising: A first acquisition module, configured to acquire a first image frame and a first semantic segmentation result of the first image frame; A second acquisition module, configured to acquire a second image frame and a second semantic segmentation result of the second image frame, where the second image frame is one of the historical frames of the first image frame; An adjustment module, configured to obtain a final semantic segmentation result of the first image frame according to the difference between the first image frame and the second image frame, and according to at least one of the second semantic segmentation result and the first semantic segmentation result.
8. An electronic device, characterized in that, Comprising: One or more processors; Wherein, the processor is configured to call instructions to cause the electronic device to execute the image processing method according to any one of claims 1-6.
9. A chip, characterized in that, Comprising: One or more processors; Wherein, the processor is configured to call instructions to cause the chip to execute the image processing method according to any one of claims 1-6.
10. A storage medium, wherein the storage medium stores instructions, characterized in that, When the instructions run on the processor, the processor is caused to execute the image processing method according to any one of claims 1-6.
11. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method according to any one of claims 1-6.
Citation Information
Cited By
Image coding method and device, equipment, storage medium and program product
CN121309825A